New AI to Assist The Visually Impaired
By Amos Miller, founder of Buzzbo · 12 October 2026 · 4 min read
Google’s new AI can describe the world to blind people in real time — but Google says not to trust it as a guide
A blind person points their phone towards a kitchen shelf and asks where the oregano is. Instead of simply naming objects somewhere in the camera frame, the phone can tell them to move the camera right, tilt it down or step backwards until the object is visible. They can then ask a follow-up question about it without beginning the process again.
That is the idea behind Guided Vision, a new accessibility feature in Google's Gemini Live that launched on 1 October.
The development is noteworthy not merely because artificial intelligence can describe an image — technology has been doing versions of that for years — but because Google is attempting to turn visual interpretation into a continuous conversation between the user, their camera and an AI system.
Guided Vision allows someone to share their Android phone camera during a Gemini Live conversation and receive spoken descriptions of what it sees. Users can ask it to read fine print, identify an object, describe colours or patterns, inspect a room or answer more specific questions about something in view. If the camera is badly positioned, the system can provide verbal instructions to help the user reframe the shot.
For a blind or low-vision user, that seemingly small difference can be important. Conventional image recognition assumes the user can point the camera at the thing they want recognised. But if you cannot see the screen, knowing whether the label, object or document is actually inside the frame is itself an accessibility problem.
Google says Guided Vision was developed in partnership with visual-interpreting company Aira. According to Google's account, Aira helped visually interpret tens of thousands of hours of material used to train Gemini Live for real-world conversational context, while more than 1,000 members of Aira's Trusted Tester network tested and refined the system across everyday situations. Specialists from Aira also worked with Google's teams on safety guardrails.
Testing and feedback involved blind and low-vision communities in countries including India, Brazil, Singapore, Indonesia and Japan. That international involvement is significant for a technology that needs to interpret not only different languages but different products, environments and everyday contexts.
The feature is now available on Android 9 and later in regions and languages where Gemini Live is supported. It can be started from Gemini itself, assigned to an Android accessibility shortcut or opened directly through TalkBack, Google's Android screen reader.
There is considerable potential here for independence. Reading an expiry date, distinguishing two similarly shaped food packages, checking clothing colours, examining unfamiliar controls or making sense of a restaurant menu are mundane activities precisely because sighted people usually perform them without assistance. For somebody who has previously needed another person or a remote visual interpreter, having another tool immediately available on the phone can change the degree of privacy and spontaneity involved.
But the most important part of Google's announcement may be its warning about what Guided Vision cannot safely do.
Google explicitly says the generative-AI system can make mistakes. It is not a medical device, a mobility aid or a replacement for a white cane. The company says it should not be used for navigation, safe-travel guidance or obstacle detection, and tells users to continue relying on established mobility aids and safe-travel practices.
That boundary is crucial.
There is an enormous difference between an AI incorrectly identifying a tin of tomatoes and incorrectly deciding that a road is clear, a staircase has ended or an obstacle is absent. The first is irritating; the second can be dangerous. Accessibility technology powered by generative AI therefore creates an unusual trust problem: the more useful and natural the system feels, the easier it may become to overestimate what it reliably knows.
Google's own warning is an acknowledgement of that problem rather than a solution to it.
There are other unresolved questions that real-world use will answer better than a product demonstration can. Performance can vary with poor lighting, clutter, moving subjects, camera quality and unusual objects. Generative AI systems can also produce confident errors. How consistently Guided Vision recognises when it lacks enough information — rather than filling the gap with an incorrect answer — will be particularly important for users relying on it repeatedly during everyday life.
Privacy deserves attention too. Using the feature necessarily means pointing a camera at surroundings that can include other people, private documents and personal spaces. Google's launch announcement explains how the system was developed and tested, but does not provide a detailed account there of every data-retention question a privacy-conscious user might have about individual camera sessions. That is an area users may reasonably want to examine in Google's applicable Gemini privacy documentation before deciding how and where they use it.
None of those cautions makes the technology insignificant. In fact, the way it was developed offers a useful lesson for the wider technology industry.
Google did not merely announce that an AI product happened to be useful to blind people. It says blind and low-vision people were involved in stress-testing it at substantial scale, while accessibility specialists contributed to its safety design. That is a much stronger model than building a mainstream product first and asking disabled users whether it works afterwards.
The real measure of Guided Vision will now move from the laboratory to people's kitchens, workplaces, shops and homes.
If it works reliably within the limits Google has set, it could give blind and low-vision Android users another powerful layer of independent access to visual information. If users begin trusting it beyond those limits, the same conversational fluency that makes it useful could become its greatest risk.
Accessible technology does not need to be infallible to be valuable. It does, however, need to be candid about what it can and cannot be trusted to do.
On that point, Guided Vision arrives with both considerable promise and an unusually important warning label.
