Perception models combine camera, depth, inertial and location signals to construct a local scene graph and track change over time. The system prioritises relationships relevant to the selected task rather than maximising recognised objects. An uncertainty-aware attention layer decides when to sonify, when to remain silent and when to state that the scene cannot be interpreted reliably. Personalisation learns which cues a user understands and which information creates overload without changing safety boundaries invisibly. Language models may explain a scene on request, but continuous operation relies on faster structured perception and deterministic audio mappings. Blind users, orientation specialists and accessibility researchers govern what counts as useful behaviour.
Near
“Can AI turn the world around a blind person into useful, navigable sound?”
The beginning
Near is an accessibility system that translates the spatial world around a blind or low-vision person into useful, learnable sound. The ambition is not to narrate every visible object or replace a cane, guide dog or human assistance. It is to create a new perceptual layer for selected relationships that matter in motion: an open doorway, the direction of a conversation partner, an obstacle at head height, the shape of a room or the movement of a nearby cyclist. The system would be designed with blind people as an instrument that can become intuitive through use. Success means less cognitive effort and greater agency, not a more impressive description of what a camera sees.
Computer vision can label objects but labels alone do not create spatial understanding. A continuous voice describing a scene is slow, cognitively exhausting and poorly suited to urgent movement. Visual models also fail unpredictably under occlusion, unusual lighting and unfamiliar environments, while confident language can conceal uncertainty. Existing navigation systems often focus on routes and landmarks but provide limited awareness of near-field geometry and change. The challenge is to select which information is valuable, encode it without masking environmental sound and communicate uncertainty before an error becomes dangerous. Any system that demands constant attention or encourages overtrust can reduce rather than increase independence.
What it could become
Near would combine a wearable or phone-based spatial sensor with an adaptive audio engine. Users choose modes for orientation, social interaction, indoor exploration or specific tasks. Instead of naming everything, the system maps stable spatial properties to a restrained sonic vocabulary: direction, distance, opening, motion and uncertainty. A training environment lets each person learn and customise mappings safely before using them in the world. The interface provides immediate control over density and silence, preserves important environmental audio and makes system confidence perceptible. Routes and object recognition may support the experience, but the core product is a continuous, low-latency representation of nearby space co-designed around real mobility practices.
For whom
- Blind and low-vision people as paid co-designers
- Orientation and mobility specialists
- Accessibility researchers and advocates
- Cities, transit operators and mapping partners
Core capabilities
- On-device computer vision and scene understanding
- Accessible routing and live public-data fusion
- Spatial audio and haptic interaction
- Confidence-aware hazard and landmark detection
The project becomes meaningful only when a new technical possibility is translated into a clear human advantage, an experience people can understand, and a system capable of earning trust over time.
Intelligence and mathematics
The local environment is represented as a dynamic metric graph embedded in three-dimensional space. Sensor fusion estimates position and motion distributions; simultaneous localisation and mapping maintains geometry; optical flow and tracking identify moving hazards; and risk-sensitive optimisation selects a small set of cues under strict latency and cognitive-bandwidth constraints. Information theory helps choose signals that reduce spatial uncertainty most efficiently, while psychoacoustic models keep cues discriminable without masking speech, traffic or echolocation. Calibration metrics must be connected to task-level outcomes such as detection time, navigation error and cognitive load. A beautiful sonification is irrelevant if its uncertainty or latency makes action less safe.
How it might live
Near should be developed as assistive technology with blind users, orientation and mobility specialists, accessibility organisations and clinical researchers from the beginning. A first product could focus on one bounded environment or task where near-field awareness provides measurable benefit, distributed through specialist partners rather than broad consumer claims. Hardware revenue or subscription software may support continued development, with institutional programmes for training and device access. Public funding and research partnerships could keep essential functionality affordable. The durable advantage would be longitudinal co-design, safety evidence and an audio language users genuinely learn—not a generic vision model or a promise to replace established mobility tools.
For me, a venture is more than an interesting technology. It needs a narrow first user, a repeated problem, a distribution path, a credible advantage and a reason to improve as more people use it. I would test those conditions before deciding whether this idea should become a company, a product, an open technology or an ongoing research programme.
Rules for making it real
- 01
Build with blind people, never merely for them.
- 02
Preserve environmental hearing and the right to silence.
- 03
Process sensitive vision and location locally whenever possible.
- 04
Never describe uncertain hazards with false confidence.
- 05
Independence is the outcome; novelty is not.
From question to company
- 01Frame
Form a paid blind-user and mobility-specialist design council.
- 02Prototype
Prototype landmark and entrance finding with open-ear spatial audio.
- 03Prove
Pilot one repeatable outdoor journey and measure cognitive load.
- 04Build
Integrate transit, works and accessibility data with a city partner.
- 05Advance
Run safety evaluation before expanding beyond supervised pilots.
What could go wrong
Serious imagination includes the possibility that an idea should change radically—or should not exist. These are the tensions the project would need to resolve:
- A missed or invented obstacle causing physical harm.
- Audio overload masking important environmental cues.
- Location and camera data exposing intimate movement patterns.
- Unequal performance across streets, weather, lighting and mobility styles.