The CBS algorithm enabled binocular vision processing and captured information on depth, surface slant, and object pose (representational image).Figure/YouTube Researchers have developed a new vision algorithm that lets humanoid robots perceive depth more like humans, using two forward-facing eyes. Instead of relying on conventional stereo cameras, the system processes visual differences between the left and right images to build a more accurate 3D understanding of a scene. The approach, developed by a team led by researchers from York University in Toronto, could improve how humanoid robots judge distances, recognize object orientation, and navigate complex environments. Researchers say the advance brings robotic vision closer to biologically inspired human depth perception and lays the groundwork for more capable humanoid visual systems. Robot depth vision The new stereo vision algorithm allows humanoid robots to calculate depth using a camera arrangement that more closely mimics human vision. Called convergent binocular stereo (CBS), it is designed for robotic heads whose two cameras can point toward the same location instead of remaining parallel. Unlike conventional stereo systems, which mainly calculate horizontal differences between images from parallel cameras, CBS accounts for both horizontal and vertical disparities created when cameras converge on a target. This allows the system to use camera movement as part of the depth-perception process. Areas with different horizontal disparities are marked, with solid lines showing camera views and dotted lines showing axes. The algorithm first fixes both cameras on a 3D point and determines the geometric relationship between the resulting images. It then creates a five-level Gaussian pyramid to analyze the images at multiple resolutions. At coarser levels, CBS detects and matches SIFT features while using epipolar constraints to estimate horizontal and vertical disparities. Gabor filters are then used to analyze texture and edge information at different orientations and scales, helping identify corresponding pixels between the two images. The system progressively refines these disparity estimates from lower-resolution images to the original resolution, narrowing the search area at each stage. The resulting horizontal and vertical disparity maps can then be converted into 3D information and depth estimates. The researchers say CBS can also capture information about surface slant, object pose, and the perceived scaling of depth and size, providing a more humanlike approach to robotic visual processing. CBS boosts accuracy To test the algorithm, the researchers created the Convergent Binocular Stereo–BenchMark dataset, or CBS-BM. It contains 49 scenes, including ordinary tabletop environments as well as repeated patterns, featureless surfaces and highly self-occluded objects. Each scene includes parallel images, along with between five and 12 convergent fixation image pairs. CBS was broadly competitive with established stereo methods and several deep-learning systems. Across the full dataset, it achieved lower mean horizontal disparity and depth errors than the parallel methods tested. Its strongest advantage appeared in scenes containing repeated patterns, where conventional systems can struggle to determine which visually similar features correspond to each other. In that subset, CBS produced a mean depth error about 0.8 metres lower than the next-best method and a horizontal disparity error about 100 pixels lower. According to a statement by the team, the approach does have limitations. Its accuracy decreases for targets at greater distances, and it depends on precise camera calibration and motor positioning. The current implementation also takes about 69 seconds to run on an AMD Ryzen 7 7700X processor, leaving substantial scope for optimization. The researchers say CBS is not intended to replace parallel stereo in applications where humanlike vision is unnecessary. Instead, it provides a foundation for humanoid robots that need their physical eye movements and visual processing to work together, bringing robotic depth perception closer to the active, convergent vision humans use. Get the latest in engineering, tech, space & science - delivered daily to your inbox.Jijo is an automotive and business journalist based in India. Armed with a BA in History (Honors) from St. Stephen's College, Delhi University, and a PG diploma in Journalism from the Indian Institute of Mass Communication, Delhi, he has worked for news agencies, national newspapers, and automotive magazines. In his spare time, he likes to go off-roading, engage in political discourse, travel, and teach languages.
New stereo vision algorithm gives humanoid robots more humanlike depth perception
Full Article
Original Source
Read the full article at Interestingengineering →KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.