US 12,393,197 B2
Systems and methods for object detection using a geometric semantic map based robot navigation
Snehasis Banerjee, Kolkata (IN); Brojeshwar Bhowmick, Kolkata (IN); and Ruddra Dev Roychoudhury, Kolkata (IN)
Assigned to Tata Consultancy Services Limited, Mumbai (IN)
Filed by Tata Consultancy Services Limited, Mumbai (IN)
Filed on Oct. 26, 2022, as Appl. No. 18/050,015.
Claims priority of application No. 202221003604 (IN), filed on Jan. 21, 2022.
Prior Publication US 2023/0236606 A1, Jul. 27, 2023
Int. Cl. G05D 1/00 (2024.01)
CPC G05D 1/0274 (2013.01) [G05D 1/0214 (2013.01); G05D 1/0238 (2013.01); G05D 1/0246 (2013.01)] 7 Claims
OG exemplary drawing
 
1. A processor implemented method, comprising:
acquiring, via an image capturing device mounted on a robotic device executed by one or more hardware processors, a set of successive egocentric images corresponding to one or more views of one or more scenes in an action space based-indoor environment, wherein the action space based-indoor environment comprises a target object to be detected;
retrieving from a knowledge store comprised in a system database, via the one or more hardware processors, a semantic relational knowledge representation, maintained in an ontological form, for each of the egocentric images from the set of successive egocentric images corresponding to the one or more views of the one or more scenes in the action space based-indoor environment, wherein the knowledge store comprises a plurality of informative data associated with (i) actuation capability of the robotic device during navigation, and (ii) the one or more scenes in the action space based-indoor environment, wherein the informative data associated with the one or more scenes in the action space based-indoor environment comprises a plurality of objects, relationship between the plurality of objects, one or more obstacle objects, one or more restricted areas, and one or more landmark points present in the one or more scenes;
generating, via the one or more hardware processors, a geometric semantic map for the action space based-indoor environment based on the semantic relational knowledge representation and a geometrical movement monitoring of the robotic device;
iteratively performing, via the robotic device executed by the one or more hardware processors, a navigation step based on the geometric semantic map until at least one of (i) the target object is detected, and (ii) geometric semantic map based analysis of the scene is completed, wherein at each of the navigation steps, a plurality of attributes comprising a relative direction of the one or more views, one or more objects in the scene with a confidence score, a landmark score computed for the navigation step, are stored, wherein the Landmark score is computed in accordance with a) probability of the target object being found in a scene, and b) a combined probabilities of objects located in the scene and their relations to the target object, wherein the landmark score is computed based on:
rotation0 indicative of rotational moves in a zone to determine whether a scene in the zone explored in totality,
identifying presence of occlusion objects hiding the target object based on dimensional relationship;
navigating to free space to go around identified occlusion to bring the target object in view;
wherein the scene is broken into a one or more polygonal segments,
wherein the step of detecting the target object from the one or more objects comprised in the scene and having a maximum relation probability value computed based on distinct co-location probabilities and distinct edges between objects to the target object is performed by determining a parameter computed as a product of (i) the confidence score of the target object, and (ii) a centroid value computed for the target object and the one or more objects with maximum probable relation to the target object, wherein decision to move towards the polygonal segment is based on the centroid having maximum summation of relation probability by combination of objects belonging to the polygon segment; and
dynamically updating, via the one or more hardware processors, the geometric semantic map and the knowledge store with information of analyzed scenes learned over a period of time at each navigational steps.