Method and system for visio-linguistic understanding using contextual language model reasoners

Patent №

US 11,699,275

Granted

2023-07-11

Filed 2021

Owner

TATA CONSULTANCY SERVICES LIMITED

Lab

AI components

6

ml · nlp · vision · speech · kr · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

17349440

This disclosure relates generally to visio-linguistic understanding. Conventional methods use contextual visio-linguistic reasoner for visio-linguistic understanding which requires more compute power and large amount of pre-training data. Embodiments of the present disclosure provide a method for visio-linguistic understanding using contextual language model reasoner. The method converts the visual information of an input image into a format that the contextual language model reasoner understands and accepts for a downstream task. The method utilizes the image captions and confidence score associated with the image captions along with a knowledge graph to obtain a combined input in a format compatible with the contextual language model reasoner. Contextual embeddings corresponding to the downstream task is obtained using the combined input. The disclosed method is used to solve several downstream tasks such as scene understanding, visual question answering, visual common-sense reasoning and so on.

Machine learningNatural languageVisionSpeechKnowledge representationAI hardwareG06V 10/25G06F 16/5846G06F 40/20G06N 3/04G06N 3/0464G06N 3/09G06V 10/764G06V 10/82+3 more

AI classification

Natural language1.00
Vision1.00
Machine learning1.00
Speech1.00
AI hardware0.99
Knowledge representation0.87
Planning0.09
Evolutionary computation0.00

Ownership

TATA CONSULTANCY SERVICES LIMITED

assignment · 565660478

Assignors

KURMA, SAI SREE BHARGAV, KALRA, KANIKA, VADAKKEEVEETIL SREELATHA, SILPA, PATWARDHAN, MANASI, KARANDE, SHIRISH SUBHASH

On an employer assignment, the assignors are typically the inventors.

From the same owner

© 2026 NYSGPT2525 LLC