MULTI-MODAL MACHINE LEARNING ARCHITECTURES INTEGRATING LANGUAGE MODELS AND COMPUTER VISION SYSTEMS

Patent №

US 11,803,710

Granted

2023-10-31

Filed 2023

Owner

SURGETECH, LLC

Lab

AI components

6

ml · nlp · vision · speech · kr · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

18191746

Improved multi-modal machine learning networks integrate computer vision systems with language models. In certain embodiments, a computer vision system analyzes at least one image to generate a computer vision output. The language model generates an output based, at least in part, on a consideration of the computer vision output. The outputs of the language model can be generated by jointly considering textual information learned by the language model and visual content extracted by the computer vision system, thereby significantly improving the accuracy, breadth, and comprehensiveness of the outputs.

Machine learningNatural languageVisionSpeechKnowledge representationAI hardwareG06F 16/583G06F 40/35G06F 16/532G06F 40/40G06V 10/764G06V 10/774G06V 10/803G06V 10/82+3 more

AI classification

Natural language1.00
Speech1.00
Machine learning1.00
Vision1.00
Knowledge representation1.00
AI hardware0.98
Planning0.00
Evolutionary computation0.00

Ownership

SURGETECH, LLC

assignment · 632310630

Assignors

LOVE, MICHAEL, LOVE, BLAKE, SOROMENHO, TIAGO

On an employer assignment, the assignors are typically the inventors.

From the same owner

© 2026 NYSGPT2525 LLC