SYSTEM AND METHOD FOR EFFICIENTLY AMALGAMATED CNN-TRANSFORMER ARCHITECTURE FOR MOBILE VISION APPLICATIONS

Patent №

US 12,373,672

Granted

2025-07-29

Filed 2022

Owner

Mohamed bin Zayed University of Artificial Intelligence

Lab

AI components

0

Assignment

None on record

Dataset

AIPD

Application

18078657

An edge computing system, computer readable storage medium and method for object detection, including processing circuitry. The processing circuitry is configured with a hybrid CNN and vision transformer backbone network in an object detection deep learning network. The backbone network receives an image, and includes a first convolutional encoder to extract local features from feature maps of the image, a second stage having consecutive second convolutional encoders, a positional encoding layer, split depth-wise transpose attention (SDTA) encoders, consecutive convolutional encoders, a third stage and a fourth stage SDTA encoder. Each of the SDTA encoders perform multi-headed self-attention by applying a dot product operation across channel dimensions in order to compute cross-covariance across channels to generate attention feature maps. The object detection neural network includes a convolutional network that produces a fixed-size collection of bounding boxes and scores for a presence of object class instances in those boxes.

G06V 10/82G06N 3/0464G06V 10/25G06V 10/454G06V 10/764G06V2201/07G06N 3/045G05D 1/0221+4 more

Ownership

Mohamed bin Zayed University of Artificial Intelligence

From the same owner

© 2026 NYSGPT2525 LLC