System und Verfahren für Maschinenlernmodelle zur Computersicht auf Vorrichtungen
Anmelder: STMicroelectronics International N.V. 🇨🇭
Details
- Veröffentlichungs-Nr.
- EP4715753
- Anmeldetag
- 22. September 2025
- Veröffentlichung
- 25. März 2026
- Rechtsraum
- EP
Abstract
A system and method are provided for implementing transformer-based computer vision models on resource-constrained devices. An input image is divided into tokens, each corresponding to a patch. A background-aware vision transformer (BAViT) classifies tokens as foreground or background using a lightweight architecture without a class token and with a linear classifier for token-wise prediction. Training utilizes an accumulative cross entropy loss that aggregates token-level losses to improve accuracy. Tokens classified as background are pruned, thereby reducing computational complexity, runtime memory, and inference latency. Foreground tokens are processed in a downstream transformer-based object detection model, such as YOLOS, to generate detection outputs. The BAViT module operates as a pre-processing stage, facilitating integration with detection models without retraining. Configurations include BAViT-small with two transformer layers suitable for edge devices, supporting applications such as security and inventory tracking.
Anmelder
- Firma
- STMicroelectronics International N.V.
- Land
- 🇨🇭 Schweiz
Schweizer Gesellschaft des Halbleiterkonzerns STMicroelectronics, die Chips und Sensoren für Automobiltechnik, Industrie und Unterhaltungselektronik entwickelt.
733 Patente in unserer Datenbank