Multimodales Grosssprachenmodell mit Audioauslöser
Anmelder: INTEL Corporation 🇺🇸
Details
- Veröffentlichungs-Nr.
- EP4718321
- Anmeldetag
- 16. Juli 2025
- Veröffentlichung
- 1. April 2026
- Rechtsraum
- EP
Abstract
Systems and methods to trigger LLM inference based on the presences of relevant audio, such as a keyword or sound event of interest. A detection head receives acoustic embeddings from an audio encoder and determines whether the audio stream includes relevant sounds (e.g., a selected audio trigger). When the audio stream does not include relevant sounds, multimodal LLM inference is bypassed, thereby saving power and protecting privacy. When relevant sounds are detected in the audio stream by the detector, the acoustic embeddings from the audio encoder are transmitted to the multimodal LLM, which proceeds to perform inference on the acoustic embeddings. The audio encoder and/or detection head can be offloaded in the hardware and implemented before the multimodal LLM in the hardware pipeline, while the multimodal LLM can be implemented in a neural processing unit.
Anmelder
- Firma
- INTEL Corporation
- Land
- 🇺🇸 USA
US-amerikanischer Halbleiterhersteller mit Sitz in Kalifornien. Entwickelt und produziert Prozessoren, Chipsätze sowie weitere Halbleiterkomponenten für Computer, Server und eingebettete Systeme.
6.136 Patente in unserer Datenbank
Vertreten von
-
Goddar, Heinz J.