Effiziente Schätzung und Verifizierung mit Frühen Ausgängen

EP4711981 18. März 2026

Anmelder: GOOGLE LLC 🇺🇸

Details

Veröffentlichungs-Nr.
EP4711981
Anmeldetag
12. September 2025
Veröffentlichung
18. März 2026
Rechtsraum
EP
Offizieller Volltext

Abstract

One example aspect is directed to a computer-implemented method (400) for performing model decoding with reduced latency. The method includes obtaining (402) a pre-trained sequence processing model comprising a plurality of layers. The method includes modifying (404) the sequence processing model to contain an adapter layer (106) that is configured to receive and process an intermediate representation generated by a particular intermediate layer of the plurality of layers to predict an output token. The method includes training (406) the adapter layer while holding the plurality of layers of the sequence processing model frozen. The method includes deploying (408) the sequence processing model for speculative decoding in which the adapter layer, the particular intermediate layer, and the plurality of layers (104) that precede the particular intermediate layer perform speculative token decoding and the plurality of layers (108) that are subsequent to the particular intermediate layer perform token verification.

Anmelder

Firma
GOOGLE LLC
Land
🇺🇸 USA
🇺🇸 Google

US-amerikanisches Unternehmen, das Internetsuche, Onlinewerbung, Softwaredienste sowie Hard- und Software für Mobilgeräte, Cloud und künstliche Intelligenz entwickelt und betreibt.

16.097 Patente in unserer Datenbank

Vertreten von

Noch Fragen?

Wir helfen Ihnen gerne weiter. Schreiben Sie uns einfach.

Kontakt aufnehmen