Optimierung der Quantisierung in Aufmerksamkeitsbasierten Neuronalen Netzen durch Straffung der Verteilungen von Eingangstensoren auf Offline-Transformationen

EP4745845 20. Mai 2026

Anmelder: Axelera AI BV 🇳🇱

Details

Veröffentlichungs-Nr.
EP4745845
Anmeldetag
14. November 2024
Veröffentlichung
20. Mai 2026
Rechtsraum
EP
Offizieller Volltext

Abstract

The invention is notably directed to a computer-implemented method of optimising quantisation in a neural network (40) having an attention mechanism. The dataflow architecture of the neural network includes quantisation stages (41<sub>Q</sub> - 47<sub>Q</sub>) downstream of respective transformation stages, which involve transformations that can be represented through transformation matrices and include offline transformations. The method first comprises optimising a stage (41, 42<sub>3</sub>) of the transformation stages for subsequent quantisation by a respective one of the quantisation stages (41<sub>Q</sub>, 42<sub>Q</sub>), by optimising an offline transformation involved in said stage in an unsupervised manner, the offline transformation representable through a dot product of an offline input tensor and an offline transformation matrix, using an objective function designed to reduce a degree of tailedness of a distribution of tensor components of a tensor resulting from said dot product. Next, said stage is precomputed at least partly by performing the optimised offline transformation as an affine transformation based on said dot product, and stored to ready the neural network for inferences. The invention is further directed to related methods and computer programs.

Anmelder

Firma
Axelera AI BV
Land
🇳🇱 Niederlande

Vertreten von

Noch Fragen?

Wir helfen Ihnen gerne weiter. Schreiben Sie uns einfach.

Kontakt aufnehmen