Aura LiteRT-LM
Mobile-QAT deployment for supported Android, mobile, and edge hardware.
View model ↗Base tier · Three deployment formats
Aura is a Gemma 4 E4B model family built around one canonical combined adapter. Choose the runtime that fits your device without changing the identity at the center of the experience.
Aura is designed to be the companion in the palm of your hand. The base model is perfect for edge device deployment on all device tiers.
Choose a format
Each release uses the same canonical combined Aura adapter and is validated independently for its target runtime.
Mobile-QAT deployment for supported Android, mobile, and edge hardware.
View model ↗Portable local inference for llama.cpp-compatible desktops, laptops, and edge systems.
View model ↗High-fidelity Transformers inference and evaluation for GPUs with sufficient memory.
View model ↗