Gemma 4 on Android: Tricks for Faster On-Device Inference

When I tried building an on-device AI app with Gemma 4, the pitch was clear: model weights on the device, no server, no API calls, works offline. Getting it to actually run fast was a different problem.

This post covers what I learned working with LiteRT-LM 0.12.0 and Gemma 4 E2B on Android in Kotlin. Some of it is configuration. Some of it is understanding what the bottleneck actually is before reaching for a fix. If you're building with Gemma 4 E2B on Android and inference feels too slow to ship, here are the tricks that actually helped.

1. Basic Setup

Add the dependency:

// build.gradle

Gemma 4 on Android: Tricks for Faster On-Device Inference

Related reading

Gemma 4 on 16GB RAM: What Actually Works for Structured AI Workflows

I was trying to Learning About Gemma 4 and It was pretty good

The Delusion of Infinite Compute: Running Gemma 4 on an i5 CPU

Your Laptop Just Got Smarter: A Complete Guide to Gemma 4's Four Models

Gemma 4: A Practical Guide for Developers

From Cloud Dependence to Device Intelligence: How Gemma 4 is Reshaping Local AI