Skip to main content

Google Antigravity SDK Runs Local AI Agents With Gemma and Ollama

Google's Antigravity SDK now supports local agent models through Gemma with LiteRT-LM and OpenAI-compatible servers such as Ollama, LM Studio, and vLLM. The first-party Gemma configuration needs more than 24 GB of VRAM or unified memory. Local execution can reduce cloud token use, but it does not guarantee that every task stays private.

On this page

Antigravity can now point its agent harness at a local server

Google added local model support to the Antigravity SDK on 23 September. The official announcement describes two routes: a first-party Gemma configuration using LiteRT-LM, and an OpenAI-compatible configuration for servers such as Ollama, LM Studio, and vLLM.

Antigravity is an agent harness rather than a new model. It coordinates planning, tool calls, subagents, policies, and workspace access around whichever model is configured. Developers can use a local model for an entire workflow or combine local and hosted agents. The Python SDK is open source under Apache 2.0, although its platform wheels include a compiled runtime.

The OpenAI-compatible route is the more flexible option. A developer supplies a local endpoint, model name, and optional API key through LocalOpenAIAgentConfig. Compatibility with an endpoint does not guarantee that the underlying model can reliably plan, call tools, or recover from errors. Those behaviours still depend on the model, prompt format, context length, and server implementation.

Google's preferred Gemma setup is not a lightweight download

The first-party path uses Gemma 4 26B A4B in a LiteRT-LM package. Google recommends more than 24 GB of VRAM or unified memory and a 64,000-token context window for agent work. That puts the reference setup in workstation and high-memory laptop territory rather than on a typical phone or entry-level computer.

Google says other .litertlm models may run, but warns that they might not work well with the SDK. The linked model card currently describes a text-only build. It lists image and audio input, along with multi-token prediction, as future work rather than current capabilities.

The model card also reports a 15.8 GB quantised model and about 15 GB of GPU memory on a MacBook Pro with an M4 Max. Its published time-to-first-token measurement is roughly 14 seconds for a 1,024-token prefill in a short benchmark. That result is vendor-supplied, uses a much smaller context than the 64,000 tokens recommended for agents, and should not be treated as a general latency promise.

Google's demo kept most tokens local, not the whole task

In Google's demonstration, a hosted Gemini 3.8 Flash agent read file names and task descriptions, then planned work for a group of local agents. The local models audited and patched the code. Google reports that 3,322 of 3,417 tokens, or 97.2 per cent, were processed locally, while 95 tokens went to the hosted planner.

That is a useful architecture example, not an independent benchmark. The percentage comes from one vendor demonstration, and the cloud model still received information about the files and requested work. A workflow can keep most token volume on the device while sending the most revealing metadata to a hosted service.

Developers should inspect each agent configuration, tool, trace, and fallback before describing a mixed workflow as private or offline. A local endpoint can still trigger networked tools, and a hosted planner can still receive prompts even if local workers produce most of the output. Antigravity starts agents with read-only workspace access by default, but projects can grant broader tools and policies.

The SDK makes model choice practical but leaves evaluation to developers

The release gives developers one orchestration layer for switching between hosted and local agents. That can help with cost controls, offline development, sensitive code, or experiments that compare multiple serving stacks. Ollama and LM Studio offer simpler desktop setup, while vLLM is better suited to machines serving models to several clients.

The current package is labelled alpha, so configuration and behaviour can change. A serious evaluation should record time to first output, total task time, peak memory, tool-call accuracy, recovery from failed commands, and every network request. Those measurements matter more than whether the endpoint accepts an OpenAI-shaped request.

The main constraint is still the model. Antigravity can route work to a local system, but it cannot make a small model reason like a larger hosted one or make a 26-billion-parameter model fit modest hardware. The useful change is narrower: developers can now test those tradeoffs inside the same agent framework instead of rebuilding the orchestration layer for each runtime.

CuriousLM runs supported AI models locally on your device. Try CuriousLM.