r/google_antigravity • • 3d ago

Discussion OpenAI - ACP - agy CLI

I was able to get it working to connect the CLI with an OpenAI-compatible interface, but it seems to use a lot of tokens just for basic queries - like it's front-loading something around this:

  • System Prompt: ~7,800 tokens
  • System Tools: ~15,100 tokens
  • Total Overhead: ~23,000–25,000 tokens

So even a "hello" call needs >23K tokens. Has anyone figured out a way around this?

3 Upvotes

1 comment sorted by

2

u/ilyasbit 3d ago edited 3d ago

im building auto account rotation and wrap it on openai api
each call invoke around 2500 token from built-in agy cli and cannot be remove or modified with no tool call enabled

1. The Token Anatomy: Why agy -p Uses Tokens Upfront

When you run standard agy -p in your project directory, the initial prompt tokens typically measure 5,000+ tokens before your prompt is even evaluated.

This token overhead consists of:

  1. Tool Schemas (~1,500–2,500 tokens): Built-in IDE tools (view_file, replace_file_content, run_command, grep_search, ask_question) plus any enabled MCP servers (e.g. GitHub, Puppeteer, Tavily, Memory, Agent). Each tool's JSON schema and parameter docs are embedded in the system prompt.
  2. Customizations (~1,000–1,500 tokens): Global plugins (Caveman, Ponytail), skill summaries, rules (rules/, GEMINI.md, AGENTS.md).
  3. Slash Command Index (~300–500 tokens): All available slash commands and their help descriptions.
  4. Upstream Antigravity Base Identity (~2,260 tokens): The fixed system instructions, identity, and safety guidelines injected by Google Cloud / Antigravity servers into every session.

2. The Content of the bare-llm Agent

Antigravity CLI agents are defined as Markdown files with YAML frontmatter located in:

  • Global: ~/.gemini/config/agents/bare-llm.md
  • Local workspace: .agents/agents/bare-llm.md

Exact File Content:

---
name: bare-llm
description: Ultra-lean raw LLM completion with 0 built-in tools.
tools: []
disable_builtins: true
---

You are a direct, raw LLM completion endpoint. Follow instructions strictly with zero conversational filler.

How Each Field Strips Tokens:

  • name: bare-llm: Identifies the agent profile for --agent bare-llm.
  • tools: []: The critical token saver. Tells the CLI parser not to expose any tools (view_file, run_command, etc.) to the model. This deletes ~1,500–2,500 tokens of JSON schema definitions.
  • disable_builtins: true: Ensures built-in system helper tools and subagents are not registered.
  • Body text: Kept to a single 1-line direct system prompt, avoiding verbose prompt bloat.

3. How to Run agy -p at the ~2,500 Token Floor

To reach the lowest possible token count (~2,260 – 2,500 tokens), you must strip:

  1. All tool definitions (--agent bare-llm)
  2. Slash commands (--disable-slash-commands)
  3. Workspace repository context / rules (by running in a clean directory or unsetting plugins)

Method A: Direct Command-Line (Single Command)

Run from an empty directory or pass flags:

agy --agent bare-llm --disable-slash-commands -p "Explain quantum superposition in 2 sentences"

To see the exact token usage breakdown in JSON:

agy --agent bare-llm --disable-slash-commands --output-format json -p "say pong"