Independent guide. Not affiliated with, endorsed by or sponsored by Google, Google DeepMind or Gemini. About this site
gemini4argon.comUnofficial guide

Guide

Connect Gemini 4 Argon to Claude with MCP

Updated · Independent coverage, not affiliated with Google

Want Claude to call on Gemini 4 Argon for a second opinion, a giant draft or a look at a video? The cleanest way is a small MCP server that gives Claude a few "ask Gemini" tools. Argon's API isn't open yet, so this guide runs today on Google's current Gemini model and switches to Argon when you change one setting.

Argon isn't in the API yet. Google has announced Gemini 4 Argon but hasn't published an API model ID. Everything below works now with gemini-3.8-flash, and moving to Argon means changing one environment variable. We track the rollout on our release date page.

Why connect Gemini to Claude?

Claude stays in charge. Gemini becomes a tool it can reach for when a job suits it:

  • A second opinion. A model from a different family reviews Claude's plan or code and tends to catch different mistakes.
  • Very long drafts. Google says Argon can write up to 1M output tokens in one response, up from 64K. That's a full report or documentation set in one pass.
  • Video understanding. Gemini accepts video, and Google reports Argon scores 91.7% on LVBench, a long-video understanding benchmark. Claude can hand it a screen recording and get a written breakdown.

How it works

MCP (Model Context Protocol) is an open standard for plugging tools into AI apps. Claude Code or Claude Desktop starts a small MCP server program in the background and talks to it over standard input and output. Ours offers four tools, each a Python function that calls the Gemini API.

Python3.10+for the MCP Python SDK
uvPython runnerinstalls packages for you
API keyGemini APIfrom Google AI Studio
ClaudeCode or Desktopeither works

Step 1: Build the Gemini MCP server

Save this as gemini_mcp.py. The comment block at the top tells uv which packages to install, so there's no separate setup. It uses the MCP Python SDK's MCPServer class (called FastMCP before version 2 of the SDK) and the Interactions API in Google's google-genai SDK.

# /// script
# requires-python = ">=3.10"
# dependencies = ["mcp>=2", "google-genai"]
# ///
import os
import time
from pathlib import Path

from google import genai
from mcp.server import MCPServer

# The one value to change when Google publishes the Gemini 4 Argon model ID.
MODEL = os.environ.get("GEMINI_MODEL", "gemini-3.8-flash")
MAX_OUTPUT = os.environ.get("GEMINI_MAX_OUTPUT")  # optional token ceiling
ROOT = Path(os.environ.get("CLAUDE_PROJECT_DIR", "."))  # set by Claude Code

client = genai.Client()  # reads GEMINI_API_KEY from the environment
mcp = MCPServer("gemini")


def _ask(prompt: str, system: str | None = None, max_tokens: str | None = None) -> str:
    args = {"model": MODEL, "input": prompt, "store": False}
    if system:
        args["system_instruction"] = system
    if max_tokens:
        args["generation_config"] = {"max_output_tokens": int(max_tokens)}
    return client.interactions.create(**args).output_text


@mcp.tool()
def ask_gemini(prompt: str) -> str:
    """Ask Gemini a question and return its answer."""
    return _ask(prompt)


@mcp.tool()
def second_opinion(work: str) -> str:
    """Get Gemini's independent review of code, a plan or a draft."""
    return _ask(work, system="You are a blunt senior reviewer. List problems first, then fixes.")


@mcp.tool()
def long_draft(brief: str, out_file: str) -> str:
    """Have Gemini write a long document, save it to out_file and return the path."""
    text = _ask(brief, max_tokens=MAX_OUTPUT)
    path = (ROOT / out_file).resolve()
    path.parent.mkdir(parents=True, exist_ok=True)
    path.write_text(text, encoding="utf-8")
    return f"Saved {len(text):,} characters to {path}"


@mcp.tool()
def watch_video(video_path: str, question: str) -> str:
    """Upload a local video file and ask Gemini about it."""
    f = client.files.upload(file=str(ROOT / video_path))
    for _ in range(60):  # wait up to about 5 minutes for processing
        if f.state and f.state.name == "ACTIVE":
            break
        time.sleep(5)
        f = client.files.get(name=f.name)
    else:
        raise RuntimeError("The video is still processing. Try again shortly.")
    return client.interactions.create(
        model=MODEL,
        store=False,
        input=[
            {"type": "video", "uri": f.uri, "mime_type": f.mime_type},
            {"type": "text", "text": question},
        ],
    ).output_text


if __name__ == "__main__":
    mcp.run()  # stdio by default, which is what Claude expects
  • Long drafts go to disk. Claude Code caps tool results at 25,000 tokens by default, so long_draft returns a file path, not the text.
  • No stored conversations. The Interactions API stores exchanges by default. One-shot tool calls don't need that, hence store=False.
  • Never print() here. Standard output carries the protocol. Use the logging module, which writes to standard error.

Step 2: Add it to Claude Code

claude mcp add --env GEMINI_API_KEY=YOUR_KEY --transport stdio gemini \
  -- uv run /absolute/path/to/gemini_mcp.py
  • The -- separates Claude Code's options from the command that starts your server. Keep --transport stdio between --env and the name, or the name gets read as another KEY=value pair.
  • This adds the server at local scope: only you, only this project, stored in ~/.claude.json. Add --scope user for every project. In a team's shared .mcp.json, never paste a key: write ${GEMINI_API_KEY} and each person's own environment fills it in.

Or add it to Claude Desktop

Open Settings from the Claude app menu, go to Developer and click Edit Config. That opens claude_desktop_config.json (in ~/Library/Application Support/Claude/ on macOS, %APPDATA%\Claude\ on Windows). Add this, then quit Claude Desktop completely and reopen it:

{
  "mcpServers": {
    "gemini": {
      "command": "uv",
      "args": ["run", "/absolute/path/to/gemini_mcp.py"],
      "env": {
        "GEMINI_API_KEY": "YOUR_KEY",
        "GEMINI_MODEL": "gemini-3.8-flash"
      }
    }
  }
}

On Windows, write paths with double backslashes. If the app can't find uv, use its full path (from which uv or where uv). Desktop has no project folder, so give the tools absolute file paths.

Step 3: Test it

  1. Run it by hand once. With GEMINI_API_KEY set, run uv run /absolute/path/to/gemini_mcp.py. The first run installs packages, then it should sit silently, waiting for a client. Press Ctrl+C.
  2. Check Claude sees it. claude mcp list should show gemini. Inside a session, /mcp shows its status and tools.
  3. Give it real work. "Get a second_opinion from Gemini on the diff you just wrote." "Use long_draft to write reference docs for this repo into docs/reference.md." "Use watch_video on demo.mp4 and list every UI glitch with timestamps."

Switching to Gemini 4 Argon

When Google publishes the Argon ID and your account has access (paid API customers are among the first), re-add the server with one more variable. --env takes several KEY=value pairs:

claude mcp remove gemini
claude mcp add --env GEMINI_API_KEY=YOUR_KEY GEMINI_MODEL=ARGON_MODEL_ID --transport stdio gemini \
  -- uv run /absolute/path/to/gemini_mcp.py

Use the exact ID from Google's models page. In Claude Desktop, edit GEMINI_MODEL and restart. For long output, also set GEMINI_MAX_OUTPUT to the most tokens you're happy to pay for. Our API page tracks the endpoint details.

Tips and caveats

  • Watch the bill. At Google's intro price of $10 per 1M output tokens, one maximum-length Argon answer costs about $10 (about $20 at the later standard price). See pricing.
  • Keys stay out of code. The script reads GEMINI_API_KEY from the environment. Never commit your Claude config files.
  • Rate limits are per project. A second key in the same project won't help. Going over returns 429 RESOURCE_EXHAUSTED: wait, then retry.
  • Know what you send. Each tool call sends that text or file to Google. Check your company's rules first.

The reverse direction

Claude Code can also run as an MCP server with claude mcp serve, which clients such as Gemini CLI can add as a stdio server. It exposes Claude Code's tools, not Claude the model, and the connecting client must confirm each tool call with you.

Claude and Gemini handle words and code. For a launch video or soundtrack, see our guide to making videos around Gemini 4 Argon.

Need visuals for what you're building?

Lumeta AI has Nano Banana 2, GPT Image 2.5, Veo 3.1, Gemini Omni and more in one account.

Try Lumeta AI

Frequently asked questions

Can I use Gemini 4 Argon in Claude Code today?

Not yet. Google hasn't published an API model ID for Gemini 4 Argon. You can set up the connection now with an MCP server running gemini-3.8-flash, then set GEMINI_MODEL to the Argon ID once it's released and your account has access.

What is an MCP server?

MCP (Model Context Protocol) is an open standard for connecting AI apps to tools. An MCP server is a small program that Claude Code or Claude Desktop runs in the background, offering tools like ask_gemini that Claude can call.

Does this work in Claude Desktop?

Yes. Add the server under mcpServers in claude_desktop_config.json (Settings, Developer, Edit Config), then fully restart the app. In Claude Code, use the claude mcp add command instead.

Will my code be sent to Google?

Only what goes into a Gemini tool call. When Claude calls ask_gemini or second_opinion, that text goes to the Gemini API. The rest of your session stays with Claude.

How much does it cost?

You pay Google for Gemini API usage by the token, on top of your Claude plan. Google's announced intro pricing for Gemini 4 Argon is $2 per 1M input tokens and $10 per 1M output tokens, rising later to $4 and $20.

Sources: Google, 9to5Google, Claude Code MCP docs, MCP local servers, MCP Python SDK, Gemini API, Gemini video understanding, Gemini rate limits, uv, Gemini CLI.