Concepts
Providers & Models
The core registry includes OpenAI OAuth and API, Grok OAuth and API, Gemini/Antigravity CLI, direct Gemini API, AtlasCloud, MiniMax, NovelAI, and registered ComfyUI workflows. Runway and Higgsfield are separate MCP-backed integrations described below.
Provider paths
provider: "oauth"— calls ChatGPT's Codex backend from the server with your ChatGPT session. A GPT-6 model plans the prompt andgpt-image-2renders it. The default path; no API key needed.provider: "api"— calls the OpenAI Responses API with the hostedimage_generationtool. RequiresOPENAI_API_KEY.provider: "grok"— callsapi.x.aidirectly with the xAI OAuth session in~/.progrok/auth.json, runs mandatory xAI Web Search and a planner pass (defaultgrok-4.3, configurable in settings or via--planner-model), then calls xAI Images API. Grok 4.5 and 4.6 are selectable.provider: "grok-api"— same xAI pipeline asgrokbut uses a directXAI_API_KEYinstead of the OAuth session.provider: "gemini-api"— calls the Google Gemini image API directly. Supports two models (nano-banana-2andnano-banana-pro), aspect ratio and resolution controls, and two auth modes: aGEMINI_API_KEYor a Vertex AI service-account JSON (VERTEX_SERVICE_ACCOUNT_JSON). When both are configured, Vertex AI takes priority unless overridden by the last-saved auth mode. Cost varies by model and resolution (see Gemini API section below).provider: "agy"— spawns the Antigravity CLI (agy -p) to generate via Google Gemini (nano-banana-2). Fixed 1024×1024 JPEG output, max 3 refs. Free (no token cost).provider: "atlascloud"— calls AtlasCloud withATLASCLOUD_API_KEYfor GPT Image 2 generation and edit, with up to 10 references.provider: "minimax"— calls MiniMax withMINIMAX_API_KEYforimage-01andimage-01-live, with one reference.provider: "nai"— calls NovelAI with a persistent token. It is text-to-image only and exposes a dedicatednegativePromptplus native sampling controls; references and edits fail closed.provider: "comfy"— executes a registered ComfyUI image workflow; each workflow is a runtime-discovered model.
The core lanes cover Classic, Node, and Agent Mode. Agent Mode is web-UI only; Classic and Node also have CLI commands.
Per-request override
Generation commands use the explicit core lane IDs below. Where a command still accepts auto, it is a routing mode, not a provider lane.
| Value | Behavior |
|---|---|
auto | Routing mode where accepted; not a provider lane. |
oauth | Force the GPT OAuth path. |
api | Force the API-key Responses path; requires a configured key. |
grok | Force the xAI OAuth path to api.x.ai; run ima2 grok login once to authorize. |
grok-api | Same xAI pipeline but authenticates with a direct XAI_API_KEY. |
gemini-api | Direct Google Gemini API. Requires GEMINI_API_KEY or VERTEX_SERVICE_ACCOUNT_JSON. Supports aspect ratio and resolution controls. No quality/format/moderation/multimode controls. |
agy | Spawn Antigravity CLI for Gemini image generation. Requires agy binary installed. 1024×1024 fixed, JPEG, max 3 refs, no quality/size/mask controls. |
atlascloud | Direct AtlasCloud API. Requires ATLASCLOUD_API_KEY. |
minimax | Direct MiniMax API. Requires MINIMAX_API_KEY. |
nai | Direct NovelAI text-to-image lane. Requires NOVELAI_API_KEY; references and edits are unsupported. |
comfy | Registered ComfyUI image workflow. Availability and models come from the runtime workflow catalog. |
Prompt Builder backend
Prompt Builder routing is independent from the current image provider. Settings > Providers
offers Auto plus GPT OAuth, OpenAI API, Grok, and Grok API, along with a backend-scoped Builder
model. Auto chooses the first ready backend in the order oauth → grok → api → grok-api.
An explicit selection pins routing and returns a typed error, such as a missing-key error, instead
of falling back. The via <backend> badge reports the answering backend.
Models
The GPT OAuth lane defaults to gpt-6-luna; the API-key lane keeps gpt-5.6-luna. Prompt Builder has a separate backend
preference; when its backend resolves to GPT, Luna is the default GPT Builder model.
| Model | Use |
|---|---|
gpt-6-luna | GPT OAuth default and default GPT Builder model when the Builder uses GPT OAuth. |
gpt-6-sol | The other everyday GPT-6 model on GPT OAuth. |
gpt-6-astra | Reasons longest and is the slowest. Selectable on the OAuth and API lanes. |
gpt-5.6-luna / gpt-5.6-terra / gpt-5.6-sol | API-key lane models (gpt-5.6-luna is its default). On GPT OAuth these ids map to gpt-6-sol (sol) or gpt-6-luna. |
gpt-5.5 / gpt-5.4 / gpt-5.4-mini | API-key lane compatibility models; on GPT OAuth they run as gpt-6-luna. |
grok-imagine-image | Compatible fast Grok image model. |
grok-imagine-image-quality | Highest-quality Grok image model ("Grok+" / Best in the UI). |
grok-imagine-image-2.0 | Default Grok image model for new sessions. |
grok-imagine-video | Base model for Ref2V, edit, and extension compatibility paths. |
grok-imagine-video-1.5 | Default Grok video model for prompt-only T2V and single-image/frame I2V, including 1080p when supported. |
nano-banana-2 | Gemini Flash image model (maps to gemini-3.1-flash-image). Used by both gemini-api and agy providers. |
nano-banana-pro | Gemini Pro image model (maps to gemini-3-pro-image). Available on the gemini-api provider only. |
The app also exposes quality (low, medium, high) and
moderation (auto, low) controls. Reasoning effort accepts none, low,
medium, high, and xhigh.
ima2 defaults set model gpt-5.5 and
ima2 defaults set reasoning high write both OAuth and API-provider default keys, so
your "default model" stays one concept across provider paths.
Gemini API provider
The gemini-api provider calls Google's image generation API directly without the
Antigravity CLI. It supports two models selectable in the UI or via --model:
nano-banana-2(default) — Gemini Flash; faster, lower cost.nano-banana-pro— Gemini Pro; higher quality, higher cost.
Auth modes. Configure either a Gemini API key (GEMINI_API_KEY env
var or via the Settings UI) or a Google Cloud service-account JSON
(VERTEX_SERVICE_ACCOUNT_JSON env var or via the Settings UI Vertex JSON input).
When both are present the server respects the last-saved mode
(geminiAuthMode: "apikey" or "vertex" in config), defaulting to Vertex
when both are configured and no mode is saved.
Aspect ratio and resolution. The direct API path exposes 10 aspect ratios
(1:1, 2:3, 3:2, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, 21:9) and 4 resolution tiers
(512px, 1K, 2K, 4K). Selections are mapped to exact pixel dimensions and passed as
aspect_ratio and image_size protobuf enum values.
The Vertex AI path ignores these controls — the Vertex endpoint does not accept
the response_format field, so output defaults to 1K / 1:1 regardless.
Per-image cost estimates (based on output token counts × official rates):
| Model | 512px | 1K | 2K | 4K |
|---|---|---|---|---|
nano-banana-2 (Flash 3.1, $60/1M tok) | $0.045 | $0.067 | $0.101 | $0.151 |
nano-banana-pro (Pro 3, $120/1M tok) | $0.134 | $0.134 | $0.134 | $0.240 |
The agy provider (Antigravity CLI) uses nano-banana-2 only, at fixed
1024×1024 with no resolution or aspect controls, and is estimated as free (no API token charge
in the cost estimator).
Grok pipeline
Grok Classic, Node, and Agent requests run a three-step pipeline: mandatory xAI Web Search,
planner pass (default grok-4.3, overridable via
IMA2_GROK_PLANNER_MODEL, settings UI, or --planner-model on video
commands) with an English final image prompt, then xAI image creation. Grok 4.3 remains available as a compatibility override.
Text-only requests use /v1/images/generations; requests with reference images,
a Node parent image, or an Agent current image use /v1/images/edits so image-to-image
context is preserved. Grok accepts up to three total input images in this path.
ima2 maps OpenAI-style sizes to xAI aspect_ratio and resolution
controls. Grok mask edit is not wired in this release and returns
GROK_MASK_UNSUPPORTED.
Model and size pickers. The UI exposes a two-button image model picker
("Grok" / Fast = grok-imagine-image; "Grok+" / Best = grok-imagine-image-quality)
and a size picker with native xAI aspect_ratio and resolution (1k/2k) values.
Billing and quota. When Grok is authorized, GET /api/quota returns
a grok object with a monthly usage bar and a
billing field (usedUsd / limitUsd) displayed as
"$used/$limit" in the QuotaCard header (e.g. "$134.80/$1500.00").
Switch Account. The QuotaCard exposes a "Switch Account" button for Grok that
starts a server-side xAI device-code flow
(POST /api/auth/switch → GET /api/auth/switch/:sessionId). The button
opens the verification URL in a new tab, displays the user code, and the server polls until
complete. The same flow is available for Codex/GPT OAuth.
Grok video
Grok video generation defaults to canonical grok-imagine-video-1.5 ("Grok V1.5").
The base grok-imagine-video model remains available for Ref2V, edit, and extension. A two-button video model
picker at the top of the video controls panel lets you switch between them. Three modes are
auto-detected from reference count: text-to-video (0 refs), image-to-video (1 ref), and
reference-to-video (2–14 refs; max 15s on 1.5, 10s on base). Controls include duration (1–15s), resolution
(480p, 720p, and 1080p for 1.5 single-image/frame I2V), and aspect ratio. The old
grok-imagine-video-1.5-preview value is accepted as a compatibility alias.
1.5 does not add Ref2V, V2V edit, or extension support, so those routes remain
base-model only. Choose the planner model in video settings or pass
--planner-model on CLI video commands.
The endpoint POST /api/video/generate streams SSE events: planning → submitted →
progress → done. From the CLI: ima2 video "prompt" --model grok/grok-imagine-video-1.5 --duration 5 --resolution 720p
(or set a persistent target once with ima2 defaults set video grok/grok-imagine-video-1.5).
MCP lanes (Runway, Higgsfield)
Since 3.0.0 the CLI routes two additional lanes through remote MCP providers:
runway (Gen-4/4.5, Veo 3.1, Seedance 2, Kling — image and video) and
higgsfield (catalog-only until a paid plan unlocks execution). Inspect live lane
status and per-model capabilities with ima2 models, then target them like any
other lane: ima2 video "prompt" --model runway/veo-3.1 --duration 8. MCP jobs are
asynchronous — the CLI submits to POST /api/mcp/generate and waits on the shared
SSE event stream until the file lands in the gallery.
API-provider defaults
When the API path is used without explicit options, these defaults apply:
| Variable | Default |
|---|---|
IMA2_API_IMAGE_MODEL_DEFAULT | gpt-5.6-luna |
IMA2_API_REASONING_EFFORT | low |
IMA2_API_IMAGE_SIZE | 1024x1024 |
IMA2_API_ALLOW_WEB_SEARCH | true |
See Configuration for the full environment table.