AI agent
Turn a natural-language prompt into an Excalidraw scene with Claude, ChatGPT or Gemini, and what it costs.
On this site these pages are documentation, not a hosted service.
excalidraw.2plot.dev carries no provider keys, so generation is disabled
here and nothing below will spend anything. Everything on the page —
streaming onto the canvas, the cost estimate, the Stop button, the daily
spend ceiling — works when you run the repo locally with a
.envholding aprovider key. This is deliberate and not a fault.
Overview
Turn a natural-language prompt into an Excalidraw scene. Pick a provider — Claude, ChatGPT or Gemini — write a prompt, hit generate; the result is dispatched to the canvas via command: updateScene (no component remount). Run the same prompt twice with different providers to compare output.
Claude and ChatGPT runs stream: elements appear on the canvas as the model closes each one, so you watch the scene being drawn rather than waiting on a spinner and receiving it all at once. The status line counts shapes and seconds as they land. Gemini has no streaming path in its SDK, so it still returns the whole scene at the end — the page keeps that behaviour rather than faking progress it cannot see.
Each provider is enabled only when its key is present, and the badges above the controls say which name was looked for. The cost line under the button prices the run before you click it: a typical figure and a ceiling, because the max-token budget bounds the spend without determining it. Gemini has no price on file here and says so rather than showing $0.00.
Live demo
Source
# File: docs/ai-agent/ai_agent.py
"""AI agent: turn a prompt into an Excalidraw scene via Claude, ChatGPT or Gemini.
Uses the command dispatch pattern (`command: updateScene`) — no component
remount, no key-hack needed on the canvas. The same page compares Claude
Opus 4.7 / Sonnet 4.6 against GPT-6 Astra and Gemini 2.5 Flash / Pro.
Ship-readiness checklist for using this in your own production app:
1. Set env vars:
ANTHROPIC_API_KEY=...
CHATGPT_API_KEY=... # OPENAI_API_KEY is accepted as a fallback
GEMINI_API_KEY=...
All three are optional — the page disables the corresponding provider
if its key is missing, and says which name it looked for.
2. This page calls the LLM synchronously inside a Dash callback. For
production-grade UX, wrap with a background task queue (Celery / RQ
/ Dramatiq) or Dash's `background=True` long-callback machinery. The
code here prioritizes clarity over throughput.
3. The prompt template is a starting point — tune `SYSTEM_PROMPT` for
your domain. Prompt caching is active on Claude calls because the
system prompt is stable across requests.
"""
from __future__ import annotations
import json
import os
import re
import time
import traceback
import uuid
from typing import Any, Dict
import dash
import dash_mantine_components as dmc
from dash import Input, Output, State, callback, clientside_callback, dcc, html, no_update
from dash_excalidraw import DashExcalidraw
from lib import scene_library, scene_stream, spend
from lib.scene_ai import ( # shared with /benchmark — see lib/scene_ai.py
CLAUDE_EFFORT,
CLAUDE_MAX_TOKENS,
CLAUDE_MODELS,
MODEL_PRICING,
EFFORT_CAPABLE,
EFFORT_LEVELS,
NO_KEYS_NOTICE,
MODEL_EFFORT,
MODEL_MAX_TOKENS,
OPENAI_MODELS,
any_provider_configured,
available_models,
call_model,
estimate_cost,
format_money,
supported_efforts,
GEMINI_MAX_TOKENS,
GEMINI_MODELS,
SYSTEM_PROMPT,
_call_gemini,
_cleanup_json,
_coerce_types,
_extract_json_block,
_parse_and_normalize,
_spend_allowed,
)
from docs._shared import canvas_frame, sync_canvas_theme
sync_canvas_theme("ai-canvas")
def offered_claude_models():
"""The models this deployment's key can actually call.
CALLED FROM A CALLBACK, NEVER AT IMPORT. Page modules are imported while
Dash registers pages, so a module-level call here put an outbound request
to api.anthropic.com on the BOOT path — every start of the app blocked on
a third party being reachable, and the answer was then frozen for the life
of the process. It was visible in the boot log, one line under
"Loading docs/ai-agent/ai-agent.md":
INFO:httpx:HTTP Request: GET https://api.anthropic.com/v1/models...
`available_models` caches, so calling it per callback costs one request
for the process and no more. `verified` False means the check could not
run — the full table is offered and the page SAYS so rather than implying
it was checked.
"""
return available_models(CLAUDE_MODELS)
# ---------------------------------------------------------------------------
# Prompt template (domain-specific instructions for producing Excalidraw JSON)
# ---------------------------------------------------------------------------
# Read once at import: on the deployed site this never changes, and locally a
# `.env` edit means a restart anyway.
ANY_KEY = any_provider_configured()
HAS_CLAUDE_KEY = bool(os.environ.get("ANTHROPIC_API_KEY"))
# CHATGPT_API_KEY is this site's name; OPENAI_API_KEY is the SDK's own and is
# accepted so a machine that already exports one needs no second copy.
HAS_CHATGPT_KEY = bool(os.environ.get("CHATGPT_API_KEY")) or bool(
os.environ.get("OPENAI_API_KEY")
)
HAS_GEMINI_KEY = bool(os.environ.get("GEMINI_API_KEY")) or bool(
os.environ.get("GOOGLE_API_KEY")
)
# ---------------------------------------------------------------------------
# Providers
# ---------------------------------------------------------------------------
def _benchmark_status(provider, model, element_count, elapsed, meta) -> str:
"""One line carrying everything a sweep needs to be compared.
Elements alone say nothing about whether a setting was worth it — the
interesting question is elements per second and per dollar, at a given
effort and budget. So the line reports the settings ACTUALLY applied
(which is not always what was asked: effort is silently dropped on models
that reject it) alongside tokens, latency and an estimated cost.
"""
plural = "s" if element_count != 1 else ""
head = f"{model} — {element_count} element{plural} in {elapsed:.0f}s"
if not meta:
return f"Generated via {provider} / {head}."
bits = [f"effort={meta['effort']}", f"max_tokens={meta['max_tokens']:,}"]
if meta.get("effort_ignored"):
bits.append("(effort ignored — model rejects it)")
out_tok, in_tok = meta["output_tokens"], meta["input_tokens"]
bits.append(f"{out_tok:,} out / {in_tok:,} in")
if meta.get("cache_read"):
bits.append(f"{meta['cache_read']:,} cached")
# MODEL_PRICING, not CLAUDE_PRICING: a ChatGPT run priced from the Claude
# table finds nothing and silently reports no cost at all, which reads as
# "this was free".
price = MODEL_PRICING.get(model)
if price and (out_tok or in_tok):
cost = (in_tok * price[0] + out_tok * price[1]) / 1_000_000
# Estimate, not a bill — see CLAUDE_PRICING.
bits.append(f"~${cost:.3f}")
if element_count:
bits.append(f"{elapsed / element_count:.1f}s/element")
# Say truncation in words. `stop=max_tokens` is accurate and means nothing
# to someone whose scene just stopped growing — and hitting the budget is
# the single most common reason a drawing comes out shorter than asked
# for, because the budget covers thinking AND output.
stop = meta.get("stop_reason")
if stop in ("max_tokens", "incomplete"):
bits.append(
f"TRUNCATED — hit the {meta['max_tokens']:,}-token budget; "
f"raise it or lower effort for a longer scene"
)
elif stop and stop not in ("end_turn", "completed"):
bits.append(f"stop={stop}")
return f"{head} · " + " · ".join(bits)
def _format_parse_error(raw: str, exc: json.JSONDecodeError) -> str:
"""Show a window of text around the failure position so the Parsed tab
actually helps you understand what broke."""
pos = exc.pos or 0
start = max(0, pos - 120)
end = min(len(raw), pos + 120)
pointer = " " * (pos - start) + "^"
return (
f"JSON parse error: {exc.msg}\n"
f" at line {exc.lineno}, column {exc.colno} (char {pos})\n\n"
f"--- context (±120 chars) ---\n"
f"{raw[start:end]}\n"
f"{pointer}\n"
f"--- end context ---\n\n"
f"Length of model response: {len(raw)} chars.\n"
f"Try switching to a different model, shortening your prompt, or "
f"asking for a simpler diagram."
)
# ---------------------------------------------------------------------------
# Layout
# ---------------------------------------------------------------------------
def _provider_status():
if not ANY_KEY:
# ONE neutral notice, not three red badges. This is the permanent
# state of the deployed site, so it has to read as "this is how the
# page is here", not as a fault. Red would be a lie: nothing is
# broken and nothing the reader can do would fix it.
return dmc.Alert(
NO_KEYS_NOTICE,
title="AI generation is disabled on this site",
color="blue",
variant="light",
)
items = []
items.append(
dmc.Badge(
"Claude: ready" if HAS_CLAUDE_KEY else "Claude: missing ANTHROPIC_API_KEY",
color="green" if HAS_CLAUDE_KEY else "red",
variant="light",
size="sm",
)
)
items.append(
dmc.Badge(
"ChatGPT: ready"
if HAS_CHATGPT_KEY
else "ChatGPT: missing CHATGPT_API_KEY",
color="green" if HAS_CHATGPT_KEY else "red",
variant="light",
size="sm",
)
)
items.append(
dmc.Badge(
"Gemini: ready"
if HAS_GEMINI_KEY
else "Gemini: missing GEMINI_API_KEY",
color="green" if HAS_GEMINI_KEY else "red",
variant="light",
size="sm",
)
)
# The "unverified" badge cannot be decided here: knowing whether the
# model list was checked means asking the provider, and this function runs
# at page-import time. The slot is filled by `_sync_models` on first
# render instead.
items.append(html.Div(id="ai-model-check"))
return dmc.Group(items, gap="xs")
component = dmc.Stack(
gap="md",
children=[
_provider_status(),
dmc.Paper(
withBorder=True,
p="md",
children=dmc.Stack(
gap="sm",
children=[
dmc.Grid(
gutter="md",
children=[
dmc.GridCol(
dmc.Select(
id="ai-provider",
label="Provider",
data=[
{"value": "claude", "label": "Claude"},
{"value": "chatgpt", "label": "ChatGPT"},
{"value": "gemini", "label": "Gemini"},
],
# Land on a provider whose key is present,
# so the first click can succeed. Falls
# back to Claude, which then reports the
# missing key by name.
value=(
"claude"
if HAS_CLAUDE_KEY
else "chatgpt"
if HAS_CHATGPT_KEY
else "gemini"
if HAS_GEMINI_KEY
else "claude"
),
),
span={"base": 12, "sm": 3},
),
dmc.GridCol(
dmc.Select(
id="ai-model",
label="Model",
# Filled by `_sync_models` on first
# render — see offered_claude_models().
data=[],
value=None,
),
span={"base": 12, "sm": 4},
),
dmc.GridCol(
dmc.Select(
id="ai-effort",
label="Effort",
description="Thinking depth",
data=EFFORT_LEVELS,
value=CLAUDE_EFFORT.get(
CLAUDE_MODELS[0]["value"]
) or "low",
),
span={"base": 6, "sm": 2},
),
dmc.GridCol(
dmc.NumberInput(
id="ai-max-tokens",
label="Max tokens",
description="Caps thinking + output together",
value=CLAUDE_MAX_TOKENS[
CLAUDE_MODELS[0]["value"]
],
min=1000,
max=128000,
step=4000,
),
span={"base": 6, "sm": 2},
),
dmc.GridCol(
dmc.NumberInput(
id="ai-seed",
label="Seed",
description="Bump for a fresh id namespace",
value=1,
min=1,
),
span={"base": 12, "sm": 1},
),
],
),
# What this run will cost, BEFORE it is triggered. The
# three controls above are all cost levers and none of
# them says so on its own — model sets the rate, max
# tokens sets the ceiling, effort sets how much of that
# ceiling gets used.
dmc.Alert(
id="ai-estimate",
color="gray",
variant="light",
p="xs",
),
dmc.Textarea(
id="ai-prompt",
label="Prompt",
placeholder="E.g. 'a flowchart for onboarding a new engineer: signup → security training → first PR → mentorship pairing'",
autosize=True,
minRows=3,
maxRows=6,
),
dmc.Group(
[
dmc.Button(
"Generate",
id="ai-generate-btn",
leftSection="✨",
color="indigo",
loaderProps={"type": "dots"},
# Nothing to generate with, so the control is
# off rather than offering a click that can
# only produce an error.
disabled=not ANY_KEY,
),
# THE BRAKES. Enabled only while a run is in
# flight, and filled (not subtle) because the
# moment you want it is the moment you should not
# have to hunt for it.
dmc.Button(
"Stop",
id="ai-stop-btn",
leftSection="■",
color="red",
disabled=True,
),
dmc.Button(
"Clear canvas",
id="ai-clear-btn",
variant="subtle",
color="red",
),
# The scene library: LOCAL-ONLY, enabled on the same
# condition as generation itself (provider keys
# present). See lib/scene_library.py.
dmc.Button(
"Save scene",
id="ai-save-btn",
leftSection="💾",
variant="light",
color="indigo",
disabled=not ANY_KEY,
),
dmc.Select(
id="ai-library",
placeholder=(
"Saved scenes (local)" if ANY_KEY
else "Scene library is local-only"
),
data=[],
searchable=True,
clearable=True,
w=380,
disabled=not ANY_KEY,
),
]
),
],
),
),
# Processing banner — driven by the callback's `running=` clause so
# feedback appears the instant the click registers, not once the
# model has returned. `running` collides with normal Outputs, so
# this element is only *shown*/*hidden* from running; its inner
# copy is static.
html.Div(
id="ai-processing-banner",
style={"display": "none"},
children=dmc.Alert(
color="blue",
variant="light",
children=dmc.Group(
gap="sm",
children=[
dmc.Loader(size="sm", color="blue", type="bars"),
dmc.Stack(
gap=0,
children=[
dmc.Text(
"Generating scene…",
fw=600,
size="sm",
),
dmc.Text(
"Sending prompt to the model and parsing "
"the response — typically 2–20 s. Opus / "
"Gemini Pro are slower than Sonnet / Flash.",
size="xs",
c="dimmed",
),
],
),
],
),
),
),
dmc.Alert(
id="ai-status",
children="Ready.",
color="gray",
variant="light",
title="Status",
),
dcc.Store(id="ai-last-raw", data=""),
# The drawing in ARRIVAL ORDER, captured from the canvas command stream
# (see the timeline section at the bottom). The timeline replays
# `elements[:i]`.
dcc.Store(id="ai-history", data={"elements": []}),
# Mirror of this page's controls for window.DeckBridge (assets/
# deck_bridge.js) — how a Stream Deck drives this page. Holds only a
# timestamp; the state itself lives on window.
dcc.Store(id="ai-deck-mirror"),
dcc.Store(id="ai-library-mount", data=0),
# The streaming pair. `ai-run` holds the id of the generation in
# flight plus how many elements the canvas has already been shown;
# `ai-stream-tick` collects whatever arrived since. 400ms is fast
# enough that shapes appear to land as they are drawn, and slow
# enough that a long generation is a few dozen polls rather than a
# few thousand.
dcc.Store(id="ai-run", data=None),
dcc.Interval(id="ai-stream-tick", interval=400, disabled=True),
dmc.Tabs(
value="canvas",
children=[
dmc.TabsList(
[
dmc.TabsTab("Canvas", value="canvas"),
dmc.TabsTab("Raw response", value="raw"),
dmc.TabsTab("Parsed envelope", value="parsed"),
]
),
dmc.TabsPanel(
value="canvas",
pt="sm",
children=dmc.Box(
style={"position": "relative"},
children=[
dmc.LoadingOverlay(
id="ai-canvas-overlay",
visible=False,
zIndex=100,
overlayProps={
"radius": "md",
"blur": 2,
"backgroundOpacity": 0.55,
},
loaderProps={
"color": "indigo",
"type": "bars",
"size": "lg",
},
),
canvas_frame(
DashExcalidraw(
id="ai-canvas",
height="640px",
UIOptions={
"welcomeScreen": False,
"canvasActions": {
"clearCanvas": True,
"export": False,
"saveAsImage": True,
},
},
),
min_height=640,
),
# THE TIMELINE. Every scene arrives one element at a
# time, so the draw order is a timeline for free:
# scrub it to step back through the drawing, or
# replay it. Works on saved scenes too.
dmc.Paper(
withBorder=True,
p="sm",
mt="xs",
children=dmc.Stack(
gap=6,
children=[
dmc.Group(
[
dmc.Button(
"Replay",
id="ai-timeline-play",
leftSection="▶",
size="xs",
variant="light",
color="indigo",
),
dmc.Text("Timeline", size="sm", fw=600),
dmc.Text(
"No scene yet.",
id="ai-timeline-label",
size="xs",
c="dimmed",
),
],
gap="sm",
),
dmc.Slider(
id="ai-timeline",
min=0,
max=1,
value=0,
step=1,
disabled=True,
color="indigo",
),
],
),
),
],
),
),
dmc.TabsPanel(
value="raw",
pt="sm",
children=dcc.Loading(
id="ai-raw-loading",
type="default",
delay_show=200,
delay_hide=300,
custom_spinner=dmc.Stack(
gap="xs",
p="sm",
children=[
dmc.Skeleton(height=18, width="40%", radius="sm"),
dmc.Skeleton(height=14, radius="sm"),
dmc.Skeleton(height=14, width="95%", radius="sm"),
dmc.Skeleton(height=14, width="80%", radius="sm"),
dmc.Skeleton(height=14, width="90%", radius="sm"),
dmc.Skeleton(height=14, width="70%", radius="sm"),
dmc.Skeleton(height=14, radius="sm"),
dmc.Skeleton(height=14, width="60%", radius="sm"),
],
),
children=dmc.ScrollArea(
style={"height": 520},
children=dmc.Code(
id="ai-raw",
block=True,
style={
"whiteSpace": "pre-wrap",
"wordBreak": "break-word",
"fontSize": 11,
},
),
),
),
),
dmc.TabsPanel(
value="parsed",
pt="sm",
children=dcc.Loading(
id="ai-parsed-loading",
type="default",
delay_show=200,
delay_hide=300,
custom_spinner=dmc.Stack(
gap="xs",
p="sm",
children=[
dmc.Skeleton(height=18, width="35%", radius="sm"),
dmc.Skeleton(height=14, radius="sm"),
dmc.Skeleton(height=14, width="88%", radius="sm"),
dmc.Skeleton(height=14, width="92%", radius="sm"),
dmc.Skeleton(height=14, width="75%", radius="sm"),
dmc.Skeleton(height=14, radius="sm"),
dmc.Skeleton(height=14, width="65%", radius="sm"),
],
),
children=dmc.ScrollArea(
style={"height": 520},
children=dmc.Code(
id="ai-parsed",
block=True,
style={
"whiteSpace": "pre-wrap",
"wordBreak": "break-word",
"fontSize": 11,
},
),
),
),
),
],
),
],
)
# ---------------------------------------------------------------------------
# Callbacks
# ---------------------------------------------------------------------------
@callback(
Output("ai-estimate", "children"),
Output("ai-estimate", "color"),
Input("ai-provider", "value"),
Input("ai-model", "value"),
Input("ai-effort", "value"),
Input("ai-max-tokens", "value"),
)
def _show_estimate(provider, model, effort, max_tokens):
"""Price this run before it happens.
Two numbers, not one, and the reason is in `estimate_cost`: `max_tokens`
is a ceiling, so a single figure has to choose between being pessimistic
(quote the ceiling, and every real run looks like a bargain) or optimistic
(quote the typical, and the bill can exceed the quote). Showing both makes
the spread itself the information — it is exactly what effort controls.
"""
if not ANY_KEY:
return (
"No provider keys on this site — nothing is spent here. "
"The estimate works locally with a `.env`.",
"blue",
)
if not model:
# First paint: the model control ships empty and `_sync_models` fills
# it. Without this the estimate would flash "no price on file for
# this model", which reads as a broken model rather than an unfinished
# render.
return "Checking which models this key can call…", "gray"
if provider == "gemini":
# Gemini has no entry in MODEL_PRICING and its billing is not ours to
# quote. Saying so beats rendering $0.00, which reads as "free".
return "No cost estimate for Gemini here — see Google's pricing.", "gray"
est = estimate_cost(model, effort, max_tokens)
if not est["priced"]:
if est["reason"] == "budget":
# Mid-edit the NumberInput hands over whatever is in the box, so
# this is a normal transient state, not an error worth shouting
# about. It used to raise here and 500 the callback.
return "Set a max-token budget to see the cost.", "gray"
return "No price on file for this model — cost unknown.", "gray"
# The budget AS PRICED, not the raw control value — they differ whenever
# the box holds something like "64.000" or a number outside the range.
budget = est["budget"]
sent = est["effort"]
effort_phrase = (
f"effort {sent}" if sent else "no effort set (the model's own default)"
)
budget_state = spend.summary()
parts = [
dmc.Text(
[
"About ",
dmc.Text(format_money(est["typical"]), fw=700, span=True),
f" for this run · at most {format_money(est['ceiling'])} "
f"if it uses the whole budget.",
],
size="sm",
),
dmc.Text(
f"{budget:,} max tokens · {effort_phrase} · "
f"typical run uses ~{est['fraction']:.0%} of the budget.",
size="xs",
c="dimmed",
),
]
# What is left today. Shown always rather than only when low: a number
# that appears for the first time when you are nearly out is a number
# nobody has learned to read.
parts.append(
dmc.Text(
f"Daily budget: {format_money(budget_state['spent'])} of "
f"{format_money(budget_state['ceiling'])} used · "
f"{format_money(budget_state['remaining'])} left (resets midnight UTC).",
size="xs",
c="orange" if budget_state["fraction"] > 0.8 else "dimmed",
)
)
# Say it out loud when the level in the selector is not the level that
# will be sent — otherwise the estimate looks wrong rather than the
# control looking inert.
if effort and effort != "none" and sent is None:
parts.append(
dmc.Text(
f"This model ignores the effort parameter, so “{effort}” "
f"costs the same as any other level here.",
size="xs",
c="dimmed",
fs="italic",
)
)
return dmc.Stack(parts, gap=2), "gray"
# Clear stays SYNCHRONOUS and is its own callback. It is instant, and routing
# it through a job queue would add a round trip to a button whose whole value
# is that it responds immediately. Splitting it also keeps `ctx.triggered_id`
# out of the background worker, where callback context is a different animal.
#
# It writes the same Output as the generator, so one of the two must declare
# allow_duplicate — this one, because it is the secondary writer.
@callback(
Output("ai-model", "data"),
Output("ai-model", "value"),
Output("ai-model-check", "children"),
Input("ai-provider", "value"),
prevent_initial_call=False,
)
def _sync_models(provider):
"""Fill the model list, and say whether it was verified.
`prevent_initial_call=False` is load-bearing: this fires on first render,
which is what lets the control ship empty and the availability check stay
off the import path.
"""
if provider == "chatgpt":
return OPENAI_MODELS, OPENAI_MODELS[0]["value"], None
if provider == "gemini":
return GEMINI_MODELS, GEMINI_MODELS[0]["value"], None
data, verified = offered_claude_models()
# With no key at all, "unverified" is noise: of course it could not be
# checked, and the notice above already says why. The badge is for the
# case where a key EXISTS and the check still could not run.
badge = (
None
if verified or not ANY_KEY
else dmc.Badge(
"model list unverified", color="yellow", variant="light", size="sm"
)
)
return data, data[0]["value"], badge
@callback(
Output("ai-effort", "value"),
Output("ai-effort", "data"),
Output("ai-max-tokens", "value"),
Output("ai-effort", "disabled", allow_duplicate=True),
Input("ai-model", "value"),
prevent_initial_call=True,
)
def _sync_model_defaults(model):
"""Move effort and budget to this model's defaults when the model changes.
Without this, switching models silently carries the previous model's
settings over — which quietly invalidates a comparison, because you would
be reading a difference between models that is partly a difference in
configuration. The effort control is also disabled outright on models that
reject the parameter, so the UI cannot offer a choice the API will 400 on.
"""
capable = model in EFFORT_CAPABLE
allowed = set(supported_efforts(model))
data = [e for e in EFFORT_LEVELS if e["value"] in allowed]
# MODEL_* rather than CLAUDE_*: switching to a ChatGPT model has to pick
# up ITS defaults, or the run would carry the previous provider's budget
# and the comparison would be measuring configuration, not models.
effort = (MODEL_EFFORT.get(model) or "none") if capable else "none"
if effort not in allowed:
effort = "none"
return effort, data, MODEL_MAX_TOKENS.get(model, 32000), not capable
# Clear stays SYNCHRONOUS and is its own callback. It is instant, and routing
# it through a job queue would add a round trip to a button whose whole value
# is that it responds immediately. Splitting it also keeps `ctx.triggered_id`
# out of the background worker, where callback context is a different animal.
#
# It writes the same Output as the generator, so one of the two must declare
# allow_duplicate — this one, because it is the secondary writer.
@callback(
Output("ai-canvas", "command", allow_duplicate=True),
Output("ai-status", "children", allow_duplicate=True),
Output("ai-status", "color", allow_duplicate=True),
Output("ai-raw", "children", allow_duplicate=True),
Output("ai-parsed", "children", allow_duplicate=True),
Output("ai-last-raw", "data", allow_duplicate=True),
Input("ai-clear-btn", "n_clicks"),
prevent_initial_call=True,
)
def _clear(_clicks):
cmd = {
"id": f"clear-{uuid.uuid4()}",
"type": "updateScene",
"payload": {"elements": []},
}
return cmd, "Canvas cleared.", "gray", "", "", ""
@callback(
Output("ai-canvas", "command"),
Output("ai-status", "children"),
Output("ai-status", "color"),
Output("ai-raw", "children"),
Output("ai-parsed", "children"),
Output("ai-last-raw", "data"),
Output("ai-run", "data"),
Output("ai-stream-tick", "disabled"),
Input("ai-generate-btn", "n_clicks"),
State("ai-provider", "value"),
State("ai-model", "value"),
State("ai-effort", "value"),
State("ai-max-tokens", "value"),
State("ai-prompt", "value"),
running=[
# ONLY the banner. The canvas overlay used to be here and is gone on
# purpose: it covered the canvas for the whole call, which is the
# exact thing streaming exists to remove. Everything else that needs
# locking is handled by `_lock_controls`, which keys off the ticker
# and therefore stays on for the whole DRAWING, not just this
# callback — which now returns in milliseconds.
(
Output("ai-processing-banner", "style"),
{"display": "block", "marginTop": 4, "marginBottom": 4},
{"display": "none"},
),
],
prevent_initial_call=True,
)
def _generate(_gen_clicks, provider, model, effort, max_tokens, prompt):
"""START a generation. Returns immediately; the canvas fills in as it draws.
This used to block for the whole call and hand back a finished scene, which
is why the page had a loading overlay: there was nothing to show until
there was everything to show. Claude and ChatGPT runs now go to a worker
thread that parses elements out of the token stream as each one closes,
and `ai-stream-tick` collects them — first shape on the canvas in a few
seconds instead of a minute of spinner.
`background=` is gone with the blocking call. It existed so a 100-second
generation could not hold a request worker and take /healthz down with it;
the generation no longer happens in a callback at all, so the risk it
guarded against is gone with it. The producer is a plain thread — see
lib/scene_stream.start for why not a background worker.
Gemini still runs synchronously: `_call_gemini` has no streaming path and
returns a whole string, so there is nothing to stream. It keeps the old
behaviour rather than pretending otherwise.
"""
idle = (no_update,) * 6 + (no_update, True)
# The button is disabled, so this is only reachable by a crafted request.
# Answered in the same words and the same colour as the notice above,
# because a caller who gets here has not done anything wrong either.
if not ANY_KEY:
return (no_update, NO_KEYS_NOTICE, "blue") + idle[3:]
if not prompt or not prompt.strip():
return (no_update, "Write a prompt first.", "yellow") + idle[3:]
# ---- THE SPEND GATE -------------------------------------------------
# This check has to live HERE, not on the page's `tier: auth`, and the
# reason is worth stating because the frontmatter looks like it covers it.
#
# 1. `lib/page_tiers.degraded_tier` makes every tier except `hidden` fail
# OPEN when Clerk is not configured. That is the right trade for
# reading documentation and exactly the wrong one for a page that
# spends money, which must fail CLOSED.
# 2. Page tiers are path-based, and every Dash callback posts to the one
# shared `/_dash-update-component` route. No path-based gate can tell
# this callback from any other.
#
# So the page tier governs who can READ the page; this governs who can
# make it BILL.
if not _spend_allowed():
return (
no_update,
"Sign in to generate — this page spends real API credits, so "
"generation is limited to signed-in visitors.",
"yellow",
) + idle[3:]
missing = {
"claude": (not HAS_CLAUDE_KEY, "ANTHROPIC_API_KEY is not set in the environment."),
"chatgpt": (
not HAS_CHATGPT_KEY,
"CHATGPT_API_KEY / OPENAI_API_KEY is not set in the environment.",
),
"gemini": (
not HAS_GEMINI_KEY,
"GEMINI_API_KEY / GOOGLE_API_KEY is not set in the environment.",
),
}.get(provider, (False, ""))
if missing[0]:
return (no_update, missing[1], "red") + idle[3:]
# The ceiling, checked HERE as well as inside the call. `stream_model`
# admits once per run, but it is a generator whose body does not execute
# until the worker thread advances it — so without this the refusal would
# arrive as a red error on a run that had already been created. Checking
# up front makes a blown budget a clean, yellow "not now".
try:
spend.check(estimate_cost(model, effort, max_tokens)["typical"])
except spend.CeilingReached as exc:
return (no_update, str(exc), "yellow") + idle[3:]
# ---- Gemini: no stream available, so keep the one-shot path ------------
if provider == "gemini":
started = time.monotonic()
try:
raw = _call_gemini(model, prompt.strip())
except Exception as exc: # noqa: BLE001 - surface any provider error
traceback.print_exc()
return (no_update, f"{provider} call failed: {exc}", "red",
str(exc), "", "", no_update, True)
try:
parsed = _parse_and_normalize(raw)
except json.JSONDecodeError as exc:
return (
no_update,
f"Parse error at char {exc.pos}: {exc.msg}. See Parsed tab for context.",
"red", raw, _format_parse_error(raw, exc), raw, no_update, True,
)
except ValueError as exc:
return (no_update, f"Parse error: {exc}", "red", raw, str(exc), raw,
no_update, True)
elements = parsed.get("elements", [])
cmd = {
"id": f"ai-{uuid.uuid4()}",
"type": "updateScene",
"payload": {
"elements": elements,
"appState": parsed.get("appState", {}),
"files": parsed.get("files", {}),
},
}
status = _benchmark_status(
provider, model, len(elements), time.monotonic() - started, None
)
return (cmd, status, "green", raw, json.dumps(parsed, indent=2), raw,
no_update, True)
# ---- Claude / ChatGPT: stream it --------------------------------------
try:
run_id = scene_stream.start(
model=model,
user_prompt=prompt.strip(),
max_tokens=max_tokens,
effort=effort,
)
except Exception as exc: # noqa: BLE001 - a bad model id, a missing key
traceback.print_exc()
return (no_update, f"{type(exc).__name__}: {exc}", "red") + idle[3:]
# Blank the canvas so the drawing starts from nothing and each shape's
# arrival is visible, rather than accumulating over the previous scene.
return (
{"id": f"ai-reset-{run_id[:8]}", "type": "resetScene", "payload": {}},
"Drawing…",
"gray",
no_update,
no_update,
no_update,
{"id": run_id, "cursor": 0, "provider": provider, "model": model},
False,
)
@callback(
Output("ai-canvas", "command", allow_duplicate=True),
Output("ai-status", "children", allow_duplicate=True),
Output("ai-status", "color", allow_duplicate=True),
Output("ai-raw", "children", allow_duplicate=True),
Output("ai-parsed", "children", allow_duplicate=True),
Output("ai-last-raw", "data", allow_duplicate=True),
Output("ai-run", "data", allow_duplicate=True),
Output("ai-stream-tick", "disabled", allow_duplicate=True),
Input("ai-stream-tick", "n_intervals"),
State("ai-run", "data"),
prevent_initial_call=True,
)
def _stream_tick(_n, run):
"""Collect whatever the worker has parsed since the last poll."""
stop = (no_update,) * 7 + (True,)
if not run or not run.get("id"):
return stop
state = scene_stream.take(run["id"], run.get("cursor", 0))
if not state["found"]:
# Expired, forgotten, or an instance restart took the store with it.
# Stop polling rather than ask forever about a run that cannot answer.
return stop
elements = state["all"]
run_next = {**run, "cursor": state["cursor"]}
if state["error"]:
scene_stream.forget(run["id"])
return (no_update, state["error"], "red", no_update, no_update,
no_update, None, True)
# updateScene REPLACES the scene, so every tick sends everything so far.
# captureUpdate NEVER while drawing: without it each tick becomes its own
# undo step, and a 40-element scene would take 40 Ctrl+Z to undo.
cmd = no_update
if state["elements"]:
cmd = {
"id": f"ai-tick-{run['id'][:8]}-{state['cursor']}",
"type": "updateScene",
"payload": {"elements": elements, "captureUpdate": "NEVER"},
}
# A gap means elements the producer wrote can no longer be read back, so
# the canvas is missing shapes it was already shown. Saying nothing is how
# this last went unnoticed for a whole session — it just looked like the
# model drawing less.
lost_note = f" · {state['lost']} lost from the buffer" if state.get("lost") else ""
if not state["done"]:
plural = "s" if len(elements) != 1 else ""
return (
cmd,
f"Drawing… {len(elements)} element{plural} so far "
f"({state['elapsed']:.0f}s){lost_note}",
"orange" if state.get("lost") else "gray",
no_update, no_update, no_update, run_next, False,
)
scene_stream.forget(run["id"])
if not elements:
return (
no_update,
"The model returned no elements. Try a different model or a "
"larger budget.",
"yellow",
no_update, no_update, no_update, None, True,
)
# One last updateScene, this time as a single undoable step.
final_cmd = {
"id": f"ai-final-{run['id'][:8]}",
"type": "updateScene",
"payload": {"elements": elements, "captureUpdate": "IMMEDIATELY"},
}
# The Raw tab shows the scene REBUILT from the streamed elements rather
# than the exact bytes off the wire: the parser consumes the deltas as
# they arrive, and keeping a second full copy of every response in memory
# to populate a tab is not worth the memory.
rebuilt = json.dumps({"elements": elements}, indent=2)
status = _benchmark_status(
run.get("provider"), run.get("model"), len(elements),
state["elapsed"], state["meta"],
)
if state.get("lost"):
status = (
f"{status} · WARNING: {state['lost']} element(s) were generated "
f"but could not be read back, so this scene is incomplete."
)
return final_cmd, status, "orange", rebuilt, rebuilt, rebuilt, None, True
return final_cmd, status, "green", rebuilt, rebuilt, rebuilt, None, True
@callback(
Output("ai-status", "children", allow_duplicate=True),
Output("ai-status", "color", allow_duplicate=True),
Output("ai-run", "data", allow_duplicate=True),
Output("ai-stream-tick", "disabled", allow_duplicate=True),
Input("ai-stop-btn", "n_clicks"),
State("ai-run", "data"),
prevent_initial_call=True,
)
def _stop(_clicks, run):
"""Press the brakes: stop the generation, keep what it has drawn.
Two things have to happen and only one of them is obvious. Disabling the
ticker stops the PAGE from polling — but the worker thread would carry on
pulling tokens to the full budget with nobody reading them, which is
exactly the cost this button exists to prevent. `scene_stream.cancel`
closes the provider stream, so billing stops at the tokens already
produced.
What is on the canvas stays there. A stop is not an undo: the shapes drawn
so far are usually the reason you are stopping.
"""
if not run or not run.get("id"):
return no_update, no_update, no_update, True
scene_stream.cancel(run["id"])
drawn = len(scene_stream.take(run["id"], 0)["all"])
plural = "s" if drawn != 1 else ""
return (
f"Stopped. {drawn} element{plural} kept; no further tokens are being "
f"generated.",
"yellow",
None,
True,
)
@callback(
Output("ai-generate-btn", "loading"),
Output("ai-generate-btn", "disabled"),
Output("ai-stop-btn", "disabled"),
Output("ai-clear-btn", "disabled"),
Output("ai-prompt", "disabled"),
Output("ai-provider", "disabled"),
Output("ai-model", "disabled"),
Output("ai-max-tokens", "disabled"),
Input("ai-stream-tick", "disabled"),
)
def _lock_controls(tick_disabled):
"""Lock the controls while a drawing is in flight.
Keyed on the TICKER, not on the generate callback's lifetime: that
callback now returns in milliseconds, so a `running=` lock would release
while the canvas was still filling in. The ticker being enabled is the
definition of "a run is in progress".
"""
drawing = not tick_disabled
# `or not ANY_KEY`: with no provider configured these stay off for good.
# Without it this callback would re-enable the button on first render and
# hand the reader a click whose only outcome is a red error.
off = drawing or not ANY_KEY
# Stop is the one control that is enabled precisely when a run is going.
return (drawing, off, not drawing) + (off,) * 5
# ---------------------------------------------------------------------------
# Timeline, scene library, and the Stream Deck bridge
# ---------------------------------------------------------------------------
#
# The timeline is CLIENTSIDE end to end. The history is captured from the
# canvas `command` prop itself — every updateScene the stream sends carries
# all elements so far, in arrival order — so none of the streaming callbacks
# above had to change. Scrubbing writes the canvas through
# `dash_clientside.set_props` rather than an Output: an Output on
# `ai-canvas.command` from a callback that (indirectly) listens to it would be
# a dependency cycle, and set_props is outside the static graph. Commands the
# timeline itself sends carry a `tl-` id prefix and are ignored by the capture.
clientside_callback(
"""
function (cmd) {
var nu = window.dash_clientside.no_update;
if (!cmd || !cmd.type) return nu;
if (String(cmd.id || "").indexOf("tl-") === 0) return nu;
if (cmd.type === "resetScene") return {elements: [], t: Date.now()};
if (cmd.type === "updateScene" && cmd.payload && Array.isArray(cmd.payload.elements)) {
return {elements: cmd.payload.elements, t: Date.now()};
}
return nu;
}
""",
Output("ai-history", "data"),
Input("ai-canvas", "command"),
prevent_initial_call=True,
)
clientside_callback(
"""
function (h) {
var n = ((h && h.elements) || []).length;
window.__aiTlShown = n; // the canvas already shows the whole scene
return [Math.max(n, 1), n, n === 0,
n ? (n + " / " + n + " elements") : "No scene yet."];
}
""",
Output("ai-timeline", "max"),
Output("ai-timeline", "value"),
Output("ai-timeline", "disabled"),
Output("ai-timeline-label", "children"),
Input("ai-history", "data"),
prevent_initial_call=True,
)
clientside_callback(
"""
function (v, h, tickDisabled) {
var nu = window.dash_clientside.no_update;
var els = (h && h.elements) || [];
if (v === null || v === undefined || !els.length) return nu;
var label = v + " / " + els.length + " elements";
if (v === window.__aiTlShown) return label;
if (tickDisabled === false) return label + " — drawing; scrub when it finishes";
window.__aiTlShown = v;
window.dash_clientside.set_props("ai-canvas", {command: {
id: "tl-" + Date.now() + "-" + v,
type: "updateScene",
payload: {elements: els.slice(0, v), captureUpdate: "NEVER"}
}});
return label;
}
""",
Output("ai-timeline-label", "children", allow_duplicate=True),
Input("ai-timeline", "value"),
State("ai-history", "data"),
State("ai-stream-tick", "disabled"),
prevent_initial_call=True,
)
clientside_callback(
"""
function (n, h) {
var nu = window.dash_clientside.no_update;
var N = ((h && h.elements) || []).length;
if (!n || !N) return nu;
if (window.__aiTlTimer) clearInterval(window.__aiTlTimer);
var i = 0, ms = Math.max(40, Math.min(250, 6000 / N));
window.dash_clientside.set_props("ai-timeline", {value: 0});
window.__aiTlTimer = setInterval(function () {
i += 1;
window.dash_clientside.set_props("ai-timeline", {value: i});
if (i >= N) { clearInterval(window.__aiTlTimer); window.__aiTlTimer = null; }
}, ms);
return nu;
}
""",
Output("ai-timeline-play", "loading"),
Input("ai-timeline-play", "n_clicks"),
State("ai-history", "data"),
prevent_initial_call=True,
)
# The deck mirror. Publishes this page's controls to window.DeckBridge on every
# change; the deck reads them back with one JS call and writes through
# set_props, so provider → model → defaults chains run exactly as by hand.
# Page actions (generate, preset, …) live in assets/deck_bridge_ai_agent.js.
clientside_callback(
"""
function (prov, mdata, mval, edata, effort, edis, mt, seed, tl, hist, ldata, lval,
tickDisabled, status, pdata) {
if (!window.DeckBridge) return window.dash_clientside.no_update;
var drawing = tickDisabled === false;
var n = ((hist && hist.elements) || []).length;
return window.DeckBridge.update(location.pathname, {
provider: {type: "select", id: "ai-provider", value: prov, options: pdata, wrap: true, disabled: drawing},
model: {type: "select", id: "ai-model", value: mval, options: mdata, wrap: true, disabled: drawing},
effort: {type: "select", id: "ai-effort", value: effort, options: edata, disabled: !!edis || drawing},
max_tokens: {type: "number", id: "ai-max-tokens", value: mt, min: 1000, max: 128000, step: 4000,
label: (mt || 0).toLocaleString(), disabled: drawing},
seed: {type: "number", id: "ai-seed", value: seed, min: 1, max: 9999, step: 1},
timeline: {type: "number", id: "ai-timeline", value: tl, min: 0, max: n, step: 1,
label: (tl || 0) + "/" + n, disabled: !n || drawing},
scene: {type: "select", id: "ai-library", value: lval, options: ldata,
disabled: !(ldata && ldata.length)}
}, (typeof status === "string" ? status : "") + (drawing ? " [drawing]" : ""));
}
""",
Output("ai-deck-mirror", "data"),
Input("ai-provider", "value"),
Input("ai-model", "data"),
Input("ai-model", "value"),
Input("ai-effort", "data"),
Input("ai-effort", "value"),
Input("ai-effort", "disabled"),
Input("ai-max-tokens", "value"),
Input("ai-seed", "value"),
Input("ai-timeline", "value"),
Input("ai-history", "data"),
Input("ai-library", "data"),
Input("ai-library", "value"),
Input("ai-stream-tick", "disabled"),
Input("ai-status", "children"),
State("ai-provider", "data"),
prevent_initial_call=False,
)
@callback(
Output("ai-library", "data"),
Input("ai-library-mount", "data"),
prevent_initial_call=False,
)
def _list_scenes(_mount):
"""Fill the library on first render — never on the public site, where the
library does not exist (and must not be created)."""
if not ANY_KEY:
return []
try:
return scene_library.options()
except Exception: # noqa: BLE001 - a broken DB must not break the page
traceback.print_exc()
return []
@callback(
Output("ai-library", "data", allow_duplicate=True),
Output("ai-status", "children", allow_duplicate=True),
Output("ai-status", "color", allow_duplicate=True),
Input("ai-save-btn", "n_clicks"),
State("ai-history", "data"),
State("ai-prompt", "value"),
State("ai-provider", "value"),
State("ai-model", "value"),
State("ai-effort", "value"),
State("ai-max-tokens", "value"),
State("ai-seed", "value"),
prevent_initial_call=True,
)
def _save_scene(n, hist, prompt, provider, model, effort, max_tokens, seed):
"""Save the generated scene, in draw order, with the settings that made it."""
if not n:
return no_update, no_update, no_update
if not ANY_KEY:
return no_update, "The scene library is local-only.", "blue"
elements = (hist or {}).get("elements") or []
try:
sid = scene_library.save(elements, prompt=prompt or "", provider=provider, model=model,
effort=effort, max_tokens=max_tokens, seed=seed)
except ValueError as exc:
return no_update, str(exc), "yellow"
return (scene_library.options(),
f"Saved scene #{sid} — {len(elements)} elements, in draw order, "
f"to {scene_library.db_path().name}.", "green")
@callback(
Output("ai-canvas", "command", allow_duplicate=True),
Output("ai-status", "children", allow_duplicate=True),
Output("ai-status", "color", allow_duplicate=True),
Output("ai-prompt", "value"),
Input("ai-library", "value"),
prevent_initial_call=True,
)
def _load_scene(scene_id):
"""Put a saved scene back on the canvas (and its prompt back in the box,
ready to iterate). The timeline picks it up like any other scene."""
if not scene_id or not ANY_KEY:
return no_update, no_update, no_update, no_update
row = scene_library.get(scene_id)
if not row:
return no_update, f"Scene #{scene_id} is not in the library.", "yellow", no_update
cmd = {
"id": f"lib-{scene_id}-{uuid.uuid4().hex[:8]}",
"type": "updateScene",
"payload": {"elements": row["elements"], "captureUpdate": "IMMEDIATELY"},
}
status = (f"Loaded #{row['id']}: {row['title']} — {row['model'] or '?'} · "
f"{row['n_elements']} elements. Scrub the timeline to replay it.")
return cmd, status, "gray", row["prompt"]
:defaultExpanded: false :withExpandedButton: true
Source: /ai-agent
Note for AI agents: This is the static, prerendered view of an interactive Dash application served because we detected a non-JS user agent. Full prose docs:
- /ai-agent/llms.txt — LLM-friendly documentation
- /sitemap.xml
- /robots.txt