Nano Banana: what it costs, and what you actually get

Every Nano Banana version run through the production channel-thumbnail module — buildThumbBrief + applyThumbnailChannelIdentity + golden craft bar + owner A/B preferences — on three different channel types. Byte-identical prompt per channel across all four models.

Generated 2026-09-03 · 12 images · $0.8498 total spend.

1 · Measured, this run

Nano Banana

gemini-2.5-flash-image
$0.039
measured all-in / image · 7.7s avg
1290 image tokens @ 1K 16:9
the original

Nano Banana 2 Lite

gemini-3.1-flash-lite-image
$0.0346
measured all-in / image · 3.3s avg
1120 image tokens @ 1K 16:9
cheapest first-party

Nano Banana 2

gemini-3.1-flash-image
$0.0689
measured all-in / image · 9.8s avg
1120 image tokens @ 1K 16:9

Nano Banana Pro

gemini-3-pro-image
$0.1407
measured all-in / image · 24.0s avg
1120 image tokens @ 1K 16:9
what the module uses today
Watch the token math. Google quotes $0.134 per 1K image for Pro, but the API's candidatesTokenCount came back at ~1460, not 1120. The extra ~340 are thinking/text tokens billed at $12/1M, not the $120/1M image rate. Bill on candidatesTokenCount and you overstate Pro by ~30%. The real all-in is $0.1407 — image tokens plus a small thinking charge.

2 · First-party price list (Google, verified 2026-09-03)

API modelMarketing nameBatch 1KStandard 1K2K4K
gemini-3.1-flash-lite-imageNano Banana 2 Lite$0.0168$0.0336
gemini-2.5-flash-imageNano Banana (1)n/a$0.039
gemini-3.1-flash-imageNano Banana 2$0.034$0.067$0.101$0.151
gemini-3-pro-imageNano Banana Pro$0.067$0.134$0.134$0.240

Nano Banana 2 also has a 0.5K tier: $0.045 standard / $0.022 batch. Nano Banana (1) is billed at 1290 tokens per ≤1024² image; every Gemini 3.x image model is 1120 tokens at both 1K and 2K — which is why NB2 Lite undercuts NB1 despite sharing the same $30/1M rate.

3 · Third-party endpoints

ProviderModelPosted priceNotes
OpenRouterNB Pro / NB2 / NB2 Lite$0.134 / $0.067 / $0.0336Verified live via their /api/v1/models endpoint — token rates are identical to Google's ($120 / $60 / $30 per 1M image tokens). Zero markup, one key for everything.verified
Google Batch APIall Nano Banana models50% of standardFirst-party. Async, up to 24h turnaround. NB2 Lite batch = $0.0168/img is the cheapest legitimate Nano Banana anywhere.verified
kie.aiNB Pro$0.09 (1K–2K), $0.12 (4K)Posted on kie.ai/nano-banana-pro. Credit-based (18 credits ≈ $0.09). Genuine 33% undercut of Google's own rate — the cheapest credible NB Pro reseller.verified
kie.aiNB2$0.04 (1K), $0.06 (2K), $0.09 (4K)Same platform, NB2 tier.reported
ReplicateNB Pro$0.14 (1K–2K), $0.24 (4K)Model JSON exposes no pricing field; figure is from kie.ai's comparison table. Roughly official rate.reported
fal.aiNB Pro$0.15 (1K–2K), $0.30 (4K)Not on fal's public pricing page (only 'Nanobanana $0.0398' is listed). $0.15 is the value pinned in this repo's own sealed contract. Most expensive endpoint in the survey.reported
fal.aiNB (1)$0.0398Listed on fal.ai/pricing. Essentially official rate.verified
WaveSpeed AINB Pro$0.14Aggregator listing only.reported
Together AI / ToapisNB Pro / NB2$0.134 / $0.0466Mirror official serverless rates, no markup.reported
MuapiNB Pro$0.12Aggregator listing only.reported

4 · The sub-cost claims (treat with suspicion)

These are relay/gateway platforms quoting prices below what Google charges them. None were verifiable at source; several are promoted through SEO content farms that cite each other.

PlatformClaimedWhy it doesn't add up
Picsart$0.013 flat (NB2 and NB Pro)Would be 10× under Google's own cost for Pro. Almost certainly a promo/credit-bundle rate or a mislabelled model, not sustainable Pro inference.
ApiPass$0.0136 (NB2 1K) / $0.0864 (NB Pro 1K)NB Pro figure is plausible-ish; the NB2 figure is below Google's wholesale.
laozhang.ai$0.05 flat, all resolutionsChina-facing relay. Alipay/TG payment, 'no VPN needed'. Below Google cost at 2K+.
APIYI$0.05 flatSame relay class, same claim.
Evolink.ai$0.025–$0.05Same relay class.
PoYo$0.12 (NB Pro 1K–2K)Credit reseller.
Apertis$0.50 flat per NB Pro requestMost expensive in the survey; sells free caching + free grounding instead.
One concrete debunk. A widely-cited "9 platforms compared" guide claims OpenRouter serves NB2 at ~$0.0039/image and NB Pro at ~$0.0161. I pulled OpenRouter's live model endpoint: its image-output rates are $0.00006 and $0.00012 per token — exactly Google's $60/$30/$120 per 1M. At 1120 tokens that is $0.067 and $0.134, not $0.0039. That article divided by the wrong token count. Its other numbers deserve the same suspicion.

5 · Same prompt, four models, three channel types

Investory Finance / editorial explainer

“The £40,000 Pension Mistake Nobody Warns You About”

identity profile investory-editorial-realism · energy bold · headline £40,000 / GONE QUIETLY · prompt 4384 UTF-8 bytes (identical for all four models)

Investory — Nano Banana
Nano Banana8.1s · $0.039 measured
Investory — Nano Banana 2 Lite
Nano Banana 2 Lite3.7s · $0.0346 measured
Investory — Nano Banana 2
Nano Banana 210.0s · $0.069 measured
Investory — Nano Banana Pro
Nano Banana Pro22.6s · $0.1404 measured
Show the exact prompt sent to all four models
Rules: 1280x720 YouTube thumbnail. The hero fills 55-75% of the frame, aggressively cropped. Typography is HUGE (owns 25-40% of the frame), ultra-bold, rendered as a designed physical object (plate, smear, strip, slab or sticker) made of the scene's material world, with one PAYOFF word 2-4x larger than the rest. HARD RULE: text NEVER covers the hero's face or eyes - beside, above or across the body only. Spelling EXACTLY as quoted - every visible word must be a correctly spelled real word. Everything must read at 120px on a phone. No play buttons, no UI, no watermarks, no extra small text. Channel "INVESTORY" (signature look, obey strictly: premium realistic financial editorial photograph, tactile real-world materials, decisive human-scale consequence, restrained navy-black and gold grade, palette #0B1220 / #D8A11A, accent #D8A11A). Scene: LAYOUT MODE: split composition; hero opposite the chosen type zone. HERO PROP (dominant, 30-50% of frame, cropped close): a man in his late fifties at a kitchen table, one hand pressed flat on a printed pension statement while the other tears the corner of the page away, his face lit hard from a low window, jaw set in the instant he understands the number. BACKGROUND (separate supporting layer behind the hero - darker, simpler, depth): an out-of-focus suburban kitchen sinking into navy shadow, a cold grey morning beyond the blinds. STORY DETAILS (symbolic, on/around the hero): the torn corner of the statement falling as a thin gold-lit sliver; a cooling mug of tea with an untouched skin on it. Headline: "£40,000" (the payoff word, HUGE) (accent color) then "GONE QUIETLY" - placed clear of all faces. Render the headline as cinematic title-card lettering with metallic bevel and rim light, embedded in the scene atmosphere, blockbuster one-sheet gravity. Small badge pill "INVESTORY" in a corner away from the text. CHANNEL IDENTITY CONTRACT (non-negotiable; reject a scene that misses any visible fact):
MUST SHOW:
- A believable adult human decision or tactile financial artifact visibly enacts the title's exact wealth mechanism.
- The scene still communicates a concrete financial consequence when the headline is covered.
- The dominant material reads as photographed paper, property, cash-flow object, or human-scale evidence rather than a digital asset.
MUST NOT SHOW:
- Generic floating coins, neon market charts, holographic dashboards, anonymous gold bars, luxury filler, video-game art, or a glossy 3D product render. USER-APPROVED GOLDEN CRAFT BAR: One unmistakable, story-specific hero dominates at phone size and is cropped decisively at the frame edge; a centered hero is optional only when the surrounding type and action read more strongly that way. Use only two or three meaningful visual elements with real foreground, hero, and background depth. Avoid a dead 50/50 picture-and-copy split: one hero contour or atmospheric layer should overlap the future type zone. Headline typography is oversized and treated as a physical visual object, never a generic clean font floating in empty space. Keep one forceful contrast system and one restrained accent; reject muddy, generic, or product-render imagery. Channel identity is compact and secondary in a corner; it must never compete with the hook. OWNER-SELECTED A/B PREFERENCES: Prefer a two-word concrete tension or payoff hook over a longer sentence; make the decisive noun visibly dominant. Stage the hero at the peak of action with readable emotion in face, eyes, or hands—not a calm pose after the event. Show cause and consequence together: one close hero action plus one smaller proof detail that completes the story. Bring one consequential hand, object, rupture, or release into the near foreground so the viewer feels the action before decoding the subject. Use a strong diagonal or edge crop to pull the eye from hook to hero to consequence in under one second. Allow a centered hero only for a genuinely stronger peak-action image; arrange supporting type around its silhouette and never turn it into a default symmetrical poster. Do not let a detached signboard or flat black text panel become the composition. Typography must share the scene's material, light, motion, or negative space while leaving the hero dominant. Keep headline contrast immediate and physical while preserving one clean hierarchy; do not let background lore compete.

Inked Histories History / illustrated narrative

“The Night Rome's Treasury Vanished”

identity profile inked-histories-engraved-action · energy spectacle · headline EMPTY / BY DAWN · prompt 4345 UTF-8 bytes (identical for all four models)

Inked Histories — Nano Banana
Nano Banana7.5s · $0.039 measured
Inked Histories — Nano Banana 2 Lite
Nano Banana 2 Lite3.3s · $0.0347 measured
Inked Histories — Nano Banana 2
Nano Banana 210.5s · $0.0691 measured
Inked Histories — Nano Banana Pro
Nano Banana Pro23.1s · $0.1413 measured
Show the exact prompt sent to all four models
Rules: 1280x720 YouTube thumbnail. The hero fills 55-75% of the frame, aggressively cropped. Typography is HUGE (owns 25-40% of the frame), ultra-bold, rendered as a designed physical object (plate, smear, strip, slab or sticker) made of the scene's material world, with one PAYOFF word 2-4x larger than the rest. HARD RULE: text NEVER covers the hero's face or eyes - beside, above or across the body only. Spelling EXACTLY as quoted - every visible word must be a correctly spelled real word. Everything must read at 120px on a phone. No play buttons, no UI, no watermarks, no extra small text. Channel "INKED HISTORIES" (signature look, obey strictly: original high-craft historical ink-and-charcoal editorial illustration on aged paper, tactile cross-hatching, dramatic chiaroscuro and restrained ember-gold proof accents, palette #141110 / #C2761F, accent #C2761F). Scene: LAYOUT MODE: split composition; hero opposite the chosen type zone. HERO PROP (dominant, 30-50% of frame, cropped close): a Roman treasury guard wrenching open an iron-banded strongbox at the peak of the discovery, the lid flung back, his lantern arm thrown wide, mouth open in the half-second before he shouts, the box interior utterly empty. BACKGROUND (separate supporting layer behind the hero - darker, simpler, depth): a vaulted stone undercroft receding into deep cross-hatched blackness, one collapsed shelf, torchlight raking the columns. STORY DETAILS (symbolic, on/around the hero): a single coin still spinning on the flagstones catching an ember-gold highlight; a snapped chain swinging from the strongbox hasp. Headline: "EMPTY" (the payoff word, HUGE) (accent color) then "BY DAWN" - placed clear of all faces. Render the headline as cinematic title-card lettering with metallic bevel and rim light, embedded in the scene atmosphere, blockbuster one-sheet gravity. Small badge pill "INKED HISTORIES" in a corner away from the text. CHANNEL IDENTITY CONTRACT (non-negotiable; reject a scene that misses any visible fact):
MUST SHOW:
- One human historical action, artifact recovery, danger, discovery, or consequence is visibly staged at the peak moment.
- Tactile cross-hatching, aged paper, charcoal depth, and a restrained ember-gold or rust proof detail anchor the illustration.
MUST NOT SHOW:
- AAA adventure-game key art, glossy 3D character models, RPG inventory props, comic-panel grids, speech bubbles, UI overlays, or decorative artifact-only scenes. USER-APPROVED GOLDEN CRAFT BAR: One unmistakable, story-specific hero dominates at phone size and is cropped decisively at the frame edge; a centered hero is optional only when the surrounding type and action read more strongly that way. Use only two or three meaningful visual elements with real foreground, hero, and background depth. Avoid a dead 50/50 picture-and-copy split: one hero contour or atmospheric layer should overlap the future type zone. Headline typography is oversized and treated as a physical visual object, never a generic clean font floating in empty space. Keep one forceful contrast system and one restrained accent; reject muddy, generic, or product-render imagery. Channel identity is compact and secondary in a corner; it must never compete with the hook. OWNER-SELECTED A/B PREFERENCES: Prefer a two-word concrete tension or payoff hook over a longer sentence; make the decisive noun visibly dominant. Stage the hero at the peak of action with readable emotion in face, eyes, or hands—not a calm pose after the event. Show cause and consequence together: one close hero action plus one smaller proof detail that completes the story. Bring one consequential hand, object, rupture, or release into the near foreground so the viewer feels the action before decoding the subject. Use a strong diagonal or edge crop to pull the eye from hook to hero to consequence in under one second. Allow a centered hero only for a genuinely stronger peak-action image; arrange supporting type around its silhouette and never turn it into a default symmetrical poster. Do not let a detached signboard or flat black text panel become the composition. Typography must share the scene's material, light, motion, or negative space while leaving the hero dominant. Keep headline contrast immediate and physical while preserving one clean hierarchy; do not let background lore compete.

Gratitude Springs Meditation / wellness

“Let The Water Take The Weight”

identity profile gratitude-springs-human-sanctuary · energy cozy_pop · headline JUST FLOAT / 10 MINUTES · prompt 4308 UTF-8 bytes (identical for all four models)

Gratitude Springs — Nano Banana
Nano Banana7.5s · $0.039 measured
Gratitude Springs — Nano Banana 2 Lite
Nano Banana 2 Lite2.8s · $0.0345 measured
Gratitude Springs — Nano Banana 2
Nano Banana 28.9s · $0.0687 measured
Gratitude Springs — Nano Banana Pro
Nano Banana Pro26.2s · $0.1404 measured
Show the exact prompt sent to all four models
Rules: 1280x720 YouTube thumbnail. The hero fills 55-75% of the frame, aggressively cropped. Typography is HUGE (owns 25-40% of the frame), ultra-bold, rendered as a designed physical object (plate, smear, strip, slab or sticker) made of the scene's material world, with one PAYOFF word 2-4x larger than the rest. HARD RULE: text NEVER covers the hero's face or eyes - beside, above or across the body only. Spelling EXACTLY as quoted - every visible word must be a correctly spelled real word. Everything must read at 120px on a phone. No play buttons, no UI, no watermarks, no extra small text. Channel "GRATITUDE SPRINGS" (signature look, obey strictly: hyperreal cinematic meditation photography, natural water and atmospheric light, tactile human or environmental serenity, restrained blue and moonlit contrast, palette #0A1A26 / #8FD3E8, accent #8FD3E8). Scene: LAYOUT MODE: centered hero at peak action; reserve asymmetric clean pockets around its silhouette for native typography. HERO PROP (dominant, 30-50% of frame, cropped close): a serene adult woman floating on her back in still dark water, face calm and turned to the sky, arms open and released, hair fanned out, the surface tension breaking into slow rings around her shoulders. BACKGROUND (separate supporting layer behind the hero - darker, simpler, depth): a moonlit spring under low mist, dark treeline dissolving into blue night, one cold band of light across the water. STORY DETAILS (symbolic, on/around the hero): concentric ripples spreading outward from her released hands; a faint breath of steam rising off the warm water into the cold air. Headline: "JUST FLOAT" (the payoff word, HUGE) (accent color) then "10 MINUTES" - placed clear of all faces. Render the headline as cinematic title-card lettering with metallic bevel and rim light, embedded in the scene atmosphere, blockbuster one-sheet gravity. Small badge pill "GRATITUDE SPRINGS" in a corner away from the text. CHANNEL IDENTITY CONTRACT (non-negotiable; reject a scene that misses any visible fact):
MUST SHOW:
- A serene human or environmental sanctuary visibly expresses the episode's exact emotional promise.
- Water, mist, light, foliage, sky, or quiet human presence is used as a tactile calm cue where it fits the topic.
MUST NOT SHOW:
- A repeated rock-pile or paired-stones hero for an unrelated topic, synthetic wellness UI, plastic 3D props, or generic spa stock. USER-APPROVED GOLDEN CRAFT BAR: One unmistakable, story-specific hero dominates at phone size and is cropped decisively at the frame edge; a centered hero is optional only when the surrounding type and action read more strongly that way. Use only two or three meaningful visual elements with real foreground, hero, and background depth. Avoid a dead 50/50 picture-and-copy split: one hero contour or atmospheric layer should overlap the future type zone. Headline typography is oversized and treated as a physical visual object, never a generic clean font floating in empty space. Keep one forceful contrast system and one restrained accent; reject muddy, generic, or product-render imagery. Channel identity is compact and secondary in a corner; it must never compete with the hook. OWNER-SELECTED A/B PREFERENCES: Prefer a two-word concrete tension or payoff hook over a longer sentence; make the decisive noun visibly dominant. Stage the hero at the peak of action with readable emotion in face, eyes, or hands—not a calm pose after the event. Show cause and consequence together: one close hero action plus one smaller proof detail that completes the story. Bring one consequential hand, object, rupture, or release into the near foreground so the viewer feels the action before decoding the subject. Use a strong diagonal or edge crop to pull the eye from hook to hero to consequence in under one second. Allow a centered hero only for a genuinely stronger peak-action image; arrange supporting type around its silhouette and never turn it into a default symmetrical poster. Do not let a detached signboard or flat black text panel become the composition. Typography must share the scene's material, light, motion, or negative space while leaving the hero dominant. Keep headline contrast immediate and physical while preserving one clean hierarchy; do not let background lore compete.

6 · What the pictures actually say

7 · The price floor for Nano Banana Pro at 1K

Every posted rate found, cheapest first. The dividing line is $0.067 — Google's own batch price. Nothing below that is Google selling you an image; it is somebody reselling under wholesale.

Provider1K priceStatusReading
apimodels.app$0.03claimSold as a “budget channel”. The same site lists a primary nanobananapro SKU separately — a cheaper channel for an identical Google model does not exist upstream.
reApi$0.0322claimSells two SKUs side by side: gemini-3-pro-image-preview (default) at $0.0322 and -official at $0.124. A 3.9× gap on one platform for “the same” model is the tell.
Velokey$0.0335claimFlat 1K/2K. Exactly −75% of list, which is under Google’s batch wholesale.
APIYI / laozhang.ai / aiandapi$0.05claimRelay class. Alipay/WeChat, “direct access from China”. Below batch wholesale.
Google Batch API$0.067floorFirst-party, 50% off standard, async up to 24h. This is the real floor — the lowest price Google itself will sell a Pro image for.
Pixapi$0.08claimJust under batch. Plausible only on a subsidised margin.
kie.ai$0.09credibleCheapest synchronous reseller with a real posted rate. Caveat: kie.ai’s own /nano-banana page quotes ~24 credits ≈ $0.12 for the same model — their two pages disagree.
Muapi$0.12crediblePosted per-image, no credit maze.
Modellix$0.1265credibleSits just under Google list.
Google direct / OpenRouter$0.134credibleList price. OpenRouter verified at exactly Google’s token rates, zero markup.
Replicate$0.14credible≈ list.
fal.ai$0.15credibleWhat this repo pays today. Most expensive endpoint found.
The bottom floor you can actually trust is $0.067. The lowest number published anywhere is $0.03 (apimodels.app), then $0.0322 (reApi) and $0.0335 (Velokey) — but reApi and apimodels each sell a separate, dearer "official" SKU for the same model ($0.124 and up). A platform charging 4× more for the honest route is telling you what the cheap route is. Below $0.067 you are buying either a subsidised loss-leader, a resold enterprise quota, or a quietly substituted cheaper model.

Also worth knowing: Google charges the same $0.134 for 1K and 2K (both are 1120 tokens). If you are paying Google or OpenRouter, always ask for 2K — the extra resolution is free. This repo's sealed contract already does.

8 · New signature type treatment: scene_forged

The brief: typography that takes the scene's shape, lighting, angle and symmetry while staying very prominent and distinct. The key constraint, learned the hard way over two failed attempts: it must stay typography — a graphic layer that borrows the scene's light — and never become a physical prop sitting in the scene.

AttemptWhat it producedWhy it failed
metal_monolith v1Colossal brushed-metal slab, support line tucked under one endProminent but detached — and the “never centred” rule actively destroyed the symmetry that made the original render work.
scene_forged v1Brass plaque on the table; letters carved into stone blocksOver-corrected. Perfect lighting and perspective, but the type became a prop: small, recessive, and an object physically covered the letters (“BY DA_N”).
scene_forged v2BelowStays type, owns 25–40% of the frame, borrows the key light / colour temperature / atmosphere, sits on the scene's axis, and no object may cover a letter.
investory scene_forged
Investoryscene_forged
inked-histories scene_forged
Inked Historiesscene_forged
gratitude-springs scene_forged
Gratitude Springsscene_forged

Nano Banana Pro at 2K. Channel badge pinned bottom-right on all three. Remaining gap: on Gratitude Springs the model chose an asymmetric composition, so the “centre on the scene's axis” clause never triggered — symmetry there is a composition instruction (centred hero), not a type one.

9 · Centred hero, and two new heist channels

A centred hero was previously a grudging exception in two places — the art director was told to "use split by default" and "choose centered_hero only when materially stronger", and the owner rules said "allow a centered hero only for…". Both now treat it as an equal, deliberately powerful option: the subject met head-on, symmetrical, one-point perspective, with converging depth and foreground tension. The only banned outcome is a flat, evenly-lit title card with a small object floating in the middle.

Two new heist channels were added to the identity resolver to test it in deliberately different visual worlds — both rendered centred, on Nano Banana Pro at 2K, using scene_forged.

Vault Breach
Vault Breachheist / technical doc
The Getaway Files
The Getaway Filesheist / retro caper
ChannelIdentity contractCentred result
Vault Breach
vault-breach-mechanism-realism
Must show one real security mechanism being defeated by human hands and tools, met head-on with converging perspective. Bans the striped-top burglar cliché, guns, money piles, holographic security UI and 3D vault renders.One-point corridor converging on the drill bit at dead centre, single caged worklight, dust jetting toward the lens, severed shackle in the near foreground. The symmetry is doing the work, not flattening it.
The Getaway Files
getaway-files-period-caper
1970s film still: 35mm grain, halation, faded print stock, period-accurate props. Bans any modern technology, digital sharpness and gun violence as subject.Symmetrical terminal lounge, subject dead centre holding the lens, two identical cases in the near foreground, tungsten halation. No anachronism in frame.

10 · New: story-interest intelligence

Every existing gate in this module checks execution — identity contract, safe zones, copy fidelity, spelling. None of them asked whether the story was worth telling. The Vault Breach render below is technically correct on every one of them and still dull, because concrete is an inert material and a measurement of an inert material carries no human stake.

thumbnailStoryInterest.ts scores the subject deterministically before any image is paid for, and buys one text-only re-plan when the concept is weak.

before
Before — 45/100 weak“18 INCHES / OF CONCRETE”
after
After — 85/100 compelling“6 MONTHS / NOBODY NOTICED”
ConceptScoreWhat the gate said
“18 INCHES / OF CONCRETE”
drill on a vault door
45 weak“the headline is a measurement of an inert material, not a stake”; “names no consequence — it describes rather than threatens”
“6 MONTHS / NEXT DOOR”
first re-plan attempt
60 weakNumber now measures human audacity, but still no consequence named — rejected again
“6 MONTHS / NOBODY NOTICED”
man breaking through, cot and crossed-off calendar behind him
85 compellingHuman at the moment of consequence + audacity marker + a named reversal
How it scores. Human presence and agency in the hero (+20, or −25 when absent); an inert material or surface as hero with nobody acting on it (−25); a headline naming consequence, reversal or jeopardy (+15, −10 when absent); a number attached to a building material rather than to something a person can lose, risk or escape with (−25) — the signature failure; audacity or scale marker (+10); irony or reversal present (+15). Below 65 it re-plans; below 40 it is inert.

A replacement is only kept if it actually scores higher, so a bad re-plan cannot overwrite a merely-weak original. The doctrine also runs as STEP 0 in the art director prompt, ahead of layout and scene invention, so a dull subject is never staged beautifully in the first place.

11 · The leak, and the guard that now catches it

Why it shipped. buildThumbBrief emitted "(the payoff word, HUGE)" inside the same quoted run it told the model to spell verbatim, and BANANA_RULES carried bare HUGE and PAYOFF as emphasis. Nano Banana 2 rendered PAYOFF onto an Investory thumbnail; Nano Banana Pro rendered HUGE onto an Inked Histories one. Not a weak-model artifact.

Why the QA gate let it through. thumbnailOcrMatchesExpected only computed missing. A leaked word adds no missing copy, so the candidate scored exact and shipped. The gate was one-directional.
Fixed in three layers. Both guards fire only on positive evidence, and the vision-transcript fallback is preserved, so unreadable-but-correct stylized type cannot false-block a good candidate.

12 · The real module, end to end

Everything above froze a scene so each model received identical bytes. These two received only a channel name, a video title and the channel's own identity — no hand-written scene, hero or headline. renderCandidate() invented the layout, hero, background, story details and copy itself, so the story-interest gate, identity contract, golden craft bar and badge signature all ran for real.

Sealed Records
Sealed Records“BURIED / DEAL”
Overbuilt
Overbuilt“NO / SEWERS”
ChannelTitle givenWhat the module decided on its own
Sealed Records
sealed-records-document-evidence
“The Secret Deal That Buried The Epstein Case For A Decade”Chose gloved hands lifting a heavily redacted non-prosecution agreement from a weathered evidence box under hard institutional light, and wrote “BURIED / DEAL”. The identity contract bans any recognizable real person, face or likeness, so the exposé is told through the record — no likeness was generated, structurally rather than by request.
Overbuilt
overbuilt-structural-critique
“Why The Burj Khalifa Is A Terrible Building”First concept scored 15/100 inert — “no human presence or agency in the hero; the headline names no consequence”. It re-planned and landed on a sanitation worker wrenching a waste hose into one of a queue of tankers, the tower hazy behind him, and wrote “NO / SEWERS”. Anchored on the documented sewage-trucking issue, not invented damage.
The story-interest gate fired unprompted, in production code. Nothing in the run told it the first Burj concept was dull — it scored the subject, rejected it at 15/100, bought one text-only re-plan and kept the replacement because it scored higher (60). That is the intelligence working on a video it had never seen, before any image was paid for.

Honest caveat: 60/100 is still “weak” — “NO SEWERS” names no consequence word, so the gate improved the concept without getting it to compelling. It kept it because 60 beats 15, exactly as designed.

Badges resolved from channel constants: Sealed Records → solid_pill, Overbuilt → engraved_plate. Both bottom-right. Leak and spelling guards clean on both.

13 · Scope expansion: sober energy and comparison layout

Two constraints the module could not express at all, found by auditing where its rules were one-sided.

WasWhy that is narrowNow
Energy: spectacle / bold / cozy_pop
module comment: “ALL are catchy — none are sleepy”
All three are LOUD. A disaster, a death toll, a grief or health topic, an investigation or a legal reckoning cannot be staged as “impossible scale” or “adorable” without reading as tasteless — and hype on serious material is what makes a channel look untrustworthy.sober — arresting through restraint, not saturation. It also flips the pipeline's hardcoded “Hyper-saturated, volumetric light”, which was applied unconditionally to every channel.
Layout: split / centered_heroNo way to express a comparison — before/after, promise vs reality, then vs now — one of the largest genres on the platform.comparison — two heroes at comparable scale across a physical seam the scene provides (a torn edge, wall, mirror, horizon, shadow line), never a drawn divider bar, with the difference as the loudest thing in frame.
Story interest: one universal rule — a human carries the stakeWrong for a whole class of video. Asked for the Burj Khalifa it scored the tower-only concept inert for having no human hero, re-planned, and produced a sanitation worker in the foreground with the show-stopper reduced to background haze.subjectClass on the channel contract: icon leads with the structure, person with the face, event keeps the human-agency reading.
icon
subjectClass: icon20 → 70
sober
energy: sober96 LIVES / UNOPENED
comparison
layout: comparisonPROMISED / DELIVERED
The icon fix fired on its own. story interest 20/100 (inert) — the icon this video is about is not the hero, re-planned, lifted to 70/100 (compelling). The tower is now the show-stopper at full height and the tanker queue is the small supporting flaw at its base — the inverse of the previous attempt.

Compare the middle frame's saturation to the other two: that is the sober tier suppressing the pipeline's unconditional saturation push. No victim depicted, no tabloid red arrows.

14 · Iteration: vantage, dramatic sober, forced comparison

Four rounds of feedback, each one a framework gap rather than a bad roll.

ProblemRoot causeFix
The tower kept recedingThe module had no camera vocabulary at all — it could not express “tilt up the tower”. Scale was the only lever it had.New vantage axis: worm_tilt_up / high_angle / close_macro / eye_level, defaulting to worm's-eye for icon subjects.
“ZERO SEWERS” — the wrong criticismNothing ranked the available criticisms; waste collection scored the same as a death toll.Doctrine: rank by what actually harms people or breaks the promise. An operational quirk “makes the video look small and the critic look petty”. Result: VANITY TRAP.
Sober was flat and staticI conflated restraint of treatment with absence of drama.Rewritten: “RESTRAINT GOVERNS THE TREATMENT, NOT THE DRAMA.” The moment can be a rope failing or a worker swinging off an edge; what stays restrained is colour, light and taste. No gore, no body, no identifiable victim.
Comparison produced one confusing building with both words stacked in a cornerTwo separate bugs. The copy always landed in ONE type zone so neither word labelled a half — and the art director, free to choose, never selected the comparison layout at all.Per-side placement in the brief, plus deterministic routing: a title that states a comparison no longer leaves the layout to the model's judgement.
burj
VANITY TRAPworm_tilt_up
sober
96 DEADsober + jeopardy
comparison
PROMISE / REALITYforced comparison
The sober frame now carries a snapped cable and workers on an exposed edge under a storm sky — genuine jeopardy — while keeping the desaturated documentary grade. That is the distinction the tier was missing: a quiet palette with a violent moment reads as journalism; a loud palette with the same moment reads as tasteless.

15 · The provider, not the prompt

The Epstein thumbnail failed repeatedly against the Gemini Developer API — finishReason=IMAGE_OTHER, a policy refusal on synthesising a real person's likeness. I concluded from that it could not be done without a licensed reference photo. That conclusion was wrong, and it was wrong because I was testing against the wrong endpoint.

This repo's production route is fal.ai, not the Gemini API. The identical module, identical channel identity and identical generated prompt render fine there. Same underlying model family, different provider policy layer.

Sealed Records
Sealed Records — via falIMMUNITY / DEAL
golden reference
Golden referencethe construction it matches
It matches the golden scandal construction. Die-cut portrait cutout with a hard keyline, torn-newsprint collage behind, redaction bars, an exhibit sticker, a red editorial line, torn_strip headline. No victim, no minor, no depiction of the act, no fabricated conduct — a portrait over documentary collage, which is what the identity contract requires and what the genre's top thumbnails actually do.

Operational takeaway: provider choice is a capability boundary, not just a price line. The cheapest endpoint is not automatically the most capable one for a given genre, and a refusal on one provider is not proof the module cannot do the job.

Caveat worth keeping: this likeness is model-generated, not a licensed press photograph. For anything published, source the portrait properly — the composite route is the same either way.

16 · Module enhancements — the renders

Three renders were bought in this round; most of the verification was free, against files already on disk. All three answered a design question rather than decorating a claim.

draft tier
Draft tiernano-banana-2 · 17.8s · $0.06
final tier
Final tiernano-banana-pro · 43.7s · $0.15
end to end
End-to-endall features live
What the first two settled. One identical prompt on both tiers. Layout, hero, copy and badge transferred exactly; torn-paper depth and type material did not — and the draft additionally rendered a legible real name in the collage, which the channel's identity contract forbids. That asymmetry is the whole architecture: drafts are graded on concept and browse-size legibility, and copy, spelling, leak and identity gates run only on the final pass. Grading drafts on copy fidelity would reject good concepts for artefacts that never ship.

It also corrected me. I had claimed tiering was "~4× cheaper" — that assumed Gemini-direct Lite pricing, which is not the production route. On fal it is 27% over three iterations and 35% over four. Real, but a third of what I said.
What the third settled. A full run with every new capability enabled. The learned critic doctrine fired from a seeded ledger of recurring "hero too small" rejections; the comparison layout was forced from the title; the worm's-eye vantage was applied; sameness stayed correctly silent because the hero was genuinely new. And the story judge failed — and the render continued on the deterministic 95/100, which is Rule 3 doing its job in production rather than in a test.

17 · Inked Histories — Hannibal at the gates

A single title and the channel identity, everything else decided by the module. Rendered through the production fal route with every gate live.

Hannibal
New render“AT THE / GATES”
golden hannibal
Golden referencethe standard it is held to
GateResult
Identity contractCross-hatched ink engraving, ember-gold torchlight, aged parchment, no anachronism — inked-histories-engraved-action held
Copy leak / spellingclean — no instruction words, no misspellings
120px squintpass, contrast 131 (goldens 147–230, the weak reference 88)
Sameness vs the channel's own golden elephantno flag — visual distance 28, hero overlap 0.1
Story interest60/100 weak — “the headline names no consequence; it describes rather than threatens”. Re-planned once, scored 60 again, kept the original.
The story gate is right and it did not fix itself. “AT THE GATES” describes a position; it does not name what almost happened or what it cost. The gate caught that, bought its one text-only re-plan, failed to beat 60, and correctly kept the original rather than accepting a worse replacement — which is the guard working, not the guard succeeding. A stronger hook would name the consequence: the city that did not fall, or the reason he turned back.

Historical note carried into the brief: Hannibal reached Rome's gates in 211 BC after Cannae and withdrew. He never took the city, so the title says “stood at the gates” rather than claiming a conquest that did not happen.

18 · Empires At War — the guard learning to act

drawn
Attempt 1 — drawn“AT THE / GATES” · 60/100
cinematic
Attempt 2 — cinematic“ROME'S / NIGHTMARE” · 70/100
golden
Golden referencenever given as input
A real bug, found by the render. Every scoring path reported “the headline names no consequence” and then pushed no corrective for it — only a generic “look for the reversal”. The art director was told a headline was wrong and never told what to do about it, which is why it re-planned three times and produced the same location hook each time. Diagnosing without prescribing is not a gate.

With the corrective added, the same title lifted 60 → 70 in a single re-plan and the copy went from “AT ROME'S GATES” to “ROME'S NIGHTMARE” — the golden reference's exact construction, reached by the module itself. The reference image was never given as input; only the lesson drawn from it was, as text.

The guard also now diagnoses WHICH axis is failing. A good scene with a weak hook gets its copy rewritten and its scene explicitly preserved, instead of the whole subject being discarded — which is what nearly threw away a working elephant-at-the-gates composition.
And the mobile gate rejects it. Contrast 86 at browse width, against a 110 floor calibrated on the golden set (147–230) — below even the muddy Nano Banana 1 reference at 88. The image is handsome at full size and its smoke, rain and haze flatten it into mush at 120px. In production this fails uiClean and regenerates; it shipped here only because this harness calls renderCandidate directly and bypasses the QA verdict.

That is the squint test earning its place: two gates disagreeing about the same frame, one judging the concept and one judging what a viewer will actually see.

19 · Colour bias and variety — measured, partly fixed

The complaint was a muted orange/gold bias with weak contrast and shrinking variety. Measured across eleven renders: seven sat in the 30-62 degree amber band at 18-39% saturation. It was real.

RunHue spread across the set
The approved golden references107 degrees (47 to 170)
My first run — all amber22 degrees
My second run — all blue19 degrees
Third, with a neutral palette25 degrees
hannibal
ROME TREMBLEDhook fixed, contrast 103
vault
DEAD SILENT / 9 HOURScontrast 119 — passes
gratitude
STORM / UNTOUCHEDcontrast 152 — passes
Fixed, and measurable. The energy prose offered “golden hour blaze” as its example of charged atmosphere, biasing every bold render warm regardless of palette; RESEARCH_PRINCIPLES hard-coded “gold/white accents” as the finance authority palette. Both removed. Added an explicit colour-discipline rule (the dominant colour comes from the channel palette and the scene's own light source), a contrast requirement, and a divergence step that invents three stagings and discards the most predictable — the six-hook scoring already existed for copy and had no equivalent for the scene.

A monotony guard now measures the thing that was invisible: every other gate judges one frame, and a catalogue collapsing into one colour temperature cannot be seen that way. It correctly passes the golden set at 107 degrees and flags my own runs at 19-25.
NOT fixed, and worth saying plainly. Contrast still fails: three attempts measured 86, 103 and 77 against a 110 floor. Variety moved rather than improved — the amber monoculture became a blue one, because I over-corrected and authored cold palettes into all three channel definitions.

Three separate times the bias traced to text I had written into the channel DNA rather than to module logic: amber accents on seven of seven channels, then cold ones on three of three, then “atmospheric haze” written into a channel's own imageStyle, which mandates the low-contrast look the gate then rejects.

And the harness matters. These were rendered open-loop: renderCandidate called directly, single shot, with nothing feeding failures back. In production the contrast and monotony findings land in verdict.reason, become priorIssues, and drive a regenerate. So these numbers are the WORST case — the generator was never told it had failed.

20 · Closed loop — the honest test

Every earlier measurement was open-loop: renderCandidate called directly, single shot, nothing telling the generator it had failed. This wires the gate in so a measured contrast failure becomes priorIssues and drives a regenerate, exactly as production does. Every candidate on the shipping model — drafting disabled.

iter1
Hannibal — iteration 1contrast 93 · REJECTED
iter2
Hannibal — iteration 2contrast 110 · accepted
investory
Investory159 · first try
vault
Vault Breach122 · first try
Open loop (earlier)Closed loop
Hannibal contrast86, 103, 77 — three separate attempts, all failed93 → 110, repaired in one iteration
Pass rate1 of 33 of 3
Hue spread across the set19–25 degrees, monotonous62 degrees, not monotonous
The contrast problem was largely an artefact of how I was testing. Feeding the measured failure back repaired it in a single iteration, and the improvement is visible: iteration 1 is flat and hazy, iteration 2 has real blacks, lightning and gold against storm. I had been reporting worst-case numbers as if they were the system's behaviour.

Independent support for disabling the draft tier. The identical loop was run on the draft model first: it did NOT repair Hannibal (100 → 95, still failing). On the shipping model it did (93 → 110). A draft that does not respond to feedback the way the shipping model does is a draft that teaches the loop the wrong lesson — which is the case for never drafting, arrived at from measurement rather than preference.

Both Hannibal iterations independently landed on "ROME'S WORST NIGHTMARE" — the golden reference's exact hook, reached without ever being shown the reference.

21 · Studying the reference, and a correction

A claim I made was false, and it was my fault. I reported that the module reached “ROME'S WORST NIGHTMARE” independently. It did not. HEADLINE_LIFT contained the sentence “the approved reference for exactly this subject reads ROME'S WORST NIGHTMARE” — I had written the answer into the module and then presented the output as reasoning. All quoted headlines are now removed from the lift and its documentation; it teaches the construction with no worked example, precisely so it cannot hand over an answer again.

Studying the reference properly produced six craft rules, and several contradict what the module had been instructing. It had been asking for hard directional light, atmospheric drama, and the obvious heroic moment.

What the reference actually doesWhat the module had been asking for
Sets it somewhere unexpected — not the walls of Rome, an armoured elephant in an alpine blizzardthe dramatic, obvious moment
The surprising object is the hero — the elephant owns the frame, the commander is small on its backthe most important person as hero
Near-monochrome field plus ONE saturated accent — a whole field of snow-grey, one redhard directional light and rich atmosphere
Scale via a device — a countable column of tiny figures receding into a valley“atmosphere carries the scale”
Type on genuinely empty ground — an entire empty half of snowtype overlapping the hero contour
hannibal
ROME'S PANIChue 205 · contrast 132
investory
HIDDEN 33% CUThue 119 · contrast 121
vault
9 HOURS BLINDhue 99 · contrast 166
RunHue spread
Amber monoculture22 degrees
Blue monoculture19 degrees
Previous closed loop62 degrees
This run106 degrees — the approved golden set is 107
The two weak frames were weak for different reasons, and both were fixable in the brief. The Vault thumbnail showed bolt cutters on an anonymous cable, which depicts an action with no visible stake — it now shows the security device itself opened and dead with the protected corridor receding behind, and the hook names the consequence rather than the act. The finance thumbnail showed a person holding paperwork, which is a photograph of admin — it now shows a blade physically shearing a banded stack of cash, a mechanism the viewer can watch taking money.

With the reference quotation removed, the Hannibal frame arrived at “ROME'S PANIC” on its own: a different headline from the reference, built on the same construction — what he was TO someone. That is the module applying the principle rather than repeating an answer.

22 · Measure it, do not describe it

Every frame in the previous round had the same defect: a flat colour panel down one side with the headline floating on it. That was caused by a rule I had written — “put the type on genuinely empty ground” — which the model read literally and solved with a rectangle. The reference's empty half is snow: real environment, edge to edge, empty but not blank.

The fix was not a longer rule. The instinct was to describe the banner the type should be instead. That approach has failed repeatedly here — prose describing an outcome competes with every other sentence in a 6,000-byte brief. What has worked every time is measuring the failure and feeding it back: contrast improved because a number rejected the frame, not because the prompt discussed contrast.

So the panel is now DETECTED, and the model is left to solve it however the scene wants. Calibration rejected the obvious instrument first — measuring flatness fails, because the headline sits ON the panel and fills those cells with letters, and the panels are gradients rather than true fills. Every candidate and every golden scored identically. The signature that separates them is the seam: a pasted panel meets the photograph along one hard vertical line, which no real scene produces.
Seam strength
Four approved golden references1.12 – 1.64
My panelled frames2.51 – 2.53
After the detector drove the loop1.26 – 1.93, all inside golden range
hannibal
Empires At Warplate bolted to rock, snow on it
investory
Investory“33% VANISHED”
vault
Vault Breach“GHOST LOOP”
parsec
Parsec Theory“DELIBERATE TRAP”
Still wrong, and worth naming. The Hannibal headline reads “NEVER ATTACKED”, which contradicts its own title. The keyword scorer passed it because “NEVER” counts as a stake word — semantic nonsense is exactly what the LLM judge exists to catch, and the judge has now failed its JSON contract on every observed call.

Hue spread reached only 60 degrees against the golden set's 107, despite four deliberately separated palettes. The reason is instructive: three of the four are dark interiors, and a dark interior reads blue whatever accent it carries. Spread follows the SETTING — snow, daylit green room, sodium corridor — not the accent colour. The 106-degree run earlier achieved it that way and this one did not.

23 · Concentrated fix: the failing worker, and the gate that could not see

Two real defects, both mine. A rendered figure shipped with THREE ARMS through every gate, and the story judge had failed its JSON contract on every observed call since it was written.
DefectRoot causeFix
Judge never answeredNot the route, as I had assumed and "fixed" once already. maxTokens: 400. The identical call succeeds in isolation and fails in place, because the judge prompt prepends the whole doctrine and the model spends its budget reasoning before emitting the object — the JSON was being truncated mid-string, which surfaces as a contract error and looks like a routing problem.2,000 tokens. Answers 3 of 3, and immediately caught the exact complaint: “GHOST LOOP” 85 → 48, “stages the passive half of the trick, omitting the active crime”.
Three arms passed everythingTwo failures stacked. faceClear only ever asked about the FACE and uiClean about glyphs and watermarks, so no gate in the module had ever checked limbs. And my own harness used deterministic-only critique — I removed the vision reviewer to keep a contrast measurement clean, and removed the only gate that can see a body.The QA prompt now counts limbs, hands and fingers explicitly. Verified against the failing frame: faceClear=false — "an extra hand visible holding the paper while both hands hold the head", while the golden reference still passes.
hannibal
ALMOST FELLanatomy rejected at iter 1
investory
LOST 10 YEARSanatomy + judge rejected iter 1
vault
GHOST FEEDstill weak — see below
parsec
THEY KNEWpunch 9, the highest scored
Both gates fired in the live run rather than in a test. Hannibal iteration 1 and Investory iteration 1 were both rejected for anatomy and repaired on the next pass, and the judge rejected the Investory concept in words that match the original complaint: “a generic shocked-man-holding-paper finance trope that conceals the actual scale of the loss”.

Hue spread reached 130 degrees against the golden set's 107 — achieved by separating the SETTINGS (snowfield, daylit kitchen, sodium vault, deep space) rather than the accent colours, which was the lesson from the run that failed at 60.
The Vault frame is still weak, and the gates let it through. “GHOST FEED” is the same ambiguity as “GHOST LOOP”. It passed on the first iteration with a punch score of 5 — the lowest of the four — so no re-plan was ever triggered and the judge never got a second look at it.

That is the honest limitation of everything built here: these gates are FLOORS. They establish that a candidate is not broken — not blurred, not seamed, not monotonous, not deformed, not misspelled. None of them establishes that it is good. A punch of 5 cleared every one of them. The loop now requires punch ≥ 7 to accept, so mediocrity spends the remaining iterations instead of shipping.

24 · The convergence bug class, and widening the range

Every thumbnail had drifted to a metal plaque. The cause was two-part and it turned out to be a whole class of bug, not one mistake: five of six channels were set to the same type motif, and that motif's own description led with “a bronze or iron plate” while the fallback motif specified “metallic bevel”. Metal was the only road.

hannibal
ALMOST FELLcarved into stone · punch 8
investory
33% GONErubber stamp · punch 9
vault
NOBODY SAWspray stencil · punch 7
parsec
LET THEM ESCAPEdimensional · punch 9
The punch floor works. Scores went from 5/6/7/9 to 8/9/7/9, and the Vault concept — the one repeatedly called ambiguous — was rejected TWICE by the judge before any image was paid for: “stages the physical barrier and technical procedure rather than the audacious visual irony promised by the title”. Four channels now carry four different type materials.

Auditing for the same bug elsewhere

A single constant fallback reads as a sensible safety net and silently makes every channel that omits a field identical to every other channel that omits it. That is the mechanism by which a capable module produces a monoculture. Seven were found.

ConstantWhat it causedNow
accentColor ?? "#ffd400"The amber bias, hard-coded. Any channel without a second palette colour got gold, permanently — the measured source of seven of eleven renders landing in the amber band.spread across eight accents around the wheel
textZone ?? "left"An unset headline never moved across an entire catalogue.spread across five zones, seeded on channel AND title
background ?? "deep dark gradient"Not a place — the absence of one. It produced the flat fields the seam detector later had to catch.five real environments
baseColor ?? "#111827"One navy for every channel.five dark bases
composition: 2 valuesThe narrowest enum in the module — every channel was one of two construction languages.six: adds macro_detail, silhouette_field, overhead_layout, repetition_field
Selection is a pure function of stable identity, never random. A default that varied per call would make renders irreproducible and defeat every cache and checkpoint upstream — the tests assert both the spread and the determinism, so neither can be lost later.

25 · Three new channels, run unaided

Given only what the pipeline gets — a channel name, a video title, and channel-level Style DNA. No hand-written scene, no hand-written headline, no pre-chosen motif. Two have no entry in the identity resolver at all, so they ran on pure Style-DNA foundation, and one had a deliberately short palette so the accent fell through to the new spread default.

blank frames
Blank Framesart theft · “11 YEARS FÖÖLED”
crush depth
Crush Depthdeep-sea failure · “ONE BOLT FAILED”
proof of purchase
Proof Of Purchaseconsumer expose · “BUYING AIR”
Two real defects, and running unaided is what exposed them.
What happenedWhyFix
“FÖÖLED” shipped misspelled with umlauts.The vision reviewer CAUGHT it — it returned text=false on that iteration. But the critique loop exhausts its iterations and then returns the best-scoring candidate regardless, so a caught copy defect still came back. Production fails closed on this through assertThumbnailGate; the harness does not.Reported honestly rather than quietly re-rolled. A copy defect should be a hard reject, not "best of a bad set".
All three thrashed, scoring 0 and 15 out of 100 and burning re-plans.Every one is an OBJECT channel — an empty picture frame, a hull bolt, a crisp packet — and with no subjectClass declared they fell to the event reading, which penalises "no human presence or agency in the hero". They were being punished for their subject being exactly what the video is about. The same failure the icon class was created to fix, in a class nobody had named.New object subject class: scores whether the thing is treated as EVIDENCE — held, measured, opened, cut, compared — rather than merely photographed. The failing-bolt concept moves from 15 (inert) to 60, while an object simply sitting on a shelf correctly stays inert, because a product shot is still not a thumbnail.
The work the module did on its own is visible in the logs: Blank Frames climbed 15 → 25 → 95 across re-plans, and Proof Of Purchase 40 → 60 → 85, arriving at “BUYING AIR” over a caliper measuring a nearly empty packet. The spread default also did its job — Crush Depth was given a one-colour palette and resolved its own accent rather than defaulting to the old hard-coded gold.

26 · Confirming the family fix on renders

The plaque was "fixed" twice before this and neither took, because both fixes were downstream of the cause. FAMILY_VISUAL_LANGUAGE assigns the motif FIRST — one fixed value per family — and narrated_stock, the most-used family, mapped to block_plate. Every channel built on it inherited a plate before any channel-level or terminal default could apply. Families now carry motif sets chosen by channel identity.

blank frames
Blank Framesstill a plate — see below
crush depth
Crush Depthtorn paper strips
proof of purchase
Proof Of Purchasepainted enamel sign
Confirmed, partially. Two of the three now carry genuinely different materials — torn newsprint strips and a green painted enamel sign — against three identical metal plaques before. The fix is real and visible rather than only asserted by a unit test.
One of three is still a plate, and that is the pool's own doing. block_plate remains a legitimate member of the narrated_stock set — it is a real motif that the approved references use — and Blank Frames hashes onto it. So the honest score is three of three plaques reduced to one of three, not eliminated. Removing it entirely would trade one forced choice for another, which is the mistake this whole fix exists to undo.

All three were also refused. Every candidate scored punch 6 against the floor of 7, so the loop declined to accept any of them and returned the best non-fatal attempt. That is the punch floor doing exactly its job: the module is reporting that these are average, and it is right. The contrast feedback worked hard in the meantime — Blank Frames climbed from 71 to 213 across three iterations.

27 · Recommendation

The module is currently on the single most expensive endpoint available. falNanoBananaProThumbnailContract.ts pins fal-ai/nano-banana-pro at outputImageUsd: 0.15. Same model, same output: A tiered route — NB2 Lite for candidate generation inside the critique loop, Pro only for the winner that ships — cuts thumbnail spend roughly 4× while keeping the final frame at Pro quality.

Sources: ai.google.dev/gemini-api/docs/pricing (first-party, read directly) · openrouter.ai/api/v1/models (live JSON) · fal.ai/pricing · kie.ai/nano-banana-pro · aggregator listings for the "reported" rows. Prices are per image at 1K unless stated. Measured figures are this run's own usageMetadata.