Skip to content
All work
2026Solo — product, frontend, AI integration

Image Banana

AI image editor — mask a region, describe the change, generate

Source-available · runs locally with your own Gemini key

Next.js 16 · React 19 · TypeScript · Tailwind CSS 4 · Zustand · Vertex AI (Gemini image) · shadcn/ui · Radix UI

Image Banana
Four real outputs from the app's AI filters — Ghibli, Toonify, Cyberpunk, Oil painting
Real outputs — filters generated in the app
Photo converted with the Ghibli filterPhoto converted with the Toonify filterPhoto converted with the Cyberpunk filterPhoto converted with the Oil painting filterGhibli

01

Context

Image Banana is an image editor where the editing tool is a sentence. Upload an image, brush or box a mask over the part you want changed, describe the result, and Gemini regenerates that region. It also expands images to any aspect ratio with AI-filled content, applies one-click style filters, and keeps a full undo history.

The interesting part isn't calling an image model — it's making generative editing feel like a normal editor: the mask is a first-class object, every generation is an undoable step, and the UI stays responsive while a request that takes seconds is in flight.

02

The hard parts

01

Treating generative edits like undoable document operations

Problem
Every AI edit produces a new version of the image, and users try variations constantly. Naive state — replace the image, lose the old one — makes comparison impossible and turns one bad generation into a restart.
Decision
The editor keeps an ordered history of image snapshots in Zustand with a lightweight thumbnail strip, so undo/redo is a pointer move, not a reconstruction. Masks are stored as their own layer and composited to a PNG at generation time, never baked into the base image.
Tradeoff
Snapshots use more memory than a command log would, and large images force a cap on history depth. In exchange, undo is instant and correct with zero re-render logic.
Outcome
Users can generate five variations, flip between them, and undo to the original without the app ever losing a version.

02

A generation pipeline that doesn't freeze the editor

Problem
Image generation is slow — seconds, not milliseconds. The canvas, brush, and history all need to stay interactive during a request.
Decision
Generation is an API route that receives the base image, the composited mask, and the prompt, and calls Vertex AI's Gemini image model server-side. The client treats the result as just another history entry; pending state is a single flag that swaps the canvas for a progress state without blocking the rest of the UI.
Tradeoff
There's no job queue — one generation at a time per client, and a page refresh mid-request loses that attempt. For a single-user editor that's the right size; for a team product I'd move generations behind a queue with server-side job state.
Outcome
Editing flow stays: mask, prompt, generate, keep or undo — with the slow part visibly bounded instead of mysterious.

03

What I'd change

If this grew past a personal tool, the first structural change would be a generation queue with persisted jobs, so a closed tab doesn't lose work and multiple users don't compete for one API quota.

04

Stack

Next.js 16 · React 19 · TypeScript · Tailwind CSS 4 · Zustand · Vertex AI (Gemini image) · shadcn/ui · Radix UI