TEST RECORD · ← all field reports

FIELD REPORTS/IMAGE/ISSUE #79

Create images and edit them from plain written instructions

Image generation and image editing usually mean two different models, and the good editors have gotten huge. Microsoft's new release does both jobs in one compact family, and it is small enough that a single consumer-class GPU covers everything in this issue.

MODELMage-Flow
PUBLISHEDJuly 23, 2026
READ TIME3 min
TESTED BYNeural Expedition
CATEGORYIMAGE

Field notes

01What it does

Mage-Flow is a pair of small models that share one brain. The generation side turns a text prompt into an image at any size from 512 to 2048 pixels and any aspect ratio, including extreme banner shapes, without cropping tricks. The editing side takes an existing image plus a written instruction: replace the background, restyle the whole frame, restore a damaged photo, or combine sources, as in "blend the object from image 2 into image 1". Both sides render readable text inside images, in English and Chinese.

What makes it worth a look is the size. Editing at this quality has meant models five to eight times larger; on published editing tests Mage-Flow matches or beats the well-known 20B editors while running in 18 to 20GB of GPU memory. The Turbo variants cut generation to four steps, which lands around half a second per image on a data-center card, fast enough to iterate on prompts like you iterate on sentences.

02How to try it

There is no hosted demo yet, so this one is for readers with an NVIDIA GPU of around 20GB. The GitHub repo installs with pip and includes a ready-made local web UI: run mage-flow-app and you get a browser page with a Text to Image tab and an Image Edit tab, weights downloading automatically on first use. For a first test, load a pet or product photo in the Edit tab, ask for "replace the background with a field of sunflowers", then give the result a second instruction and see how well the edits chain.

03Caveat

The release is days old: weights are up and complete, but there is no demo Space, so trying it requires a local GPU and a Python install rather than a one-click page. The comparison tables are Microsoft's own runs, and the largest open editors still lead a few of the boards, so treat "matches the 20B models" as close, not settled.

04What you can do with it

  • Turn one product shot into a set of variants: new backgrounds, new lighting, same product.
  • Restore and sharpen old or damaged photos with a written instruction.
  • Generate tall or wide poster and banner shapes directly instead of cropping a square.
  • Make posters and mock ads with readable text baked into the image.
  • Blend a subject from one photo into the scene of another.

Try the demo

View model page

Read this issue on neuralexpedition.com →