Runs on your device · WebGPU

Turn a screenshot into text

Drop in a screenshot or a photo and get its words as text you can edit, copy or save, line breaks and paragraphs kept. A small vision model reads it in your browser, so the image never leaves this page.

Downloads once: Florence-2 base (Microsoft, MIT licence), shared with the alt text generator: about 360 MB with WebGPU, 230 MB without

How it works

  1. 1Add screenshots or photos

    Up to 10 at a time. Paste a screenshot straight from the clipboard, or choose files: PNG, JPEG and WebP work in every browser.

  2. 2Empty margins go, tall images are split

    Blank borders are trimmed, and a tall phone screenshot is cut into strips between lines of text, so the small type isn’t squashed.

  3. 3Each line is read where it sits

    Florence-2 reads the text along with a box for every line. The boxes put the line breaks back and keep a gap between paragraphs.

  4. 4Check, copy or download

    Fix anything misread in the box, then copy it, save it as a .txt file, or copy every image’s text in one go.

What it reads well, and what it doesn’t

The model is good at clear printed type on a plain background, which covers most screenshots. It gets worse as the text gets smaller, fancier or more crooked. Here is what to expect:

ImageHow it does
Phone screenshot of a caption, comment thread or DMGood. Line breaks come out as they are on screen.
Quote card or text-heavy carousel slideGood, if the font is a normal one. Script fonts confuse it.
Photo of a menu, sign or posterFair. Straight-on and sharp helps; angles and glare cost words.
Wide desktop screenshot with small textWeak. Crop it to the part you need first.
HandwritingPoor. Expect missing and invented words.
Hindi, Arabic, Chinese and other non-Latin scriptsNot supported. It reads printed Latin-alphabet text, English best.
Tables and spreadsheetsRows come out as lines; columns don’t line up.

Tip: Small type reads better when it fills more of the image. If only one part of a big screenshot matters, crop to it before you add it.

Why tall screenshots are read in strips

Florence-2 resizes every image to a 768 × 768 square before reading it. An iPhone screenshot is 1170 × 2532 pixels, so in one piece its height would shrink to about 30% while its width shrinks to 66%: every letter gets squashed flat. Cut into two strips of roughly 1170 × 1266, each is scaled evenly to about 60% and the letters keep their shape.

The cuts are made through the gaps between lines, found by looking for rows with no detail in them. When there is no gap nearby, as in a photo of a page, the strips overlap slightly and any line read twice is kept once.

Ways people use it on Instagram and Facebook

  • Copying a caption or a long comment you screenshotted, to quote it with credit or reply to it properly.
  • Turning a carousel of text slides into a draft for a newsletter, blog post or Facebook post.
  • Getting the words off an event poster or a flyer someone sent you in a DM.
  • Saving the questions people asked in a Story question box from your own screenshots.
  • Pulling the text out of your own old quote graphics to reuse or translate.

Other people’s DMs and comments are theirs: ask before you repost what someone wrote to you privately.

Questions people ask

Are my screenshots uploaded?

No. The model downloads to your browser once, and the images are read there. Nothing goes to DMFast or any server, which matters when the screenshot is of private messages.

Does it work with handwriting?

Not well. The model was built for printed text, so handwritten notes come back with missing or invented words. Neat block capitals do better than joined-up writing.

Can it read Hindi, Arabic, Chinese or other scripts?

No. It reads printed text in the Latin alphabet, English best. Other scripts come out as nonsense or not at all, so use a tool made for that language.

How much text can one image have?

Each strip is read with a budget of about 50 lines, and a tall image is split into as many as 12 strips. If a page is very dense, it says so when text may be missing: crop it into parts and read each one.

Why is the first run slow?

The model has to download first: about 360 MB on browsers whose WebGPU supports half-precision maths, or 230 MB for the WebAssembly version. It’s kept for next time. After that, a screenshot takes a few seconds with WebGPU and a minute or more without.

More free tools

Related guides

Answer every comment and DM, automatically

When someone comments a word like GUIDE, DMFast replies and sends them your link in a DM. Free plan, no card.

Start free