<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Redspring Blog</title><description>News, updates, and notes on our process from the Redspring team.</description><link>https://redspring.xyz/</link><language>en-us</language><item><title>Narrative Games in the Age of AI</title><link>https://redspring.xyz/blog/narrative-games-in-the-age-of-ai/</link><guid isPermaLink="true">https://redspring.xyz/blog/narrative-games-in-the-age-of-ai/</guid><description>What we learned building a murder mystery with Lemony Snicket: on authority, constraint, and using AI as one wire in the sculpture.</description><pubDate>Thu, 23 Jul 2026 12:00:00 GMT</pubDate><content:encoded>&lt;h2 id=&quot;what-we-learned-building-a-murder-mystery-with-lemony-snicket&quot;&gt;What we learned building a murder mystery with Lemony Snicket&lt;/h2&gt;
&lt;p&gt;For most of human history, there was a stark line between audience and story. The audience heard the story, read the book, and eventually watched the movie. The story happened while you observed.&lt;/p&gt;
&lt;p&gt;Interactive storytelling changed that. There were choose-your-own-adventure books (remember those?), then RPG and narrative games. Now, AI is changing things again.&lt;/p&gt;
&lt;p&gt;This week, The Atlantic released an &lt;a href=&quot;https://www.theatlantic.com/games/lemony-snicket-suspicious-incident-dubious-park/&quot;&gt;interactive murder mystery&lt;/a&gt; co-developed by Redspring, written by Lemony Snicket and illustrated by Michael Kupperman. Players step into the role of detective in a closed park, talk to a cast of eccentric characters, and try to figure out who committed the murder. The characters are powered by AI, but the story architecture is still human.&lt;/p&gt;
&lt;div class=&quot;lg:relative lg:left-1/2 lg:-translate-x-1/2 lg:w-[min(72rem,calc(100vw-4rem))]&quot;&gt;
&lt;p&gt;&lt;a href=&quot;https://www.theatlantic.com/games/lemony-snicket-suspicious-incident-dubious-park/&quot;&gt;&lt;img alt=&quot;The scene of the crime in Dubious Park: a park full of suspects, from the Zookeeper at the Reptile House to Detective Aiden deCamp standing over the victim&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot; fetchpriority=&quot;auto&quot; width=&quot;2522&quot; height=&quot;1658&quot; src=&quot;https://redspring.xyz/_astro/dubious-park.CvpCeTyN_1FDdI2.webp&quot; &gt;&lt;/a&gt;&lt;/p&gt;
&lt;p class=&quot;text-gray-700 text-sm text-center&quot;&gt;Dubious Park, illustrated by Michael Kupperman&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;We sat down with &lt;a href=&quot;https://www.linkedin.com/in/caleb-madison-89198853/&quot;&gt;Caleb Madison&lt;/a&gt;, former director of games at The Atlantic and the person who originally conceived of the project, to talk about what we learned over the last year building a human-authored game with characters you can interrogate in real time.&lt;/p&gt;
&lt;h2 id=&quot;the-problem-of-authority&quot;&gt;The problem of authority&lt;/h2&gt;
&lt;p&gt;Caleb has been making games since he was fourteen, mostly crosswords. To him, the appeal is simple: games create a bounded arena with a final, true answer (something we don’t often find in the real world). A game designer leads a player towards a certain outcome, ideally in a dynamic and entertaining way.&lt;/p&gt;
&lt;p&gt;We were all captivated by the idea of bringing AI into that world. Can we use this new medium as a way to tell stories that you can engage with, that feel surprising, and that pique your curiosity because they evolve?&lt;/p&gt;
&lt;p&gt;That idea runs directly into a fundamental challenge of AI today: uncertainty.&lt;/p&gt;
&lt;p&gt;In a traditional story, the author is the source of truth. If the novel says it was a dark and stormy night, the reader accepts it. But in an AI-mediated experience, where you discover the story through conversation, who is speaking with authority? The author? The system? The character? The model?&lt;/p&gt;
&lt;p&gt;This became our central design problem. The original concept had psychologically layered suspects who could be caught in lies. It was an Agatha Christie-style cross-examination where players would triangulate the truth by finding contradictions. Caleb described why that didn’t work: “If this one thing is a lie, then how many other things are lies? It becomes a hall of mirrors.”&lt;/p&gt;
&lt;p&gt;Lemony Snicket, it turned out, understood this intuitively. His creative instincts pushed the game away from forensic interrogation and toward something more expansive. We were soon building an oddball ensemble in a fully realized world, where the characters had distinct personas and the players were fully inhabiting the space.&lt;/p&gt;
&lt;p&gt;Lemony helped us realize that the solution was to give the characters a strong enough identity that players trust them even when they surprise you.&lt;/p&gt;
&lt;h2 id=&quot;building-the-game&quot;&gt;Building the game&lt;/h2&gt;
&lt;p&gt;Early in the project, we developed a character builder that let an author train a character through dialogue, refining its persona until it feels consistent. It was a good idea, but it wasn’t quite right for this project.&lt;/p&gt;
&lt;p&gt;Working with a writer like Lemony Snicket, who already had a rich and specific creative universe, the character builder felt stifling. Instead, the best approach was to give him room to define the characters in longhand, in his own voice, with his own instincts. Our job became translating those authored characters into personas that could hold up inside an interactive system.&lt;/p&gt;
&lt;p&gt;That required a lot of testing. We put the game in front of as many people as we could, watched their conversations, and learned which paths worked and which led nowhere. Coordinating characters, world knowledge, player knowledge, and the investigatory path toward the solution were all refined through that process until the experience felt natural.&lt;/p&gt;
&lt;p&gt;Players never see most of what makes the game work, and that’s part of what makes it a success. The fact that it feels simple is a big achievement. Underneath that simplicity is a system that we iterated on for months.&lt;/p&gt;
&lt;h2 id=&quot;ai-is-just-one-wire-in-the-sculpture&quot;&gt;AI is just one wire in the sculpture&lt;/h2&gt;
&lt;p&gt;The question Caleb gets asked most about this project is some version of: isn’t AI just a way to cut corners? His answer, in this case, is definitely not.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;“I see creative work and its relationship to tools or the medium of creativity as pretty consistent, whether it’s oil paint, film, or the computer. The amount of work that the creative person, or the people on the back end, put into the thing is directly proportional to how good it is. There’s no hack or way of getting around that.&lt;/p&gt;
&lt;p&gt;We didn’t use AI as a shortcut for creativity here. Our desire, our curiosity, was to figure out how to integrate AI and LLMs thoughtfully into a larger storytelling experience.&lt;/p&gt;
&lt;p&gt;This game would not be fun without Lemony’s story, which he put so much work and effort into; without Michael Kupperman’s beautiful illustrations; and without the hours and hours of collaborative work that all of us have done to fit this together in a way that feels satisfying.&lt;/p&gt;
&lt;p&gt;My sense is that a game, in many ways, is like a big, complex digital sculpture. AI is like a little thing hanging on one of the wires in the sculpture. It’s not the base of the sculpture at all. It’s ornamental in a way that I hope lends itself to satisfaction and play.&lt;/p&gt;
&lt;p&gt;This is an attempt to use AI as a medium in and of itself: as a material integrated into other creative materials, the visual arts and the storytelling arts, to be one part of the mosaic of the expression of a story and an experience.”&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id=&quot;ai--creativity&quot;&gt;AI ≠ creativity&lt;/h2&gt;
&lt;p&gt;Towards the end of our conversation, Caleb voiced the most useful thing we took away from this project: making AI art means chiseling away at a vast space of possibility.&lt;/p&gt;
&lt;p&gt;AI can generate almost anything. Creativity comes in when deciding what AI &lt;strong&gt;won’t&lt;/strong&gt; do in a specific context. For the game, it was what persona a character holds to, what knowledge it has access to, and what the boundaries of the world are. We found that the more precisely we defined those constraints, the more the experience felt authored rather than generated.&lt;/p&gt;
&lt;p&gt;The principle of constraining AI applies beyond games. Any time you build something with AI at its center, the quality of the output is determined by the choices you make about what the model is and isn’t allowed to do. Your choices as the system architect are the creative center of the output.&lt;/p&gt;
&lt;h2 id=&quot;a-peek-into-a-bigger-room&quot;&gt;A peek into a bigger room&lt;/h2&gt;
&lt;p&gt;In the gaming world, we’re still in the early stages of AI integration. Caleb compared the current moment to the Lumière brothers projecting footage of a train arriving at a station: audiences ran because they didn’t yet understand the boundaries of the medium. (Probably an &lt;a href=&quot;https://www.atlasobscura.com/articles/did-a-silent-film-about-a-train-really-cause-audiences-to-stampede&quot;&gt;urban legend&lt;/a&gt;, but the comparison stands). Today, audiences and users are still figuring out where AI ends and where the human begins. That confusion is going to take time to resolve, and that’s why we’re continuing to experiment.&lt;/p&gt;
&lt;p&gt;Building this game left us with a set of better questions. What does authorship mean when a reader can interrogate the characters? What does truth mean in a story where the player discovers it through conversation? What does constraint mean when the system could, in theory, say almost anything?&lt;/p&gt;
&lt;p&gt;We don’t have tidy answers to any of those. But we know how to ask the questions more precisely than we did before. That feels like the right place to be when a medium is this young. We’ll be watching this space closely. As Caleb said, “I’m very fascinated by interactive storytelling, puzzles, games, storytelling proper, the classic storytelling of the 20th century and previous eras, and the intersection of all of them in the beautiful intermedia future that we are all barreling toward at breakneck speeds.”&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://www.theatlantic.com/games/lemony-snicket-suspicious-incident-dubious-park/&quot;&gt;&lt;em&gt;Lemony Snicket’s Suspicious Incident in Dubious Park&lt;/em&gt;&lt;/a&gt; is live. Go solve a murder!&lt;/p&gt;</content:encoded></item><item><title>OCR’ing 100,000 pages with open-source VLMs on Modal</title><link>https://redspring.xyz/blog/vlm-ocr-bench/</link><guid isPermaLink="true">https://redspring.xyz/blog/vlm-ocr-bench/</guid><description>We OCR&apos;d 100,000 pages with open-source vision-language models in &lt;1hr for $223, roughly 9–27× cheaper than the comparable-quality proprietary APIs.</description><pubDate>Thu, 25 Jun 2026 08:00:00 GMT</pubDate><content:encoded>&lt;p&gt;We wanted to answer a simple question: what does it actually take to OCR a large document corpus using open-source vision-language models?&lt;/p&gt;
&lt;p&gt;We picked a workload of 100,000 pages. That number was mostly arbitrary, but it was large enough to be interesting, and large enough that we had to figure out how to make it not a financially ruinous endeavor.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;tl;dr - three things that surprised us:&lt;/strong&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Self-hosting open-source OCR is cheaper and far less painful than you might think.&lt;/strong&gt; Our full 100k-page run with &lt;code&gt;dots.ocr-1.5&lt;/code&gt; finished in 56.5 minutes for $223 — about $2.27 per 1,000 pages. Staging weights, renting GPUs, and standing up the serving stack took hours, not days; &lt;a href=&quot;https://modal.com&quot;&gt;Modal&lt;/a&gt; did most of the heavy lifting, and we never touched a CUDA driver or a container registry.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The cheapest GPU per second may be the slowest to finish and unacceptable at scale.&lt;/strong&gt; An &lt;code&gt;L4&lt;/code&gt; can land in the same dollars-per-page band as an &lt;code&gt;H100&lt;/code&gt; while taking 5–6× longer to drain the queue. Cost-per-page doesn’t matter if you blow your SLA.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Benchmark your workload, not the leaderboard.&lt;/strong&gt; Two public-ranking favorites lost once we measured them on the shape of &lt;em&gt;our&lt;/em&gt; data.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;This post is for engineers, technical founders, and anyone who has mostly lived on hosted APIs from the big model providers and wants a concrete sense of what self-hosting actually looks like in practice.&lt;/p&gt;
&lt;figure class=&quot;not-prose my-6&quot;&gt;
  &lt;a href=&quot;https://redspringxyz.github.io/ocr-results-viewer/&quot; target=&quot;_blank&quot; rel=&quot;noopener&quot;&gt;
    &lt;img src=&quot;https://redspring.xyz/ocr-side-by-side.png&quot; alt=&quot;The OCR results viewer showing a rendered source page next to OCR output from chandra-ocr-2, dots-ocr-1.5, nemotron, and nemotron-fp8, with per-model latency and token counts.&quot; class=&quot;w-full rounded-lg border border-gray-200&quot; loading=&quot;lazy&quot;&gt;
  &lt;/a&gt;
  &lt;figcaption class=&quot;mt-2 text-center text-sm text-gray-500&quot;&gt;Side-by-side viewer to compare every model&apos;s OCR output against the rendered original. &lt;a href=&quot;https://redspringxyz.github.io/ocr-results-viewer/&quot; target=&quot;_blank&quot; rel=&quot;noopener&quot; class=&quot;underline&quot;&gt;Check it out →&lt;/a&gt;&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h3 id=&quot;why-open-source-models&quot;&gt;Why open-source models?&lt;/h3&gt;
&lt;p&gt;Control, cost, and quality that is &lt;em&gt;good enough&lt;/em&gt; for the vast majority of document workflows. We focused on VLMs (rather than text-only OCR) because real documents are messy: handwritten forms, scanned PDFs, awkward multi-column layouts, and text baked into images all break naive text extraction.&lt;/p&gt;
&lt;p&gt;Running your own model turns a fixed price sheet into a set of knobs. The weights, serving engine, batch strategy, GPU type, and quantization are all yours to tune. You stop paying a per-token markup and start paying for GPU-seconds, which flips the economics once volume is non-trivial.&lt;/p&gt;
&lt;p&gt;There’s a strategic angle too, and it’s the one that’s easy to underweight until it bites you. If a hosted API is the foundation of your product, you don’t own the cost curve, the model lifecycle, or the deployment path. Prices move, default models change underneath you, and the exact model you tuned around can be deprecated on the vendor’s schedule, not yours. Self-hosting trades convenience for control over the variables that matter in production.&lt;/p&gt;
&lt;h3 id=&quot;the-three-axes-that-matter&quot;&gt;The three axes that matter&lt;/h3&gt;
&lt;p&gt;Everything below comes back to three measurements:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Throughput&lt;/strong&gt; — pages per second per GPU. Sets wall-clock time and how many GPUs you need.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Cost&lt;/strong&gt; — dollars per page. Sets whether the project is viable at scale.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Quality&lt;/strong&gt; — does the model actually read the page right? A page OCR’d cheaply is worthless if it silently garbles the text. We treat this qualitatively (more on why in &lt;a href=&quot;#model-selection&quot;&gt;Model Selection&lt;/a&gt;) and ship a viewer so you can check our work.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;We’ll take them roughly in the order they mattered to us: &lt;em&gt;which&lt;/em&gt; model, then &lt;em&gt;how fast and how cheap&lt;/em&gt;, then &lt;em&gt;how close to correct&lt;/em&gt;, and finally how all of that stacks up against the proprietary APIs.&lt;/p&gt;
&lt;h2 id=&quot;who-we-are&quot;&gt;Who we are&lt;/h2&gt;
&lt;p&gt;We’re the founders of &lt;a href=&quot;https://redspring.xyz&quot;&gt;Redspring&lt;/a&gt;, an AI and product-development consultancy. We build applied ML systems, including large-scale document and data processing, so this was equal parts client-relevant and an excuse to satisfy our own curiosity.&lt;/p&gt;
&lt;h2 id=&quot;overview&quot;&gt;Overview&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;&lt;a href=&quot;#model-selection&quot;&gt;Model and engine selection&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#modal-basics&quot;&gt;Modal basics&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#benchmark-methodology&quot;&gt;Benchmark methodology&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#results&quot;&gt;Results&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#what-we-learned&quot;&gt;What we learned&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;#conclusion&quot;&gt;Conclusion&lt;/a&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id=&quot;model-selection&quot;&gt;Model Selection&lt;/h2&gt;
&lt;p&gt;We picked models based largely on how they performed on OCR-oriented public benchmarks, including &lt;a href=&quot;https://huggingface.co/datasets/allenai/olmOCR-bench&quot;&gt;olmOCR-bench&lt;/a&gt;, &lt;a href=&quot;https://99franklin.github.io/ocrbench_v2/&quot;&gt;OCRBench v2&lt;/a&gt;, and the OCR-related leaderboards collected by LLM-Stats (&lt;a href=&quot;https://llm-stats.com/benchmarks/omnidocbench-1.5&quot;&gt;OmniDocBench 1.5&lt;/a&gt;, &lt;a href=&quot;https://llm-stats.com/benchmarks/ocrbench&quot;&gt;OCRBench&lt;/a&gt;). Because these benchmarks cover different tasks and not every model appears on every leaderboard, we treated them as directional rather than definitive.&lt;/p&gt;
&lt;p&gt;Here’s the rough ranking we used to decide what to test:&lt;/p&gt;



































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Composite rank&lt;/th&gt;&lt;th&gt;Model&lt;/th&gt;&lt;th&gt;Benchmark evidence used&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;&lt;code&gt;1&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;code&gt;datalab-to/chandra-ocr-2&lt;/code&gt;&lt;/td&gt;&lt;td&gt;olmOCR-bench&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;code&gt;2&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;code&gt;rednote-dots-ocr-community/dots.ocr-1.5&lt;/code&gt;&lt;/td&gt;&lt;td&gt;olmOCR-bench&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;code&gt;3&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;code&gt;NVIDIA Nemotron Nano V2 VL&lt;/code&gt;&lt;/td&gt;&lt;td&gt;OCRBench v2&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;code&gt;4&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;code&gt;Qwen3.5-122B-A10B&lt;/code&gt;&lt;/td&gt;&lt;td&gt;LLM-Stats OCRBench&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;code&gt;5&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;code&gt;Qwen3.5-35B-A3B&lt;/code&gt;&lt;/td&gt;&lt;td&gt;LLM-Stats OCRBench&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;This is not meant to be a universal ranking. It ignores cost, latency, serving complexity, and the specifics of the workload. It was just a sensible starting point.&lt;/p&gt;
&lt;h3 id=&quot;why-we-didnt-compute-a-single-accuracy-number&quot;&gt;Why we didn’t compute a single accuracy number&lt;/h3&gt;
&lt;p&gt;The obvious move here is to report CER/WER against ground truth and crown a winner. We deliberately didn’t, for two reasons. First, that work already exists and is done well — the public benchmarks above are the right place for a rigorous, apples-to-apples score, and we’d only be reproducing them worse. Second, a single aggregate number hides exactly the failures that matter in production. A model can post an excellent average CER while occasionally and &lt;em&gt;confidently&lt;/em&gt; rewriting a clause, and a quiet hallucination is far more dangerous downstream than uniformly fuzzy text.&lt;/p&gt;
&lt;p&gt;So we used the public benchmarks to &lt;em&gt;pick candidates&lt;/em&gt;, then judged the shortlist the way you’d actually judge it before shipping: by reading the output. We built a &lt;a href=&quot;#assessing-quality-the-side-by-side-viewer&quot;&gt;side-by-side viewer&lt;/a&gt; and spent real time in it. The verdict and per-model recommendations live there, next to the evidence. If quality is mission-critical for your workload, do the same — pull a representative sample, run it through your top 2–3 candidates, and have someone who knows the domain eyeball the results.&lt;/p&gt;
&lt;h2 id=&quot;modal-basics&quot;&gt;Modal basics&lt;/h2&gt;
&lt;p&gt;&lt;a href=&quot;https://modal.com&quot;&gt;Modal&lt;/a&gt; is a good fit for this work because it makes GPU-backed serving much less painful than doing the same thing on AWS/GCP/Azure from scratch. On a hyperscaler, serving a single GPU-backed model typically means: picking a GPU instance family with the right capacity and availability, wrestling CUDA drivers into the base image, wiring up autoscaling and cold-start handling, standing up a container registry, piping in secrets via IAM or a parameter store, and mounting some kind of network volume so you aren’t re-downloading 50+ GB of weights on every boot. Each of those is a half-day task in its own right, and most of them are about cloud plumbing rather than the model you actually want to serve.&lt;/p&gt;
&lt;p&gt;Modal collapses most of that into a single Python file. The basic pattern is simple:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;stage model weights into a persistent Volume&lt;/li&gt;
&lt;li&gt;mount that Volume into a GPU-backed container&lt;/li&gt;
&lt;li&gt;start the serving engine during container startup&lt;/li&gt;
&lt;li&gt;expose the service over a normal HTTP interface&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;For batch-style setup work, like downloading Hugging Face weights into durable storage, we used one-off functions. For serving, we used a GPU-backed class that launched the engine as a subprocess during startup and exposed a web server once the container was ready.&lt;/p&gt;
&lt;p&gt;Here’s the rough shape:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; style=&quot;background-color:#24292e;color:#e1e4e8; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;python&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;volume &lt;/span&gt;&lt;span style=&quot;color:#F97583&quot;&gt;=&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt; modal.Volume.from_name(&lt;/span&gt;&lt;span style=&quot;color:#9ECBFF&quot;&gt;&quot;dots-ocr-1.5-assets&quot;&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;, &lt;/span&gt;&lt;span style=&quot;color:#FFAB70&quot;&gt;create_if_missing&lt;/span&gt;&lt;span style=&quot;color:#F97583&quot;&gt;=&lt;/span&gt;&lt;span style=&quot;color:#79B8FF&quot;&gt;True&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#B392F0&quot;&gt;@app.function&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;(&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#FFAB70&quot;&gt;    image&lt;/span&gt;&lt;span style=&quot;color:#F97583&quot;&gt;=&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;image,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#FFAB70&quot;&gt;    secrets&lt;/span&gt;&lt;span style=&quot;color:#F97583&quot;&gt;=&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;[modal.Secret.from_name(&lt;/span&gt;&lt;span style=&quot;color:#9ECBFF&quot;&gt;&quot;hf-secret&quot;&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;)],&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#FFAB70&quot;&gt;    volumes&lt;/span&gt;&lt;span style=&quot;color:#F97583&quot;&gt;=&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;{&lt;/span&gt;&lt;span style=&quot;color:#9ECBFF&quot;&gt;&quot;/models&quot;&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;: volume},&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#F97583&quot;&gt;def&lt;/span&gt;&lt;span style=&quot;color:#B392F0&quot;&gt; upload_model_to_volume&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;() -&gt; &lt;/span&gt;&lt;span style=&quot;color:#79B8FF&quot;&gt;None&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;    snapshot_download(&lt;/span&gt;&lt;span style=&quot;color:#79B8FF&quot;&gt;...&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;, &lt;/span&gt;&lt;span style=&quot;color:#FFAB70&quot;&gt;local_dir&lt;/span&gt;&lt;span style=&quot;color:#F97583&quot;&gt;=&lt;/span&gt;&lt;span style=&quot;color:#9ECBFF&quot;&gt;&quot;/models/dots.ocr-1.5&quot;&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;    volume.commit()&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;And for a long-lived service:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; style=&quot;background-color:#24292e;color:#e1e4e8; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;python&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#B392F0&quot;&gt;@app.cls&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;(&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#FFAB70&quot;&gt;    image&lt;/span&gt;&lt;span style=&quot;color:#F97583&quot;&gt;=&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;vllm_image,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#FFAB70&quot;&gt;    gpu&lt;/span&gt;&lt;span style=&quot;color:#F97583&quot;&gt;=&lt;/span&gt;&lt;span style=&quot;color:#9ECBFF&quot;&gt;&quot;L4&quot;&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#FFAB70&quot;&gt;    volumes&lt;/span&gt;&lt;span style=&quot;color:#F97583&quot;&gt;=&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;{&lt;/span&gt;&lt;span style=&quot;color:#9ECBFF&quot;&gt;&quot;/models&quot;&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;: model_volume, &lt;/span&gt;&lt;span style=&quot;color:#9ECBFF&quot;&gt;&quot;/root/.cache/vllm&quot;&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;: vllm_cache_volume},&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#FFAB70&quot;&gt;    scaledown_window&lt;/span&gt;&lt;span style=&quot;color:#F97583&quot;&gt;=&lt;/span&gt;&lt;span style=&quot;color:#79B8FF&quot;&gt;300&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#FFAB70&quot;&gt;    min_containers&lt;/span&gt;&lt;span style=&quot;color:#F97583&quot;&gt;=&lt;/span&gt;&lt;span style=&quot;color:#79B8FF&quot;&gt;0&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#F97583&quot;&gt;class&lt;/span&gt;&lt;span style=&quot;color:#B392F0&quot;&gt; DotsOcrVLLM&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#B392F0&quot;&gt;    @modal.enter&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;()&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#F97583&quot;&gt;    def&lt;/span&gt;&lt;span style=&quot;color:#B392F0&quot;&gt; startup&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;(self) -&gt; &lt;/span&gt;&lt;span style=&quot;color:#79B8FF&quot;&gt;None&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#79B8FF&quot;&gt;        self&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;.process &lt;/span&gt;&lt;span style=&quot;color:#F97583&quot;&gt;=&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt; subprocess.Popen([&lt;/span&gt;&lt;span style=&quot;color:#9ECBFF&quot;&gt;&quot;vllm&quot;&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;, &lt;/span&gt;&lt;span style=&quot;color:#9ECBFF&quot;&gt;&quot;serve&quot;&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;, &lt;/span&gt;&lt;span style=&quot;color:#79B8FF&quot;&gt;...&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;])&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#B392F0&quot;&gt;    @modal.web_server&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;(&lt;/span&gt;&lt;span style=&quot;color:#79B8FF&quot;&gt;8000&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;, &lt;/span&gt;&lt;span style=&quot;color:#FFAB70&quot;&gt;startup_timeout&lt;/span&gt;&lt;span style=&quot;color:#F97583&quot;&gt;=&lt;/span&gt;&lt;span style=&quot;color:#79B8FF&quot;&gt;900&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;, &lt;/span&gt;&lt;span style=&quot;color:#FFAB70&quot;&gt;label&lt;/span&gt;&lt;span style=&quot;color:#F97583&quot;&gt;=&lt;/span&gt;&lt;span style=&quot;color:#9ECBFF&quot;&gt;&quot;dots-ocr-1-5-vllm&quot;&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#F97583&quot;&gt;    def&lt;/span&gt;&lt;span style=&quot;color:#B392F0&quot;&gt; serve&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;(self) -&gt; &lt;/span&gt;&lt;span style=&quot;color:#79B8FF&quot;&gt;None&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;:&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#F97583&quot;&gt;        pass&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;A few details mattered. Each of these corresponds to a platform primitive that, if it didn’t exist, you’d have to build yourself on a generic cloud:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Volumes.&lt;/strong&gt; A Modal Volume is a persistent, mountable filesystem shared across functions and containers. Model weights for a VLM run 20–200+ GB; downloading them on every container start would blow the cold-start budget and rack up egress from model stores like Hugging Face. Staging weights into a Volume once and mounting it read-only at runtime means each new container sees the weights as local disk in a fraction of a second.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Secrets.&lt;/strong&gt; Modal Secrets are named, server-side key-value bundles you reference by name and inject as environment variables at runtime. We used them for things like &lt;code&gt;HF_TOKEN&lt;/code&gt; so credentials never landed in the image, the repo, or local shell history. Rotating a secret is a single CLI call; no rebuild required.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scaling controls.&lt;/strong&gt; &lt;code&gt;min_containers&lt;/code&gt;, &lt;code&gt;max_containers&lt;/code&gt;, and &lt;code&gt;scaledown_window&lt;/code&gt; let you trade idle cost against cold starts — a warm container eats GPU-time until it scales down, a cold one pays the model-load penalty on the next request. For benchmarking we set &lt;code&gt;max_containers=1&lt;/code&gt; so the system would stay on a single container rather than auto-scaling to absorb our synthetic load, which would have made throughput measurements unreadable. We also usually killed containers right after a run to avoid paying for idle time. In production, the same two knobs let you pre-warm for SLA-sensitive traffic or go fully scale-to-zero for batch windows.&lt;/p&gt;
&lt;p&gt;That’s enough infrastructure to make the experiments reproducible without turning this post into a deployment manual.&lt;/p&gt;
&lt;figure class=&quot;not-prose my-6&quot;&gt;
  &lt;img src=&quot;https://redspring.xyz/modal-screenshot.png&quot; alt=&quot;Modal Apps dashboard showing the deployed OCR inference endpoints — chandra-ocr-2, nemotron-fp8, nemotron-vl, glm-ocr, and dots-ocr-1.5 — each with its GPU type and web endpoint.&quot; class=&quot;w-full rounded-lg border border-gray-200&quot; loading=&quot;lazy&quot;&gt;
  &lt;figcaption class=&quot;mt-2 text-center text-sm text-gray-500&quot;&gt;Each model became its own Modal app exposed over a web endpoint and pinned to the GPU type we were sweeping.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id=&quot;benchmark-methodology&quot;&gt;Benchmark Methodology&lt;/h2&gt;
&lt;p&gt;The basic idea was straightforward: load a corpus of PDFs, send one OCR request per page until the sample run is complete, and measure throughput under different concurrency levels. Combined with measured pages-per-second-per-GPU and Modal’s on-demand pricing, that gave us a reasonable way to estimate cost and fleet size for a 100,000-page run before we accidentally emptied our bank accounts.&lt;/p&gt;
&lt;p&gt;Concretely, we:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;ran the same models across multiple GPU types&lt;/li&gt;
&lt;li&gt;swept through concurrency levels to determine max GPU throughput&lt;/li&gt;
&lt;li&gt;evaluated the results quantitatively (speed, cost) and qualitatively (output fidelity)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;One important setup detail: we pre-rasterized the PDF pages into images as a cheap one-off job before the main OCR benchmarks. That kept the benchmark focused on the part we actually cared about — model serving throughput and GPU cost — rather than letting PDF rendering dominate or add noise to the measurements. In a real production pipeline, you would still need to account for rendering, storage, and orchestration, but those costs are much smaller and easier to optimize than GPU-backed model inference.&lt;/p&gt;
&lt;p&gt;We did &lt;strong&gt;not&lt;/strong&gt; run &lt;em&gt;every&lt;/em&gt; model/configuration as a literal 100,000-page end-to-end job. Instead, we used representative benchmark samples and combined the measured throughput with public pricing to estimate full-run cost. For planning and comparison, that was good enough. As we got closer to the final run, we moved from single-container spot checks to multi-container &lt;code&gt;H100&lt;/code&gt; runs of 5,000–20,000 pages so the extrapolation would reflect production fleet behavior rather than individual GPU performance. After those multi-container H100 runs gave us a configuration we trusted, we ran the full 100,000-page job with that setup.&lt;/p&gt;
&lt;h2 id=&quot;results&quot;&gt;Results&lt;/h2&gt;
&lt;p&gt;The rest of this section walks through these three axes in order: &lt;strong&gt;quality&lt;/strong&gt; (is the output any good?), &lt;strong&gt;throughput and cost&lt;/strong&gt; (can we afford it at scale?), and &lt;strong&gt;proprietary-API comparison&lt;/strong&gt; (is it worth self-hosting at all?).&lt;/p&gt;
&lt;h3 id=&quot;assessing-quality-the-side-by-side-viewer&quot;&gt;Assessing quality: the side-by-side viewer&lt;/h3&gt;
&lt;p&gt;Before any of the cost numbers mean anything, you have to answer the obvious question: does this model actually read the page right? So we built a small static viewer that lets you flip through individual pages and compare every model’s OCR output against the rendered original, side-by-side. Best on desktop.&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://redspringxyz.github.io/ocr-results-viewer/&quot;&gt;Open the OCR Results Viewer in a new tab →&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Our verdict after spending real time in the viewer:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Reach for Chandra when fidelity is paramount.&lt;/strong&gt; It’s our top pick on accuracy, and the lead is most visible on tables and complex formatting where the others start dropping cells or mangling structure. Good fit for compliance-sensitive documents, archive digitization, or anything downstream that assumes the text is trustworthy.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Reach for &lt;code&gt;dots.ocr-1.5&lt;/code&gt; when throughput or cost is the binding constraint.&lt;/strong&gt; It’s a very close second on quality and a runaway winner on cost-per-page — the right call when the downstream consumer is another model (RAG indexing, classification, summarization, agentic workflows) that can absorb a little OCR noise.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Both are genuinely strong; the pick is about which axis you’re optimizing. And don’t sleep on the Qwen variants in the viewer — they weren’t purpose-built for OCR and still hold up better than we expected. The viewer exists precisely so you don’t have to take our verdict, the public leaderboards, or the cost tables below on faith.&lt;/p&gt;
&lt;h3 id=&quot;self-hosted-cost-to-process-100k-pages&quot;&gt;Self-hosted cost to process 100k pages&lt;/h3&gt;
&lt;p&gt;The headline run used a fleet of &lt;strong&gt;60 H100s&lt;/strong&gt; handling 200 concurrent requests. It submitted 100,056 pages, completed 98,159, and sustained 28.96 submitted pages/s — finishing in &lt;strong&gt;56.5 minutes&lt;/strong&gt; at &lt;strong&gt;$223.10&lt;/strong&gt; of measured GPU cost. Normalized over successful pages, that’s &lt;strong&gt;$2.27 per 1,000 pages&lt;/strong&gt;, at a failure rate just under 2%.&lt;/p&gt;
&lt;p&gt;We didn’t jump straight to 60 GPUs. Before committing to the full job, we ran a few smaller multi-container &lt;code&gt;H100&lt;/code&gt; sweeps with &lt;code&gt;dots.ocr-1.5&lt;/code&gt; to establish a throughput-and-cost baseline we trusted enough to extrapolate from:&lt;/p&gt;
&lt;div class=&quot;not-prose my-6 overflow-x-auto rounded-lg border border-gray-200&quot;&gt;
  &lt;table class=&quot;min-w-[720px] w-full text-left text-sm&quot;&gt;
    &lt;thead class=&quot;bg-gray-50 text-xs font-semibold uppercase tracking-wide text-gray-500&quot;&gt;
      &lt;tr&gt;
        &lt;th class=&quot;px-4 py-3&quot;&gt;Profile&lt;/th&gt;
        &lt;th class=&quot;px-4 py-3&quot;&gt;Sample&lt;/th&gt;
        &lt;th class=&quot;px-4 py-3&quot;&gt;Throughput&lt;/th&gt;
        &lt;th class=&quot;px-4 py-3&quot;&gt;Observed cost&lt;/th&gt;
        &lt;th class=&quot;px-4 py-3&quot;&gt;100K projection&lt;/th&gt;
      &lt;/tr&gt;
    &lt;/thead&gt;
    &lt;tbody class=&quot;divide-y divide-gray-200 bg-white text-gray-700&quot;&gt;
      &lt;tr&gt;
        &lt;td class=&quot;px-4 py-3 font-medium text-gray-900&quot;&gt;40×H100, c=160&lt;/td&gt;
        &lt;td class=&quot;px-4 py-3&quot;&gt;5,000 submitted, 4,911 ok&lt;/td&gt;
        &lt;td class=&quot;px-4 py-3&quot;&gt;18.95 pages/s&lt;/td&gt;
        &lt;td class=&quot;px-4 py-3&quot;&gt;$11.37 total&lt;/td&gt;
        &lt;td class=&quot;px-4 py-3&quot;&gt;~88–90m, ~$227–232&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
        &lt;td class=&quot;px-4 py-3 font-medium text-gray-900&quot;&gt;40×H100, c=120&lt;/td&gt;
        &lt;td class=&quot;px-4 py-3&quot;&gt;5,000 submitted, 4,980 ok&lt;/td&gt;
        &lt;td class=&quot;px-4 py-3&quot;&gt;21.77 pages/s&lt;/td&gt;
        &lt;td class=&quot;px-4 py-3&quot;&gt;$10.04 total&lt;/td&gt;
        &lt;td class=&quot;px-4 py-3&quot;&gt;~77m, ~$202&lt;/td&gt;
      &lt;/tr&gt;
    &lt;/tbody&gt;
  &lt;/table&gt;
&lt;/div&gt;
&lt;p&gt;The next table is a compact view of model x GPU sweeps. It does not list every concurrency setting we tested; instead, it normalizes each run down to the best per-GPU throughput and the estimated &lt;code&gt;100K&lt;/code&gt;-page cost that mattered for choosing the final profile.&lt;/p&gt;
&lt;div class=&quot;not-prose my-6 overflow-x-auto rounded-lg border border-gray-200&quot;&gt;
  &lt;table class=&quot;min-w-[760px] w-full text-left text-sm&quot;&gt;
    &lt;thead class=&quot;bg-gray-50 text-xs font-semibold uppercase tracking-wide text-gray-500&quot;&gt;
      &lt;tr&gt;
        &lt;th class=&quot;px-4 py-3&quot;&gt;Model / profile&lt;/th&gt;
        &lt;th class=&quot;px-4 py-3&quot;&gt;Throughput&lt;/th&gt;
        &lt;th class=&quot;px-4 py-3&quot;&gt;100K cost&lt;/th&gt;
        &lt;th class=&quot;px-4 py-3&quot;&gt;Notes&lt;/th&gt;
      &lt;/tr&gt;
    &lt;/thead&gt;
    &lt;tbody class=&quot;divide-y divide-gray-200 bg-white text-gray-700&quot;&gt;
      &lt;tr&gt;
        &lt;td class=&quot;px-4 py-3 font-medium text-gray-900&quot;&gt;Dots OCR 1.5 on H100&lt;/td&gt;
        &lt;td class=&quot;px-4 py-3&quot;&gt;0.39 pages/s/GPU&lt;/td&gt;
        &lt;td class=&quot;px-4 py-3&quot;&gt;~$281&lt;/td&gt;
        &lt;td class=&quot;px-4 py-3&quot;&gt;Best single-GPU Dots profile&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
        &lt;td class=&quot;px-4 py-3 font-medium text-gray-900&quot;&gt;Dots OCR 1.5 on A100&lt;/td&gt;
        &lt;td class=&quot;px-4 py-3&quot;&gt;0.19 pages/s/GPU&lt;/td&gt;
        &lt;td class=&quot;px-4 py-3&quot;&gt;~$312&lt;/td&gt;
        &lt;td class=&quot;px-4 py-3&quot;&gt;Cheaper GPU, about half H100 throughput&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
        &lt;td class=&quot;px-4 py-3 font-medium text-gray-900&quot;&gt;Dots OCR 1.5 on L4&lt;/td&gt;
        &lt;td class=&quot;px-4 py-3&quot;&gt;0.07 pages/s/GPU&lt;/td&gt;
        &lt;td class=&quot;px-4 py-3&quot;&gt;~$325&lt;/td&gt;
        &lt;td class=&quot;px-4 py-3&quot;&gt;Similar cost band, ~5–6× slower&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
        &lt;td class=&quot;px-4 py-3 font-medium text-gray-900&quot;&gt;Chandra on A100&lt;/td&gt;
        &lt;td class=&quot;px-4 py-3&quot;&gt;0.06 pages/s/GPU&lt;/td&gt;
        &lt;td class=&quot;px-4 py-3&quot;&gt;~$935&lt;/td&gt;
        &lt;td class=&quot;px-4 py-3&quot;&gt;Best quality, much slower and costlier&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
        &lt;td class=&quot;px-4 py-3 font-medium text-gray-900&quot;&gt;Nemotron-FP8 on H100&lt;/td&gt;
        &lt;td class=&quot;px-4 py-3&quot;&gt;0.17 pages/s/GPU&lt;/td&gt;
        &lt;td class=&quot;px-4 py-3&quot;&gt;~$646&lt;/td&gt;
        &lt;td class=&quot;px-4 py-3&quot;&gt;Quantization helped, still behind Dots&lt;/td&gt;
      &lt;/tr&gt;
      &lt;tr&gt;
        &lt;td class=&quot;px-4 py-3 font-medium text-gray-900&quot;&gt;Qwen3.5-35B-FP8 on H100&lt;/td&gt;
        &lt;td class=&quot;px-4 py-3&quot;&gt;0.05 pages/s/GPU&lt;/td&gt;
        &lt;td class=&quot;px-4 py-3&quot;&gt;~$2,032&lt;/td&gt;
        &lt;td class=&quot;px-4 py-3&quot;&gt;Not viable for this OCR workload&lt;/td&gt;
      &lt;/tr&gt;
    &lt;/tbody&gt;
  &lt;/table&gt;
&lt;/div&gt;
&lt;p&gt;These estimates used &lt;strong&gt;Modal’s pricing&lt;/strong&gt; at the time of publishing and the formula:&lt;/p&gt;
&lt;div class=&quot;bg-gray-50 p-4 rounded-md border border-gray-200 my-4&quot;&gt;
  &lt;strong&gt;Estimated run cost formula:&lt;/strong&gt;
  &lt;div class=&quot;font-mono text-sm mt-2 px-2 py-1 bg-white rounded&quot;&gt;
    run cost ≈ &lt;span class=&quot;font-semibold&quot;&gt;(pages / throughput)&lt;/span&gt; × &lt;span class=&quot;font-semibold&quot;&gt;GPU_count&lt;/span&gt; × &lt;span class=&quot;font-semibold&quot;&gt;GPU_price_per_sec&lt;/span&gt;
  &lt;/div&gt;
  &lt;div class=&quot;text-xs text-gray-600 mt-1&quot;&gt;
    &lt;em&gt;Where &quot;throughput&quot; is measured in pages/sec/GPU.&lt;/em&gt;
  &lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;Pricing assumptions used here:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;L4 = $0.000222/s&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;L40S = $0.000542/s&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;A100&lt;/code&gt; (40GB) &lt;code&gt;= $0.000583/s&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;A100-80GB = $0.000694/s&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;H100 = $0.001097/s&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;H200 = $0.001261/s&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;B200 = $0.001736/s&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;proprietary-api-price-sheet-comparison&quot;&gt;Proprietary API price-sheet comparison&lt;/h3&gt;
&lt;p&gt;We include proprietary API pricing as context, not as an argument for the APIs. The thing worth scrutinizing isn’t the per-page number; it’s the dependency you take on when one vendor owns both the model and the price sheet.&lt;/p&gt;
&lt;p&gt;The shape of the comparison is a vise. The frontier models that match or beat our open-source picks on quality — Claude Opus 4.8, GPT-5.5 — cost roughly $4,000 and $6,000 for the same 100k-page run, 18–27× our self-hosted cost. The proprietary models that get close to our cost, like &lt;code&gt;gpt-5.4-nano&lt;/code&gt; and &lt;code&gt;gpt-4.1-mini&lt;/code&gt;, aren’t quality-competitive: in our review they were the ones most prone to quietly rewriting text. The example below is representative. The source sentence says researchers run registration and fusion &lt;em&gt;concurrently&lt;/em&gt;; &lt;code&gt;gpt-5.4-nano&lt;/code&gt; invented a different claim entirely, lifting the phrase “dynamic gradient sparsity property” from the &lt;em&gt;next&lt;/em&gt; sentence and presenting it as the method. Nothing in the output signals that it’s wrong — which is exactly what makes it dangerous.&lt;/p&gt;
&lt;div class=&quot;not-prose my-6 space-y-3&quot;&gt;
  &lt;div class=&quot;rounded-lg border border-green-200 bg-green-50 p-4&quot;&gt;
    &lt;div class=&quot;mb-2 flex items-center gap-2 text-xs font-semibold uppercase tracking-wide text-green-700&quot;&gt;
      &lt;span&gt;✓ Correct&lt;/span&gt;
      &lt;span class=&quot;font-normal normal-case text-green-600&quot;&gt;(matched by &lt;code&gt;chandra-ocr-2&lt;/code&gt; and &lt;code&gt;dots.ocr-1.5&lt;/code&gt;)&lt;/span&gt;
    &lt;/div&gt;
    &lt;p class=&quot;text-sm text-gray-800&quot;&gt;
      To address the complexities of image registration task, certain researchers endeavor to &lt;span class=&quot;rounded bg-green-200 px-1 font-medium text-green-900&quot;&gt;execute registration and fusion processes concurrently&lt;/span&gt;.
    &lt;/p&gt;
  &lt;/div&gt;
  &lt;div class=&quot;rounded-lg border border-red-200 bg-red-50 p-4&quot;&gt;
    &lt;div class=&quot;mb-2 flex items-center gap-2 text-xs font-semibold uppercase tracking-wide text-red-700&quot;&gt;
      &lt;span&gt;✗ Hallucinated&lt;/span&gt;
      &lt;span class=&quot;font-normal normal-case text-red-600&quot;&gt;(&lt;code&gt;gpt-5.4-nano&lt;/code&gt;)&lt;/span&gt;
    &lt;/div&gt;
    &lt;p class=&quot;text-sm text-gray-800&quot;&gt;
      To address the complexities of image registration task, certain researchers endeavor to &lt;span class=&quot;rounded bg-red-200 px-1 font-medium text-red-900&quot;&gt;implement a dynamic gradient sparsity property&lt;/span&gt;.
    &lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;To sanity-check the earlier token-budget estimates, we also ran a sample of our own pages through several proprietary models and recorded the provider-reported costs. Extrapolating from those observed costs, this is what the same 100,000-page workload would look like.&lt;/p&gt;






































































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Model&lt;/th&gt;&lt;th&gt;Observed cost/page from sample&lt;/th&gt;&lt;th&gt;Extrapolated &lt;code&gt;100K&lt;/code&gt; cost&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;OpenAI &lt;code&gt;gpt-5.4-nano&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;code&gt;$0.0012&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;code&gt;~$120&lt;/code&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;OpenAI &lt;code&gt;gpt-4.1-mini&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;code&gt;$0.0020&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;code&gt;~$200&lt;/code&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;OpenAI &lt;code&gt;gpt-5.4-mini&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;code&gt;$0.0046&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;code&gt;~$460&lt;/code&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Anthropic &lt;code&gt;Claude Haiku 4.5&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;code&gt;$0.0051&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;code&gt;~$510&lt;/code&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;OpenAI &lt;code&gt;gpt-4.1&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;code&gt;$0.0069&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;code&gt;~$690&lt;/code&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;OpenAI &lt;code&gt;gpt-5.4&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;code&gt;$0.02&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;code&gt;~$2,000&lt;/code&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Anthropic &lt;code&gt;Claude Sonnet 4.6&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;code&gt;$0.02&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;code&gt;~$2,000&lt;/code&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Anthropic &lt;code&gt;Claude Opus 4.6&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;code&gt;$0.03&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;code&gt;~$3,000&lt;/code&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Anthropic &lt;code&gt;Claude Opus 4.8&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;code&gt;$0.04&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;code&gt;~$4,000&lt;/code&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;OpenAI &lt;code&gt;gpt-5.5&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;code&gt;$0.06&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;code&gt;~$6,000&lt;/code&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;OpenAI &lt;code&gt;gpt-5.4-pro&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;code&gt;$0.65&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;code&gt;~$65,000&lt;/code&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;OpenAI &lt;code&gt;gpt-5.5-pro&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;code&gt;$0.94&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;code&gt;~$94,000&lt;/code&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;Two things to take away from the table. First, read it by quality tier, not by price: the cheap rows ($120–$510) are &lt;em&gt;not&lt;/em&gt; substitutes for &lt;code&gt;dots.ocr-1.5&lt;/code&gt; or Chandra, because they’re the ones that hallucinate. Second, once you filter to genuinely comparable quality, the proprietary cost sits an order of magnitude or more above our $223 self-hosted run. Batch APIs can narrow that gap for jobs you’re willing to defer, but a batch endpoint is a different product the moment you need predictable sub-hour completion.&lt;/p&gt;
&lt;h3 id=&quot;what-the-api-comparison-misses&quot;&gt;What the API comparison misses&lt;/h3&gt;
&lt;p&gt;The broad takeaway is simpler than a head-to-head matrix: self-hosted VLMs on Modal are often competitive on run cost for comparable-quality OCR, and the dollar figure is the least of it. With the propietary APIs, the cheaper model you built around can get more expensive or retired out from under you. Self-hosting trades that exposure for stronger guarantees and choices &lt;em&gt;you&lt;/em&gt; can make around:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;The deployment knobs.&lt;/strong&gt; You choose the model, GPU, engine, concurrency, batching strategy, and fleet size. With an API, the highest-leverage knobs belong to the vendor.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The deadline trade-off.&lt;/strong&gt; Finishing faster is a visible capacity decision: add GPUs, pick a different GPU class, tune the server. With APIs, sub-hour throughput depends on account tiers, rate limits, and provider capacity.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Where the data goes.&lt;/strong&gt; For OCR workloads, the input is often the sensitive part. Running in your own Modal account changes the privacy and compliance conversation.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;If OCR is becoming core infrastructure for your product or business, owning the serving path matters.&lt;/p&gt;
&lt;h2 id=&quot;what-we-learned&quot;&gt;What we learned&lt;/h2&gt;
&lt;h3 id=&quot;1-benchmark-your-workload-not-somebody-elses&quot;&gt;1. Benchmark your workload, not somebody else’s&lt;/h3&gt;
&lt;p&gt;It’s tempting to read a leaderboard, take the top row, and move on. Sometimes that’s right. Sometimes it’s exactly wrong — two of our public-ranking favorites lost once we measured them on our own pages.&lt;/p&gt;
&lt;p&gt;What actually decides the winner is the shape of &lt;em&gt;your&lt;/em&gt; workload: document type, token mix, output length, preprocessing path, concurrency, and the precise quality threshold you need. OCR is a particularly nasty case because it pairs heavy image inputs with long text outputs, so a model that’s fast on chat-style traffic can fall over here. The leaderboard tells you who to test; your own pages tell you who to ship.&lt;/p&gt;
&lt;h3 id=&quot;2-gpu-price-and-sla-have-to-be-evaluated-together&quot;&gt;2. GPU price and SLA have to be evaluated together&lt;/h3&gt;
&lt;p&gt;On paper, an &lt;code&gt;L4&lt;/code&gt; or &lt;code&gt;L40S&lt;/code&gt; is far cheaper per second than an &lt;code&gt;H100&lt;/code&gt;. In practice that’s a trap, because per-second price and time-to-finish trade off against each other. Our measured &lt;code&gt;dots.ocr-1.5&lt;/code&gt; rows make it concrete: the &lt;code&gt;L4&lt;/code&gt; landed in roughly the same &lt;em&gt;cost-per-page&lt;/em&gt; band as the &lt;code&gt;H100&lt;/code&gt; while running ~5–6× slower. So the cheap GPU buys you nothing on cost and costs you a one-to-two-hour job turning into an overnight one. The right pick is a function of your deadline: a loose SLA makes a low-end GPU perfectly sensible, a tight one justifies the H100 even at a higher sticker price. We chose H100s for the final run on that basis — best cost/time trade-off, not cheapest line item.&lt;/p&gt;
&lt;h3 id=&quot;3-quantization-is-worth-taking-seriously&quot;&gt;3. Quantization is worth taking seriously&lt;/h3&gt;
&lt;p&gt;We started out mildly skeptical that FP8 would hold quality without leaving performance on the table. The data didn’t support the skepticism. The clearest case from the later sweeps is &lt;code&gt;Nemotron&lt;/code&gt;, holding everything else constant:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;BF16&lt;/strong&gt; on &lt;code&gt;H100&lt;/code&gt;: ~0.10 pages/s, ~$1,100 per 100k pages.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;FP8&lt;/strong&gt; on &lt;code&gt;H100&lt;/code&gt;: ~0.17 pages/s, ~$650 per 100k pages.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;That’s a ~70% throughput gain and a ~40% cost cut from a single knob. It wasn’t enough to make Nemotron beat &lt;code&gt;dots.ocr-1.5&lt;/code&gt; for this workload, but that’s the narrow point. The broad one is that quantization is a real, near-free lever on both throughput and cost, and it belongs in the production search space rather than the “maybe later” pile.&lt;/p&gt;
&lt;h2 id=&quot;conclusion&quot;&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;We went in curious about two things: how good open-source OCR VLMs have actually gotten, and how much effort it really takes to run them at scale. The answers turned out to be “very” and “less than we expected.” 100,000 pages, 56 minutes, $223, on infrastructure we stood up in an afternoon.&lt;/p&gt;
&lt;p&gt;Hosted APIs are the right default for prototypes and one-off product work, and we’ll keep using them there. But “default” isn’t the same as “neutral.” The moment you care about throughput, cost control, batch windows, data residency, or tuning the system around a specific workload, the calculus shifts toward owning the serving path. Self-hosting isn’t always the cheapest line on a price sheet, but for comparable-quality OCR at volume it was both cheaper &lt;em&gt;and&lt;/em&gt; more controllable — and the operational know-how you build along the way is itself a moat that compounds.&lt;/p&gt;
&lt;p&gt;That’s the part we’d underline for anyone weighing this: owning the model layer gives you direct control over the variables that decide whether something works in production — performance, cost, quality, and predictability — instead of renting them from a vendor whose incentives aren’t yours.&lt;/p&gt;
&lt;p&gt;And we don’t think $223 is the floor. Newer serving engines, more aggressive quantization, and more careful tuning all point downward from here. We’ll keep tracking the state of the art; our guess is there’s still meaningful headroom left.&lt;/p&gt;
&lt;p&gt;Useful references:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://modal.com/llm-almanac/how-to-benchmark&quot;&gt;Modal’s benchmarking guide&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://github.com/vllm-project/guidellm&quot;&gt;guidellm&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://modal.com/pricing&quot;&gt;Modal’s Pricing page&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;</content:encoded></item><item><title>We Benchmarked Frontier Reasoning Models on The Atlantic&apos;s Bracket City</title><link>https://redspring.xyz/blog/bracket-city-bench/</link><guid isPermaLink="true">https://redspring.xyz/blog/bracket-city-bench/</guid><description>A deep dive into how frontier AI models perform on complex word puzzles, revealing surprising insights about reasoning efficiency vs. accuracy.</description><pubDate>Tue, 15 Jul 2025 12:00:00 GMT</pubDate><content:encoded>&lt;p&gt;I’ve been obsessed with The Atlantic’s new word game, Bracket City. It’s a puzzle where clues hide inside nested brackets. You start with something like &lt;code&gt;[[&quot;___ of Arabia&quot;] who starred in &quot;[monosyllabic Michael [capital of Mississippi] album] Boys&quot;]&lt;/code&gt; and work your way inward—first solving ”___ of Arabia” (Lawrence) and “capital of Mississippi” (Jackson), revealing &lt;code&gt;[Lawrence who starred in [monosyllabic Michael Jackson album] Boys]&lt;/code&gt;. Each solved clue unlocks new brackets until you reveal a final statement.&lt;/p&gt;
&lt;p&gt;I’m ashamed to admit that when I get stuck, I’ll paste a screenshot into ChatGPT o3 and ask for help. It’s freakishly good. At one point, I noticed in the chain of thought that o3 often tries to reason through the entire puzzle. So I got curious: I started pasting screenshots of completely unsolved puzzles and asking o3 to work through the whole thing. Often, it would nail the solution perfectly.&lt;/p&gt;
&lt;p&gt;That’s when it hit me: if these models can solve complex word puzzles, how do they actually compare? Which ones excel at this kind of deep, recursive reasoning?&lt;/p&gt;
&lt;video width=&quot;100%&quot; controls class=&quot;mb-0&quot;&gt;
  &lt;source src=&quot;https://redspring.xyz/bracket-city.mp4&quot; type=&quot;video/mp4&quot;&gt;
  Your browser does not support the video tag.
&lt;/video&gt;
&lt;p class=&quot;text-gray-700 text-sm&quot;&gt;Claude 4 Opus solving a recent Bracket City puzzle&lt;/p&gt;
&lt;h2 id=&quot;building-the-benchmark&quot;&gt;Building the Benchmark&lt;/h2&gt;
&lt;p&gt;My first approach seemed obvious: feed screenshots to each model and see who wins.&lt;/p&gt;
&lt;p&gt;It worked okay, but the inference endpoints timed out constantly. Models would get halfway through reasoning and just… stop. Even when they didn’t timeout, the visual parsing was inconsistent—some models would misread brackets or lose track of nested structures entirely.&lt;/p&gt;
&lt;p&gt;The breakthrough came when I reimplemented Bracket City’s game logic as LLM tool calls:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; style=&quot;background-color:#24292e;color:#e1e4e8; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;js&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#F97583&quot;&gt;const&lt;/span&gt;&lt;span style=&quot;color:#79B8FF&quot;&gt; tools&lt;/span&gt;&lt;span style=&quot;color:#F97583&quot;&gt; =&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt; {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;  makeGuess: {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;    description: &lt;/span&gt;&lt;span style=&quot;color:#9ECBFF&quot;&gt;&quot;Make a guess for a specific bracket clue&quot;&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;    parameters: {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;      clue: &lt;/span&gt;&lt;span style=&quot;color:#9ECBFF&quot;&gt;&quot;The clue text inside brackets (without the brackets)&quot;&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;      guess: &lt;/span&gt;&lt;span style=&quot;color:#9ECBFF&quot;&gt;&quot;Your guess for the answer to this clue&quot;&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;    },&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;  },&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;  getHint: {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;    description: &lt;/span&gt;&lt;span style=&quot;color:#9ECBFF&quot;&gt;&quot;Get a hint (first letter) for a difficult clue&quot;&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;    parameters: {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;      clue: &lt;/span&gt;&lt;span style=&quot;color:#9ECBFF&quot;&gt;&quot;The clue text inside brackets (without the brackets)&quot;&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;    },&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;  },&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;  revealClue: {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;    description: &lt;/span&gt;&lt;span style=&quot;color:#9ECBFF&quot;&gt;&quot;Reveal the full answer for a clue (last resort)&quot;&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;    parameters: {&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;      clue: &lt;/span&gt;&lt;span style=&quot;color:#9ECBFF&quot;&gt;&quot;The clue text inside brackets (without the brackets)&quot;&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;    },&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;  },&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;};&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;This approach worked beautifully. Models could now interact with puzzles programmatically, maintaining state while focusing purely on the reasoning challenge. I gave each model up to 50 steps to solve a puzzle, with the same scoring rules as human players: start at 100 points, lose 2 for wrong guesses, 5 for hints, and 15 for reveals.&lt;/p&gt;
&lt;p&gt;The system prompt emphasized working from the innermost brackets outward—a key strategy for solving these puzzles efficiently. Models that understood this recursive pattern performed significantly better.&lt;/p&gt;
&lt;h2 id=&quot;the-results&quot;&gt;The Results&lt;/h2&gt;
&lt;p&gt;I tested 16 frontier models across 20 different Bracket City puzzles. The results revealed a fascinating tradeoff between accuracy and efficiency:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The Top Scorer: o3-high&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Average score: 92.11/100&lt;/li&gt;
&lt;li&gt;Success rate: 100%&lt;/li&gt;
&lt;li&gt;Average time per puzzle: &lt;strong&gt;11 minutes&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;The Efficiency Champion: Claude 4 Opus&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Average score: 88.9/100&lt;/li&gt;
&lt;li&gt;Success rate: 100%&lt;/li&gt;
&lt;li&gt;Average time per puzzle: &lt;strong&gt;3 minutes&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;While o3-high technically won with the highest score, it came at a steep cost—taking nearly 4x longer than Claude 4 Opus to achieve only marginally better results. This raises a crucial question: is a 3.2% improvement in accuracy worth quadrupling your inference time?&lt;/p&gt;
&lt;p&gt;Here’s the complete leaderboard:&lt;/p&gt;




























































































































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Rank&lt;/th&gt;&lt;th&gt;Model&lt;/th&gt;&lt;th&gt;Average Score&lt;/th&gt;&lt;th&gt;Success Rate&lt;/th&gt;&lt;th&gt;Avg Time (seconds)&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;1&lt;/td&gt;&lt;td&gt;o3-high&lt;/td&gt;&lt;td&gt;92.11&lt;/td&gt;&lt;td&gt;100%&lt;/td&gt;&lt;td&gt;660.41&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;2&lt;/td&gt;&lt;td&gt;claude-4-opus-20250514-32k-thinking&lt;/td&gt;&lt;td&gt;88.9&lt;/td&gt;&lt;td&gt;100%&lt;/td&gt;&lt;td&gt;183.47&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;3&lt;/td&gt;&lt;td&gt;grok-4-07-09&lt;/td&gt;&lt;td&gt;86.15&lt;/td&gt;&lt;td&gt;95%&lt;/td&gt;&lt;td&gt;309.49&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;4&lt;/td&gt;&lt;td&gt;claude-4-sonnet-20250514-32k-thinking&lt;/td&gt;&lt;td&gt;85.4&lt;/td&gt;&lt;td&gt;100%&lt;/td&gt;&lt;td&gt;172.5&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;5&lt;/td&gt;&lt;td&gt;claude-3.7-sonnet-20250219-32k-thinking&lt;/td&gt;&lt;td&gt;70.75&lt;/td&gt;&lt;td&gt;90%&lt;/td&gt;&lt;td&gt;152.3&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;6&lt;/td&gt;&lt;td&gt;gemini-2.5-pro-preview-06-05&lt;/td&gt;&lt;td&gt;70.3&lt;/td&gt;&lt;td&gt;100%&lt;/td&gt;&lt;td&gt;1185.29&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;7&lt;/td&gt;&lt;td&gt;gemini-2.5-pro-preview-05-06&lt;/td&gt;&lt;td&gt;62.8&lt;/td&gt;&lt;td&gt;95%&lt;/td&gt;&lt;td&gt;61.27&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;8&lt;/td&gt;&lt;td&gt;gpt-4.1&lt;/td&gt;&lt;td&gt;45.75&lt;/td&gt;&lt;td&gt;65%&lt;/td&gt;&lt;td&gt;32.8&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;9&lt;/td&gt;&lt;td&gt;grok-3-beta&lt;/td&gt;&lt;td&gt;40.25&lt;/td&gt;&lt;td&gt;70%&lt;/td&gt;&lt;td&gt;124.88&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;10&lt;/td&gt;&lt;td&gt;claude-3-5-sonnet-20241022&lt;/td&gt;&lt;td&gt;39.35&lt;/td&gt;&lt;td&gt;45%&lt;/td&gt;&lt;td&gt;31.63&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;11&lt;/td&gt;&lt;td&gt;o3-mini-high&lt;/td&gt;&lt;td&gt;39.13&lt;/td&gt;&lt;td&gt;67%&lt;/td&gt;&lt;td&gt;1617.94&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;12&lt;/td&gt;&lt;td&gt;gemini-2.5-flash-preview-04-17&lt;/td&gt;&lt;td&gt;35.85&lt;/td&gt;&lt;td&gt;60%&lt;/td&gt;&lt;td&gt;140.4&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;13&lt;/td&gt;&lt;td&gt;o4-mini-medium&lt;/td&gt;&lt;td&gt;29.53&lt;/td&gt;&lt;td&gt;41%&lt;/td&gt;&lt;td&gt;974.77&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;14&lt;/td&gt;&lt;td&gt;o3-mini-medium&lt;/td&gt;&lt;td&gt;25.8&lt;/td&gt;&lt;td&gt;55%&lt;/td&gt;&lt;td&gt;558.54&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;15&lt;/td&gt;&lt;td&gt;gpt-4o&lt;/td&gt;&lt;td&gt;23.9&lt;/td&gt;&lt;td&gt;45%&lt;/td&gt;&lt;td&gt;87.37&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;16&lt;/td&gt;&lt;td&gt;qwen3-235b-a22b&lt;/td&gt;&lt;td&gt;20.94&lt;/td&gt;&lt;td&gt;33%&lt;/td&gt;&lt;td&gt;1510.99&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;The most shocking results come from the bottom of the table. OpenAI’s “reasoning-optimized” mini models are disasters—o3-mini spent an absurd 27 minutes per puzzle to achieve a pathetic 39.13 average score. That’s nearly 9x longer than Claude 4 Opus for less than half the performance. Meanwhile, humble GPT-4.1 managed a respectable 45.75 score in just 33 seconds.&lt;/p&gt;
&lt;h2 id=&quot;the-time-performance-paradox&quot;&gt;The Time-Performance Paradox&lt;/h2&gt;
&lt;p&gt;The benchmark reveals a critical insight about modern AI systems: more thinking time doesn’t necessarily mean better thinking. Look at these striking comparisons:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Gemini 2.5 Pro (06-05)&lt;/strong&gt; scored 70.3 but took nearly 20 minutes per puzzle&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Gemini 2.5 Pro (05-06)&lt;/strong&gt; scored 62.8 but finished in just 1 minute&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;o3-mini (high)&lt;/strong&gt; spent 27 minutes thinking to score worse than GPT-4.1’s 33-second performance&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This suggests that many models are stuck in inefficient reasoning loops rather than making meaningful progress. Claude’s models consistently demonstrate the best balance—they think efficiently, exploring productive paths rather than spinning their wheels.&lt;/p&gt;
&lt;h2 id=&quot;why-this-matters&quot;&gt;Why This Matters&lt;/h2&gt;
&lt;p&gt;For real-world applications, the time-performance tradeoff is crucial. Consider these scenarios:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Customer support&lt;/strong&gt;: Would you rather wait 11 minutes for a 92% accurate response or 3 minutes for an 89% accurate one?&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Code debugging&lt;/strong&gt;: Is a slightly better bug fix worth 4x the debugging time?&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Research assistance&lt;/strong&gt;: Do you need the absolute best answer, or the best answer you can get in a reasonable timeframe?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In most cases, Claude 4 Opus represents the sweet spot—near-peak performance at practical speeds. The 3.2% accuracy gap between it and o3-high is negligible for most use cases, but the time difference is substantial.&lt;/p&gt;
&lt;h2 id=&quot;the-surprising-winners-and-losers&quot;&gt;The Surprising Winners and Losers&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Winners:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Claude models&lt;/strong&gt; dominated the top spots with consistent speed and accuracy&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Grok-4&lt;/strong&gt; (86.15 score, 5.2 min) showed X’s model can compete with the best&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;GPT-4.1&lt;/strong&gt; delivered respectable performance at lightning speed&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Losers:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;o3-mini and o4-mini&lt;/strong&gt; are embarrassments—slow AND inaccurate&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Gemini 2.5 Pro (06-05)&lt;/strong&gt; took 20 minutes for middling results&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Qwen3-235b&lt;/strong&gt; spent 25 minutes per puzzle for the worst performance&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The failure of OpenAI’s mini models is particularly damning. These were supposedly optimized for reasoning tasks, yet they performed worse than general-purpose models while taking 30-50x longer. It’s a cautionary tale about the dangers of optimizing for the wrong metrics.&lt;/p&gt;
&lt;h2 id=&quot;the-real-lesson&quot;&gt;The Real Lesson&lt;/h2&gt;
&lt;p&gt;This benchmark teaches us that in AI, as in life, perfection is often the enemy of good. o3-high’s marginal victory comes at such a steep time cost that it’s rarely the right choice for practical applications. Claude 4 Opus emerges as the real winner—fast enough for interactive use, accurate enough for serious work.&lt;/p&gt;
&lt;p&gt;The results also expose the hollowness of some “reasoning-optimized” claims. True reasoning capability isn’t about thinking longer—it’s about thinking better. Claude’s models demonstrate this perfectly, efficiently navigating solution spaces while others get lost in computational dead ends.&lt;/p&gt;
&lt;p&gt;For developers building AI products, the message is clear: optimize for the full user experience, not just benchmark scores. A slightly less accurate model that responds in seconds will create more value than a marginally better one that makes users wait.&lt;/p&gt;
&lt;p&gt;Want to test your own models or see the full results? The benchmark code is available at [&lt;a href=&quot;https://github.com/redspringxyz/bracket-city-benchmark&quot;&gt;github.com/redspringxyz/bracket-city-benchmark&lt;/a&gt;].&lt;/p&gt;</content:encoded></item></channel></rss>