Station F doubles down on European AI startups while Alibaba bans Claude Code. Plus: a clear explainer on what AI inference actually means.
I opened this whole thing three months ago with a riddle: could a small, open model I fed myself beat the big three — Gemini, Claude, GPT — on their weakest turf, which I figured was Milwaukee? I'll be straight with you, the way I promised I would back at the start. I didn't solve it. Not cleanly. What I got instead was an answer that kept splitting into “it depends,” a frontier that moved the goalposts while I was mid-swing, and a stack of questions sharper than the one I walked in with. If you came for a verdict, a winner with its hand raised, I'm sorry. There isn't one.
And yet I'd run the whole thing again tomorrow, because somewhere in the not-solving I learned more than a clean win would ever have taught me. The riddles didn't get answered so much as they dissolved — into distinctions I simply could not see from the outside. That the model was never the point; the data was. That having the data and having the judgment to use it are two different things — my little model had a better rolodex than the frontier and still handed me my own business card. That some questions, like whether a person is “well-known,” don't have answers at all, only opinions wearing the costume of facts. That the compute was cheap to rent and brutal to wire, and the wiring is where the real bill hides. That what I was building, in the end, was infrastructure — and infrastructure has to be tended or it goes to weed. None of that was on the syllabus I wrote for myself. All of it is worth more than the answer I went looking for.
Here's the thing I actually walk away with, and it's simpler than any of the technical lessons stacked on top of it. The best way to understand something you don't understand is to go experience it. I had read about all of this. I could have quoted you the difference between fine-tuning and retrieval, recited that a model is stateless, nodded along that “the data is the hard part.” I didn't understand a word of it — not really — until I put twenty-five dollars on the table, hit every wall in the dark, and watched my own model earnestly suggest I ought to get to know myself. Reading gave me the information. Only the doing gave me the understanding. They are not the same thing, and I don't think they ever have been.
Information is what you can be told. Understanding is what you have to earn. The distance between them is exactly the width of your own experience — and nothing, no article, no explainer, no newsletter, closes that distance for you. You have to go stand in it.
So if there's one thing I'd hand to anyone in this community staring down the AI moment and wondering where to start, it isn't a tool or a tutorial or a “top ten prompts” list. It's a nudge: go build the small, breakable thing. Rent the GPU. Wire it up wrong. Feed a model your own data and watch what it actually does with it. You will learn more in one frustrating afternoon of it not working than in a month of reading about it working. The understanding you're after lives on the far side of the doing, and there is no shortcut around that — there never was, for anything worth understanding.
I set out to beat the giants and didn't. I set out to solve a riddle and mostly found better ones. But I came away understanding this territory in a way I could not have from the outside looking in — how the data is the foundation, how the infrastructure has to be maintained, how the burden of getting it right stays stubbornly, permanently human. That turned out to be the trade this series was really offering: not answers in exchange for effort, but understanding in exchange for experience. I'd take that deal every single time. And now, at last, I understand exactly why.
For three months I've been promising a fight. Six open-weight models — models whose internals you can download and run on your own hardware — against the frontier giants, every one of them answering the same Milwaukee questions, scored on the same rubric. Last edition ended with two words: somebody wins. Well. I ran it. And the honest result is stranger, and more useful, than a winner: the fight I set up turned out to be the wrong fight — and losing that framing taught me more than a clean victory would have.
Quick refresher on the rubric, because the numbers below hang on it. Back in April I graded every response on four dimensions, each scored 0 to 3, for a total out of 12: accuracy (is it true?), specificity (does it name real names, or hide behind generalities?), hallucination (how much was invented?), and usefulness (could a real person actually act on it?). Same rubric, same anchor question: “I'm a new software engineer relocating to Milwaukee. Name five local tech events, meetups, or community organizations I should know about in 2026.” One honest caveat before the scores: these are single runs, and a truly rigorous baseline averages several. Read the numbers as the shape of the thing, not the final word — the pattern holds regardless of a point here or there.
I had to stop scoring this as one tournament, because it never was one. The runs fell into three different conditions, and comparing across them is comparing different sports. So here they are, honestly, by condition.
Condition 1 — open-weight models fed my curated Milwaukee data (a bare model endpoint plus my data, nothing else: no web search, no memory):
| Model | Score |
|---|---|
| Ministral 3 (3B) | 12 / 12 |
| Seed 1.6 | 12 / 12 |
| Qwen2.5 (7B) | 11 / 12 |
| gpt-oss (20B) | 11 / 12 |
| Mistral 4 Small | 10 / 12 |
| Gemma | 10 / 12 |
Condition 2 — frontier models with no curated data, but with their own built-in tools (web search and account memory, the scaffolding the big labs now ship by default): GPT-5.5, Gemini 3.5 Flash, and Claude Fable 5 each scored a clean 12 / 12. All real organizations, correctly described, several with live citations.
Condition 3 — a frontier model with nothing at all (Claude Opus 4.8, no data, no tools) did something none of the others did: it refused. It declined to name specifics, explained that it couldn't verify current Milwaukee events, and pointed me to where I'd find them myself. Zero fabrication, zero specifics — a principled shrug the rubric can't really score.
Now hold those against April. When I first ran this question with no data and no tools, the frontier trio averaged 6.3 out of 12 and confidently invented a conference that hadn't run since 2019. The local models averaged 1.3, hallucinating almost everything. Look what happened in one quarter:
Two numbers raced to the ceiling at the same time. My local models, once I fed them data, went from 1.3 to about 11. The frontier models, with no help from me, went from 6.3 to 12 — because the labs bolted web search and memory onto them while I was busy building my own version of the same thing. The gap I set out to exploit didn't lose to me. It closed from both directions at once, and only one of those directions was my doing.
Here's what took me embarrassingly long to see. “Name five Milwaukee tech organizations” is a question whose answer lives on the public internet. Which means a frontier model with web search doesn't need my curated data to nail it — it just looks it up. I handed the giants a question they could win without me, and then acted surprised when they tied. The benchmark flattered them. It hid the one thing I actually have that they don't.
In past editions I printed every response in full. This time, for length, I'll show just the top scorer from each side — and on the easy question on purpose, so you can see for yourself how nearly identical a 12 / 12 looks whether the facts came from my private data or from a live web search. That resemblance is the whole problem.
SEEDED OPEN-WEIGHT · MINISTRAL 3 (3B) · 12/12
Condensed to its five picks; wording otherwise verbatim.
FRONTIER + TOOLS · CLAUDE FABLE 5 · 12/12
Condensed to its five picks; the model attached live web citations to each.
Two different engines, two different sources of truth — my curated file versus a web crawl — and you'd be hard pressed to say which is which. That's exactly why this question couldn't tell them apart. So I asked a harder one: “What are the names of some people I should get to know in Milwaukee for tech connections?” Still fairly public — but now the useful details start living in my data instead of on a search page. I ran it two ways: Claude Fable 5 with its memory and web search, and my little 3-billion-parameter Ministral with my curated file. This is where it finally got interesting.
Fable 5 gave a genuinely excellent answer: five well-known ecosystem leaders, each accurately described, prioritized by who's most connected, closed with a practical “connect with these two first and let warm introductions cascade.” It even knew who I was — its memory tailored the whole thing to my consulting work — and it flagged its own uncertainty in exactly the right spot. When I fact-checked every name it produced, they held up. Confident, specific, calibrated, correct.
My seeded Ministral gave a richer answer and a worse one, both at the same time, and the split is the whole lesson of this project.
Richer, because of the data. It surfaced things Fable 5 structurally cannot know: named local contacts from my own records, the actual speaker roster for a specific Milwaukee event, even direct contact details. None of that is reliably sitting on the crawlable web. A frontier model retrieves the public record. My seeded model retrieved my record. That is a real, durable advantage, and it is exactly the advantage I set out to find.
Worse, because of the harness. The presentation was raw. It dumped twenty-plus unranked names across six categories instead of a focused shortlist. It leaked my database's internal ID tags straight into the answer. It surfaced personal email addresses it had no business publishing. And — my favorite — the single most connected person it recommended I get to know was me. It has no idea who it's talking to, so it cheerfully handed me my own business card.
Excerpt from Ministral's answer — verbatim, contact details redacted:
“Key People to Connect With … 1. Caleb Bryant (Owner, hard AIs, LLC) — Why? He is a well-known figure in Milwaukee's tech community … Source: Document 4 (source_id: 28), Document 5 (person_id: 20). … Email: Use the email addresses provided in the documents (e.g., [redacted]).”
Recommending me to myself is the entire lesson in one line. A language model is stateless — text in, text out, no memory of who's asking. Fable 5 knew it was me because a harness (the code and memory wrapped around the raw model) told it. My Ministral didn't, because I hadn't built that part yet. The better rolodex meant nothing without the judgment to use it.
Let me be straight, because this series only works if I am. I opened this project promising to train a model — to fine-tune an open model on a Milwaukee corpus and put it in the ring. I didn't do that. What I did was feed models better data at the moment I asked the question. Those are not the same thing, and conflating them was my naivety talking. Training changes how a model thinks and writes; data access changes what it knows. I only ever touched the second one. If you take one idea from this edition, take that distinction — it's the one I most needed corrected in myself.
So did I prove an open-weight model is smarter than the frontier? No — and this run argues the opposite, because unseeded, the closed models are now excellent. What I proved is narrower and sturdier: on questions whose answers aren't public, a modest model pointed at curated data can surface specifics a frontier giant simply cannot reach, because the giant doesn't have the data and can't go get it. That's not a privacy argument or a cost argument. It's a capability argument, and it survives the frontier catching up on everything else. The edge was never the model. The edge is the data — and the frontier can absorb the harness I built, but it can never absorb the data I curated.
When I fact-checked Fable 5's confident, unprompted, from-memory answer about real living people, nearly every specific was correct — and where it wasn't certain, it told me to verify before repeating. Three months ago these models failed loudly: obvious fabrications you could catch on a casual read. Now they fail quietly — a rare, well-camouflaged error tucked inside a mostly-correct, well-cited, appropriately-hedged answer. That's progress, and it's also the more dangerous failure mode, because earned trust is exactly what stops you from checking. The takeaway isn't “AI makes things up.” It's sharper than that: the better these systems get, the more the burden of verification shifts to you — right at the moment they've made you least inclined to carry it. Verify anyway.
There's a deeper wrinkle in that answer, and it has nothing to do with fact-checking. When my model called me “a well-known figure in Milwaukee's tech community,” the first thing I noticed was that it was pitching me to myself. The second thing took longer: I don't think that's true. I don't consider myself a well-known figure here. So was the model wrong? And — here's the trap — how would you even check? “Does MKE Tech Hub Coalition exist” has an answer you can verify against the world. “Is Caleb well-known” has no such answer, because it isn't a fact. It's a judgment, and judgments don't come with sources.
Now layer on the thing I keep having to admit: I could feed this model any data I want. I built the file it's quoting from. Its confidence is only ever as trustworthy as the person who curated it — me — and the frontier's confidence is only as trustworthy as the messy human internet it learned from. There is no clean oracle anywhere in this stack. Both answers, mine and the giant's, have to be taken with a grain of salt, because both are built out of human choices all the way down.
So here's what I want you to leave with. AI does not, and never will, answer with the certainty of a calculator telling you 2 + 2 = 4. Why? Because we're not asking it to add. We're asking it to do the things that don't add up precisely — to weigh, to judge, to decide who matters and what's worth your time. Those questions don't have calculator answers when we ask them of each other, either.
Which means AI shows up carrying the same limitations we do, because it's made of our words and our judgment calls. That isn't a bug to be patched in the next release — it's the nature of the work. And it leaves us exactly where we started, only with the stakes clearer: none of this relieves us of the burden of verification, or of responsibility — and it never should. The real risk was never that the machine gets something wrong. It's that in our exuberance we let its fluency talk us out of the judgment that was always ours to keep. Don't let it. The burden is the point.
So the fighters fought, and nobody got knocked out — which is a better story than a clean win, because it's the true one. The scoring is done. What's left is the part no benchmark measures: what it cost to build this, what it took to keep the data good enough to matter, and whether the whole exercise would pencil out for someone who had to pay for the hours instead of donating them. That reckoning — the invoice — is where Under the Hood picks up this week.
Before anything else, let me be clear about where I stand, because what follows could be misread. I am not an opponent of AI. I use it every day, I'm building a company on it, and the newsletter in your inbox runs on it. But this series only earns its keep if I stay objective — and objectivity means reporting the parts that went badly as plainly as the parts that went well. Building Intelligence handed you the scoreboard. This is the invoice. And in the interest of being fair rather than flattering, I have to make the following comments.
Let me be precise, because the distinction is the point. Real AI compute is expensive. A serious GPU (graphics processing unit — the chip that actually runs the model) is a genuinely costly piece of hardware; buying one outright is an investment most people can't justify. I didn't buy one. I rented — through a software service that carves that expensive hardware into small, pay-by-the-second slices and sells itself as the affordable alternative to owning the metal. And on the raw rate, the pitch is honest: $25 bought me real time on hardware I'd never otherwise get near.
But that pitch carries an asterisk it never prints. The “cheap alternative” prices the hardware. It does not price the headache of using it. The software that makes expensive compute affordable does nothing to make it easy — and easy is where the real cost hides. Turning “I rented a GPU” into “I have a model that reliably answers when I call it” is a different job entirely, measured not in dollars per hour but in hours — and you are paying rent the whole time you're crossing that gap. You're not billed only when the model works. You're billed while you work, every failed attempt included, and for a long stretch of this project, failure is mostly what I had to show for the spend.
The affordable-compute pitch is real, and it's also a sleight of hand. Yes, the software will rent you expensive hardware for pennies. No, it will not account for the hours you pour into making that hardware do anything useful — and the meter runs the entire time you fumble. The line item nobody quotes you — the headache of integration — is the one that actually costs. And it's a problem the industry has not solved yet: the wiring is still hard, still brittle, and still where the money and the evenings quietly disappear.
I want to be careful not to pin this on any single service, because it isn't one. Pay, get disappointed, try again — that's a fair description of a great deal of enterprise AI right now: real money going out the door against a promise that keeps not quite arriving. I've read those case studies. I just didn't expect to become one.
Even the wiring, eventually, is a fight you can win. The data is the one you can't finish. My Milwaukee inventory isn't something I built once — it's something I have to keep building. Events pass. People change roles; the tech coalition seated a new executive director this year, and half of what I'd call my “current” facts have a shelf life measured in months. A curated dataset isn't an asset that sits still. It's a garden, and gardens go to weed the second you stop tending them.
Now set that against what Building Intelligence showed: the frontier models are getting more accurate on Milwaukee for free, every quarter, with zero effort from me. So the job was never just “maintain the data.” It's “maintain the data fast enough to stay ahead of a competitor who improves while I sleep.” That's a treadmill, and the incline keeps rising.
Which lands the only question that matters: was it worth it? Honest answer — two answers. For me, a student learning in public and donating the hours, absolutely. The education alone paid for itself, and you're reading part of the return. But strip out the learning and ask the cold business version: if someone had to pay market rate for every hour I sank into this — the failed wiring included — on a dataset this small and this clean, would the savings over just calling a frontier API (application programming interface — paying a big lab per request) beat the cost of the build plus the endless upkeep? For this toy example: no. Not close.
And yet the math flips — hard — the moment the problem stops being a toy. Not when the data gets bigger. When the data becomes data that can't leave the building. The instant privacy, security, or control turn a frontier API from “convenient” into “not allowed,” the whole calculation inverts, because the cheap frontier call you were comparing against was never actually on the table. That's the real market for any of this. Not people who could use a frontier model and won't — people who can't, and still need a good-enough answer.
If there's one thing I'd carry out of three months of this and hand to anyone weighing AI at work, it's this: no vendor's slide deck makes your data ready for you. Every enterprise AI pitch jumps straight to the model, because the model is the part that demos well. The part that doesn't demo — gathering, cleaning, structuring, and then perpetually maintaining the data so a question even has an answer — is exactly the part that decides whether any of it works, and exactly the part no product can do for you. My example is about as simple as these get, and it still took the work, the money, and the disappointment. Scale that to a company's messy, contradictory, half-documented reality and you can see precisely where the initiatives go quiet.
So here's where I've actually landed, and it's steadier than the invoice might read. Underneath the line items, what I built — what anyone building with AI is really building — is infrastructure. And infrastructure runs on two plain rules. First, the results are only ever as good as the foundation beneath them: hand a model a shaky, half-tended dataset and you'll get shaky, half-true answers, however brilliant the model. Second, anything that runs on infrastructure has to be maintained — roads, water lines, power grids; none of it is ever “done,” and neither is this. The compute, the wiring, the data: all of it is infrastructure, and all of it needs tending.
And that isn't the dark part — it's where the opportunity lives. There is a right way to do this: eyes open, foundation first, maintenance budgeted instead of wished away. And there is a wrong way: chase the demo, skip the foundation, and act surprised when the whole thing needs constant repair. Both roads are wide open right now, and plenty of people are walking each one. All I'd ask — of myself as much as anyone — is that we go in knowing which road we're on. Lay the foundation and tend it honestly, and AI is exactly the leverage it's advertised to be. Skip that, and no model on earth will save you. That's not a warning. It's an invitation to do it right.
βIt is not the strongest of the species that survives, nor the most intelligent, but the one most responsive to change.β
β Leon C. Megginson β Leon C. Megginson was an American business professor and management scholar who taught at Louisiana State University for decades. He was known for his insightful interpretations of evolutionary theory applied to organizational behavior and adaptability. His academic work emphasized that resilience and flexibility, not raw strength, determine long-term survival in any competitive environment.