In March I started playing a parlour game with two machines. I named an adjective, Claude and ChatGPT each finished the line "as ___ as…", and the three of us argued about whose was better. For lethal, Claude produced a surgeon's silence in the recovery room, running to thirty words. I produced "as lethal as a sniper's sigh". Three syllables did the work of thirty, and the machines conceded the round.
The arguments were better than the lines. Somewhere in those adjudications a game appeared: a real one, for the App Store, in which you write similes against the dead greats and a judge tells you the truth about your line, a score and a written critique from a reader who flatters nobody. It is called As…As… and it ships under my Medlara imprint. This piece is about how it got built. Since I built it by talking to Claude, it is also about what has come to be called vibe coding: making software by describing and ruling in plain English while an AI writes the actual code. I want to tell you what that covers when the project is real, what goes wrong, and why you, with no engineering department and possibly no engineering, should try it.
Here is the game as you will meet it. Each morning one word goes out to every player at once, carried by a line from one of the greats and a dare: "Shelley said the following. Can you do better." You have until midnight, and your entry is a living draft, rewriteable all day, on the bus or in the bath. At midnight the judge reads the whole field in one sitting, the great's line hidden among the entries, and in the morning you learn where you placed and whether, in the judge's blind opinion, you outwrote Shelley. Between dailies there is a practise room with the rest of the dictionary, and a line good enough there is nominated for the living book itself, Wilstach's dictionary continued a century on, where your simile is shelved beside Twain's for as long as the book lasts.
The method, plainly
Almost none of the app's code came from my fingers. (It is written in Swift, the language of iPhone apps, which I read haltingly and do not write.) My hands are in the art and in the rulings. The work happens in two rooms. In one, Claude in chat: this is where the game is argued into existence, decision by decision, in sessions that read less like programming and more like an editorial meeting that occasionally produces a specification. In the other, Claude Code, a version of Claude that works directly on my own computer: it opens the project's files, writes and rewrites the code, and runs the results. The two never meet. They are coordinated by handover documents, briefs written at the end of a chat session and handed to the builder as instructions, with a standing order to halt and report on surprises rather than improvise.
Two documents rule the project. The specifications live in Apple Notes, of all places, because that is where I can read them in bed; when the big July redesign happened, the whole thing went into a note called Rewrite, numbered rulings in my own voice. The history lives in a log that every build session must update before it closes. An argument becomes a note, and the note becomes a build. The log records what each build did, and everything below is in that log, which is why I can tell it honestly.
What honesty requires
The demo culture around AI-built software is a highlight reel. Here is the rest of the footage.
The scoring was fake for months. The first playable builds ran on a component literally named RandomJudge: the verdict cascade, the gold stars, Shakespeare's forty voiced sayings, the whole ceremony staged over a number drawn at random while the real judging machinery was still being argued about. The theatre was built before the truth was, and playing it that way taught me the theatre carries more of the game than an engineer would ever admit.
Claude overbuilds. Left alone with a scoring system it produced radar charts and heptagon visualisations nobody had asked to keep, handsome and useless, and the brief had to be amended to strike them. The model's default setting is eagerness, and eagerness is expensive to review; you learn to write briefs the way you would write for a brilliant contractor who bills by the surprise.
Things break, constantly and unremarkably. A chat session died mid-write and took an afternoon's hand-scoring with it. Dropbox quietly corrupted the project's saved history and was banished from the workflow. One bad afternoon left a live key to the AI service sitting in the machine's memory of typed commands, and it had to be cancelled and reissued. My own hand-correction of the garbled scan of Frank Wilstach's 1916 Dictionary of Similes, the book that seeds the game's word-hoard, its corpus, went in the bin the moment Claude found a cleaner copy online. I will not itemise the rest, because the itemising would miss the point. In this way of working, iteration is not a phase you pass through on the way to being right; it is the state of being. Versions are cheap, a reversal costs an afternoon, and nothing is precious except the rulings. The log records reversal after reversal with complete equanimity, and by about week three, so did I.
Iteration is not a phase you pass through on the way to being right; it is the state of being.
Taste is the hard part
A simile game is a machine for judging, and judging is taste, and taste was the one thing I refused to let the model assert. Early on Claude offered scores with the easy confidence of a machine that has read everything. I threw them out as scaffolding.
The replacement was slower. We separated what is fact about a line from what is opinion. That a string is a simile at all is close to objective and gets stored as fact. Whether the line is any good is opinion, and opinions are stored as disposable annotations, each stamped with the version number of the marking scheme, the rubric, under which it was given, so that when my taste changes the entire corpus can be rescored in one overnight run and nothing else moves.
Taste is the one thing you have that the machine does not.
The rubric itself was fitted, not asserted. I hand-scored a few hundred lines, spread deliberately across the quality spectrum, and the seven axes came out of my scores rather than going in ahead of them: aptness, reach, recognition, depth of mapping, particularity, economy, sound. My anchor pair, for calibration and for the soul: "as final as Chicxulub" against Twain's "as final as death's decision". The asteroid wins, and the reasons it wins, written down, became the standard.
The judge was rebuilt several times, and each rebuild was a position on what judgment is. Verdicts went blind: both lines shown unattributed, in random order, so no reputation could lean on the scales. The shipping judge is a small model that runs on the phone itself, voiced as Wilstach the lexicographer: you submit your line to the man who spent his life collecting them. Above him sits a court of appeal, the largest model in the stack, permitted only to affirm or to raise, never to lower; the product of an appeal is the letter, proof the line was understood.
The economics of one person
Nothing here would matter if the sums failed, and for a solo builder the sums are now a design surface. The everyday judge runs on the phone and costs nothing. The daily competition is judged by the most expensive model in the catalogue and stays affordable by construction: every entry in the game rides one overnight run, the judge reads the whole field comparatively, and the arithmetic lands under a penny per entry, inside a subscription of about ten dollars a year: a personal literary judge on retainer for the price of a paperback. Appeals are bought singly, priced as the luxury they are.
I did not produce this architecture and neither did Claude. It fell out of arguments in which I kept saying too expensive and the model kept proposing until the shape held. The first paid corpus run went ahead under a written five-dollar ceiling with an instruction to abort if the projection exceeded it. If you take one operational habit from this piece, take the ceiling: a number goes into the brief, and the model must respect it and report against it.
Unbuilding
The bravest decisions delete, and the log's happiest entries are subtractions. The timer went when the first true sixty-second rounds showed the clock was measuring typing speed rather than wit, and wit is what is judged. A bronze-coin score presentation lasted one evening (the log's verdict: "looked like shit"). The deepest cut came in June, when the game's entire first half, a campaign against twenty-four dead writers with engraved portraits whose eyes follow your finger around the board, was frozen so the simpler half could ship alone. Months of craft went into the archive without ceremony.
Yes, folks will cheat, but if that generates a better simile, so be it.
What survived is sparer and stranger than the frozen half, and it is the game described at the top: one word, one day, one reading. Overnight judging turns a delay into a ritual; the submission screen says: "Entered. The judge reads tonight." The great's line rides the night's judging as an ordinary contestant, so outwriting Shelley is a judged result rather than a compliment; in the first real overnight run, Shelley's own line came back scored 31 out of 33, blind, close to where the rubric's anchors said he belonged. On the question of whether players will cheat by drafting with an AI of their own, the ruling stands in the spec verbatim: yes, folks will cheat, but if that generates a better simile, so be it.
The invitation
The point of this confession is not that As…As… is remarkable. It is that the method is now available to anyone, and I want to talk you into trying it.
Start with the reframe. In the old world, software was expensive to attempt: a wrong version cost a quarter and a salary, so you specified everything in advance and prayed. In this world a wrong version costs an afternoon, so you can stop dreading wrongness and start spending it. Iterative is not a stage of the work; it is the state of being, the water this world swims in, and once you accept that, throwing work away stops feeling like failure and starts feeling like steering. I froze a month of portrait craft in June and slept fine.
What the method needs from you is not technical. It needs a domain you know cold, because the model amplifies judgment and cannot substitute for it. Mine is the sentence; yours might be bird ringing, conveyancing, sourdough, the tax treatment of woodland. The right project is the thing you already argue about at dinner, because the whole method is that argument, moved indoors and given hands.
It needs rulings in writing. The single biggest difference between a toy session and a shipping product was the moment my decisions started going into documents the model treats as contracts. Wishes produce demos; rulings produce software. Keep one log, updated at the close of every session; this essay was written out of mine.
Wishes produce demos; rulings produce software.
A ceiling on spend belongs in every brief, a number the model must respect and report against, along with the standing order to halt on surprises. The model's failure mode is enthusiasm, and yours will be trust; the written documents keep both of you honest.
Above all, bring the taste discipline, because taste is the one thing you have that the machine does not. Sit down with a hundred examples of the thing you love and score them by hand, and let the standard be written down from your judgments rather than handed to you by a machine that has read everything and believes nothing. What comes back is your own taste, made explicit and arguable, and the mirror is worth the sitting.
So start tonight, and start smaller than I did. Open a chat and name your domain. Ask what good means there, refuse the answer until it matches what you already know in your hands, and write the first ruling down. Do not ask for code; code is the last thing you will need, and by the time you need it the machine will write it faster than you can describe it. The building turned out to be the easiest part of my year. What no machine wrote is the sentence at the top of the note: we are going to do a significant rewrite of the AsAs game. Yours will say something different.
When the game launches, a word will go out to everyone at once each morning. Somewhere a player will spend the whole day rewriting one line, and at midnight the judge reads. I would like that player to be you, twice over: once in the game, where the day's word and Shelley's dare will be waiting, and once in your own domain, where the next impossible thing waits for someone who knows it cold. As…As… arrives on the App Store as soon as Apple's paperwork allows, which no ruling of mine can hurry. Bring one good line.
Discussion