Project status
Editing practice sites: what works today, and what’s coming
Worth testing right now
- Ask it questions, not just edits — this is the hunt today — we just caught the assistant answering “what are our hours?” as if the site had no hours in it. The site holds 227 facts about its hours. The answering tool only ever knew how to look up phone and email, and everything else came back empty as though the site were blank. Ask it about hours, a doctor, a policy, a price — anything you can see with your own eyes. Every question that comes back empty maps a hole, and this one hole is our leading explanation for why an edit takes four minutes instead of one.
- Push the six that work — change wording, add a sentence, take a change back, ask where something appears — those four are proven. Two more landed this week and nobody outside the build has touched them yet: adding a real bullet to a list, and changing shared text on one page without disturbing its twin. Being the first person to drive those two is the single most useful hour anyone can give us. Odd phrasings and changes touching several pages are the good kind of testing; headlines especially, since our own audit found a headline built from several pieces that is unproven on this path.
- Still gold — asks with several parts in one message — “change X, remove Y, and fix Z.” That fix is built but parked mid-build, so examples from before it lands still shape it. So is any ask that earns an honest “can’t do that yet” — each one calibrates the build order; a popup test-drive this week literally reordered the queue.
- Most valuable of all — an answer that goes vague — “scanned 11,706 facts… found nothing” instead of a clear yes or an honest “not yet.” We’ve now measured this in four places: form labels, page code, the data search engines read, and plain questions about the site. Hitting one and telling us is worth more than a dozen clean passes.
- Worth exactly one try, then stop — moving a section. Not because it works — it doesn’t — but because we measured that asking in plain words doesn’t even reach the honest “not built yet”; it dies in machinery first. Seeing that yourself and telling us is useful. New pages and photos: the assistant declines honestly, no need to spend your evening there.
- Reporting — the exact words you used, and what came back — to Robert. That’s the whole protocol.
One pilot practice · private previews only · nothing goes live on its own.
Robert is reviewing the wording before this goes to the whole team. The facts are current — please hold off forwarding for now.
Last updated Aug 12, 2026, 8:58 PM UTC
What you can ask for
“Ready” means proven — an independent reviewer tried hard to break it, and a person who had never seen the assistant used it cold, successfully, in their own words. “Landed” means the first half of that: built, independently attacked, and on the site’s working copy — but no newcomer has driven it yet. It works; it isn’t proven. We’d rather show you both facts than round one of them off.
| What you want | Status | Plainly | If you ask today |
|---|---|---|---|
| Works today | |||
| Change any wording on the site | Ready | “Change ‘Restrain policy’ to ‘Leash & Carrier Policy’” Done, with a receipt. |
Works — with one exception we found by auditing ourselves: a headline built from several pieces is unproven on this path. Queued to fix. |
| Add a plain line of text — a sentence, not a bullet (that one’s below) | Ready | “Add a note that we include a free nail trim” It lands where you said. |
Works. |
| Take a change back | Ready | “Undo that” Restores the site exactly as it was, down to the last letter. Landed today, full cycle: taking back a change from ANY chat — built in hours after a teammate hit the wall, then an independent adversary found seven edge-case flaws, then every one was fixed and turned into a permanent automatic check. Six proofs now guard it on every future change, staged on the real incident’s own bytes. The assistant lane picks it up next. |
Works. |
| Ask where something appears | Ready | “Where does the old phone number still show?” Honest, complete answers. |
Works. |
| Add formatted items — a real bullet with its dot | Landed, not cold-tested | “Add a bullet to the services list” You get a real bullet, dot and all — not a plain sentence pretending to be one. Landed overnight and merged. An independent check first caught it offering choices it couldn’t honour on about half the site’s real lists, then caught the honest version forgetting to say where the text lives — both were fixed, and the last round found nothing left to break. And where a list is built in a way the assistant can’t safely copy, it now says so and points at exactly where that text lives, instead of offering you a choice it can’t keep. |
Should work. Nobody outside the build has driven it yet — you’d be the first. If you get an honest “not built yet” instead, that gap is itself a finding: tell Robert. |
| Change shared text on just one page | Landed, not cold-tested | “Change this line on the cats page only — leave the dogs page alone” The two become separate truths from then on, linked both ways; “change it everywhere” still folds them back in one ask. Full cycle in a day: built in 23 minutes, then an independent adversary caught three places its receipts could mislead — including failing to name the pages deliberately left alone. Every one was fixed, turned into a permanent automatic check, and merged. |
You’ll be asked an honest question first — nothing changes without you. No cold newcomer has driven this one yet either. |
| Coming next — what each one changes for you | |||
| Create a whole new page by asking | Being verified | When this lands: you ask for a page — “a page for our dental services” — and get one, in plain words. The second verification round came back with more to fix — honest reds. Every one was ruled on the same day and the fix build is running right now. Nothing ships until a round comes back clean. |
You’ll get an honest “not built yet.” Tell Robert you asked — it helps. |
| New pages come out fully editable — no dead corners | Being verified | When this lands: nothing on a newly created page is quietly uneditable — it’s changeable in plain words, all the way through. The newest verification round found two more honest reds — both were ruled the same day, and the next fix is being built right now. Nothing ships until a round comes back clean. |
Not something you ask for — it’s a guarantee being proven behind the scenes. |
| Ask for several changes in one message | Being built | When this lands: one ask carrying several changes gets every part answered in that first reply — one receipt covering all of it, and one undo that takes the whole ask back. Most of it is built. The machine building it died three times on the same infrastructure fault, so the work was rescued intact and parked; it gets finished by hand rather than dispatched a fourth time. Parked mid-build, work safe |
It may act on part of what you asked. Say which part went unanswered — that’s exactly the report that ordered this fix. |
| Move or reorder sections | In line | When this lands: you move a section up or down by saying so, without filing a request — and move one into another section’s box just as easily. The two open design calls were resolved today: the approach dissolved into the site’s own record system, so a move is a recorded rearrangement rather than surgery on the page. Nothing here waits on a decision anymore — it builds next in line, the moment new pages clear verification. Resolved today |
Worse than a plain no, and we own that: our audit measured that asking in plain words dies in machinery before it ever reaches the honest “not built yet.” Only a rigid, technical phrasing gets that far. Fixing the plain-words path is part of this build. |
| Take something OFF the page — a paragraph, a section | In line | When this lands: “remove that third paragraph” takes it off the page without losing it — it stays parked in the site’s own records, restorable with one word. Design finished hours after a real teammate hit this exact wall live, and its build lane is now queued in parallel: it runs beside the move work, not behind it. Designed today |
You’ll get an honest “can’t remove it yet” — plus the honest workaround: rewrite it so it earns its place. |
| Photos, video, forms on the site, redirects, and the rest | Not built yet | When this lands: swapping a photo or updating a form is the same plain-words ask as changing a headline. Mapped and ordered — each one gets built, independently attacked, and cold-tested before it ever says “Ready” here. Photo swapping and form editing finished their design pass today and joined the build queue; video, redirects and the rest are still ahead of their turn. |
Usually an honest “not built yet” — but not always, and the exceptions are the ugly kind. A form-label ask can come back vague instead of refused. And our audit caught an image swap that half-succeeded and reported itself green: it changed 603 places and left 1,019 copies behind, and said so. Half a change that calls itself done is the worst thing on this page. If you hit either, tell Robert — they’re top of the fix list. |
What’s useful to try
The rows marked Ready work right now — on the pilot practice’s private preview, not on any client’s live site. If you want to see one, ask Robert for a demo, or tell him the exact words you’d use and he’ll try it on the pilot.
And if you ask for something that isn’t built yet, say so anyway. That’s not a dead end — it’s the signal that orders this list. The two rows now first and second in line got there because testers hit those exact walls.
On timing: we don’t put dates on these, because each one must survive an independent attempt to break it and a cold test by a real person before it says “Ready” — we’d rather be late than wrong. What we promise is the order, and it’s the order above. One honest number: an edit currently takes about four minutes to land; getting it under one minute is a formal part of “done” (below), and today we count that as failing — not “nearly there.” And one more honest note from this week’s first real teammate test-drive: if you ask for several changes in one breath, today the assistant may act on part of it — and it owes you the fate of every part, up front. The fix that turns that into one answer, one receipt and one undo for the whole ask is in the build oven right now.
Overall progress
Of every kind of edit a human engineer could make to a site — a full inventory, no exclusions — this many already work in plain words at the assistant.
Works today (31)Still to come (27)
The twelve rows above group all 58 into the kinds of ask you’re likely to have. Worth knowing what’s in the remaining 27: photos, video, on-site forms, and redirects are still in there — if those are most of what your clients ask for, that’s exactly why they’re queued and listed, never silently missing.
What “done” means
The bar we hold ourselves to. All five, or it isn’t done.
- 1
Everything editable, or honestly labeled. Every piece of the site can be changed in plain words — or it carries a visible “not yet” with a reason. Nothing is silently uneditable.
- 2
Independently attacked. Every capability survived a reviewer whose whole job was to break it.
- 3
Cold-tested by a human. A person with no training used it successfully, in their own words.
- 4
Receipts true, undo exact. What it says happened, happened. Undo restores the original perfectly.
- 5
Under a minute. A routine edit lands in under sixty seconds. Today it takes about four — we count this one as failing, out loud, until it’s measured green.
For new sites
“Born fully editable” means a brand-new site from the factory — the assembly line that builds sites in this system — meets all five on day one, checked automatically at birth, not assumed. The factory’s own health checks are being repaired now to make that possible; none of that work touches any live client site.
What changed recently
- Aug 12
The assistant’s own answering tool has been blaming the site for its own gap. Asked “what are our hours?”, it replied as though there were no verified answer — while the site holds 227 facts about its hours. It only ever knew how to look up three things: phone, email, telephone. Everything else came back looking exactly like absence. An assistant reading that reasonably concludes the hours aren’t there, and then either says something false or burns minutes hunting through 11,706 facts by hand. It was found in ten minutes by someone using the tool the way a teammate would — four scored test-drives never caught it, because none of them asked about hours. A first fix is built and waiting on independent review.
The ruling behind that fix reshapes the whole system, and it is one sentence: the site’s records are a place, not a speaker. A book doesn’t read itself to you. Every stock phrase, apology and template the machinery had grown gets deleted, and the assistant does all the speaking, in its own words, every time. What stays is not speech: the machinery still refuses an unsafe change and still records who did what, because it must never be the one to certify its own work. That’s a locked drawer and a ledger, not a voice.
The redirect and video work came back clean from its fourth independent review — and the story is the landing, not the fix. Its cure deleted far more than it added: 142 lines removed, including a second permission mechanism it had quietly grown of its own. Then, mid-landing, reading the conflicts showed that finishing by hand would have silently re-opened the faked-permission hole that took two rounds to close. The landing was stopped and routed to a proper one. Standing rule out of it: anything built before a change to the trust machinery gets merged with proofs, never by hand.
Every kind of change a person could want to make now has a plan behind it. The last five areas — files and photos arriving, the insides of forms, links and settings, changes scheduled for a future date, and multi-step recipes — were designed and ruled in a single night. Nothing on the list is unowned any more: each one is working, being built, under review, or written down as a deliberate limit with its reason on its face.
We now score a test-drive on five numbers instead of a feeling: was every sentence true, did the ask actually land, how long it took, how much the person had to carry themselves, and how clearly it spoke. One false sentence fails the entire walk no matter how well the rest went. By that bar, four of four recorded walks fail today. That is the honest floor we climb from, not a setback — and Robert and Alie score one transcript a week so the ruler stays calibrated to their taste rather than ours.
The gate we turned out not to have: nothing mechanically stopped code from landing on the main line with its checks red. Our own discipline was the only thing holding, and it slipped — the main line sat red for about 77 minutes across four pushes and nothing objected. Named out loud rather than quietly patched; switching on real enforcement is Robert’s call to make.
Apostrophes survive. It sounds trivial and it wasn’t — wording with an apostrophe in it used to come apart on the way in. It landed with a 21-case letter-exact corpus behind it, and a piece of formatting left unclosed now refuses the whole change rather than half-applying it.
Asking what’s waiting to go live works for real now — in production, proven against the real site. Every filed change from every chat, with who asked, when, in their own words, and a code to take each one back. That missing list is exactly what broke both of the first two cold test-drives.
The moment where the assistant asks “are you sure?” now has a real backbone under it. An independent reviewer manufactured a yes — built one from outside the system, having never been shown the question — and drove a change through with it. The fix: asking the question writes its own record, and your answer is looked up against that record instead of taken on trust. A yes can’t be faked, used twice, or borrowed from someone else’s change. It landed after a second reviewer round found nothing left to break.
We now keep a written list of every limit we have deliberately chosen — each one with its reason and with what would make us revisit it. The law behind the list: a limit you can’t be told about isn’t a limit, it’s a wall. Its first job is the handful of asks that still come back with a silent “found nothing” instead of a reason.
The worst defect of the week was caught before any teammate met it. Asking to REMOVE a rule that forwards an old web address to a new one showed the confirmation for ADDING one — “this would permanently redirect…”, with the button labelled publish — and confirming performed the removal anyway. On the online-booking link, that is a booking funnel switched off behind a button that said the opposite. An independent reviewer found it. The fix deletes the duplicated machinery rather than patching it a third time.
Real bullets merged. Adding a bullet to a list — dot and all — came through its independent round with nothing left to break, including the honest part: where a list is built in a way the assistant can’t safely copy, it says so and names exactly where that text lives. It reaches the assistant with the next release step.
The audit’s worst finding is a silence. Ask to change a form label you can plainly see, and today the answer can be “scanned 11,706 facts… found nothing” — no reason, no honest no. Three places still do this: the insides of forms, page code, and the machine-readable data search engines read. Silence is worse than a spoken no, so killing these three is now the top of the work list.
We audited every “this is done” the project has claimed, and the biggest false one was ours. The speed claim was wrong three ways: the number measured the machinery’s own slice rather than the time a teammate actually waits; the test runs behind it pushed straight to the live branch with no review; and the checks were red when both “passing” commits landed. Standing rule out of it — a speed number measures the whole wait or it is not one. The four minutes quoted on this page is the honest number.
- Aug 11
Work is underway on the step that actually publishes an approved change. Today an approved edit sits on the private preview until someone publishes it by hand — that’s why some of the first test-drive’s approved wording is still waiting. The records side is built and an independent adversary is attacking it now; wiring it into the assistant comes next. Putting anything onto a real client’s live site stays behind a separate step only Robert can take.
The no-dead-corners guarantee finished its eighth independent verification round and came back with two more honest reds. One stayed invisible until the round insisted on running the check on our shared build server instead of the reviewer’s own machine — where it turned out to be red every single time. The other: the new checks meant to stop a problem coming back were themselves written to spot particular spellings, so three versions of the same problem rode through green. Both were ruled the same day and the next fix is building.
The speed work was built and then independently attacked. The gain is real on the bench — the machine’s own share of a landed edit roughly halved — but that is only a slice of the four minutes a teammate actually waits, and the attack found two genuine flaws, including a wording change that used to land and now dead-ends behind a reply claiming it changed things it hadn’t. It goes back for fixes. Under a minute stays counted as failing until it is measured green.
Factory repair kept moving overnight: two more batches of fixes merged and a third landed green, waiting on one last check. The factory is the assembly line that builds new sites in this system — none of this work touches a live client site.
Undo now works across chats: a teammate asked to take back an earlier session’s change and the assistant couldn’t see it — by end of day the record-reading and take-it-back machinery landed, proven on that exact incident, with an independent adversary attacking it before it reaches the assistant’s hands. Operating change behind it: small capabilities now get built directly and verified by one sharp adversary, with the heavy multi-round verification reserved for the machinery that could actually lie.
The three design questions the verification rounds forced upstairs were all ruled tonight, and both fix builds went straight into the oven. The heart of the rulings: a fix must cure the whole family of a problem, not just the example that got caught — and every proof must be tied to the exact thing it proves, using the measurement’s own record of what it measured.
Verification news, both directions: new-page creation’s second round found more to fix — honest reds, next cure round in spec. The no-dead-corners guarantee’s rebuilt proof passed its build and entered its own verification round tonight.
The two open design calls on moving sections were resolved — the approach dissolved into the site’s own record system, so a move is a recorded rearrangement rather than surgery on the page. Move and reorder builds next; nothing on this list is waiting on a decision anymore.
The multi-part-ask fix went into the build oven: one request carrying several changes will get every part answered in its first reply, with one receipt and one undo for the whole ask.
Photo swapping and form editing finished their design pass and joined the build queue — photos wire up machinery that already exists; forms get named, swappable, and honestly routed.
The expert law shipped to the assistant: it now carries the whole ask — composing missing pieces itself, answering every part of a request in its first reply, and offering finished work to accept or veto, never homework.
First real teammate test-drive (Alie): her new popup line landed on the preview. Two fixes came out of it — every part of a multi-part ask will get its own answer in the first receipt (fix in flight), and “remove this paragraph” jumped the queue with its design finished the same day.
Robert ruled: the under-a-minute speed requirement is part of “done” — full stop.
First cold-user test PASSED, 13 of 13 checks: a person who had never seen the assistant edited the site in their own words, and undo restored the original exactly.
Cleanup from that test landed: made-up clinic dates a test had typed were removed. They only ever existed on our own private preview — never on any client’s live site. Three plain-language rules were also added to how the assistant replies.
Creating new pages finished building; its second independent verification round runs today.
The “no dead corners” guarantee: verification found the proving test too weak — the test is being rebuilt, then re-run.
Factory repair started (first batch of fixes out for review). The factory is the assembly line for NEW sites in this system — no live client site is touched by this work.