AI

Operating agentically: an operating model (solo studio edition)

This one is the glue.

I’ve already written the short Notes — prompts as job descriptions, humans on the dangerous verbs, vibe versus assisted, and the identity piece about still feeling like a developer. This is how those pieces run together when I’m not writing a cute 900-word slice. This is what it looks like when I’m running a week.

It’s longer than my usual Notes on purpose. Treat it as a field manual you can steal from, not a vendor OS brochure, and definitely not a claim that I scaled AI for two hundred engineers. I didn’t. It’s a solo studio: one adult human in the room, with named agents that help me move faster without turning me into a spectator holding a fancy transcript. “Solo” doesn’t mean “no help.” It means I’m still the one who decides what help means.

The loud proof is my semi-automated job search tracker — SQLite, CLI verbs, agents in the loop, me on apply or no-apply. Product bench is the quieter mirror of the same manners, though the stakes are different some days.

Read the sections in order if you want the spine, or don’t. Your call. But the whole thing rests on one idea: instruments, not vibes. That includes the instrument that decides who gets to say “ship.” Some people call that instrument my brain. Some days it just feels like another tool on the bench.

Roles and routing

Before I care which model is smartest this week, I ask what phase we’re in — discover, build, verify, document, ship. I ask what a wrong answer would cost, who’s allowed to touch the outside world, and what “done” looks like in one plain sentence. If I can’t answer those, it isn’t time to build yet. It’s time to go think, talk to a human, talk to an agent, or talk to myself until I actually can.

I use a roster, and I’ll be honest — it still makes some engineers flinch, because it sounds like a startup org chart stickered onto a chat window. Fair. It can become that if I’m careless. For me the names are mnemonics, not personalities I’m inventing feelings for (except for fun, which I will absolutely do in the next paragraph).

Steve routes. Jules builds. Andi handles the chore-shaped work — UI, technical docs, cleaning up after Jules gets bored and skips writing a decorator. Renee hunts defects and is basically She Who Must Be Obeyed when something breaks in a stupid way; she doesn’t fix it, she tells Jules and Andi what they’d better fix before Trish finds out. Spoiler: I always find out. I read the diffs before I say yes to the commit. That’s in the house rules.

Page writes things down, does research, and makes sure the technical docs actually match the user docs so the whole project sings in one key. Tesla owns the sharp edges — sync, auth, SQLite, the stuff that bites. Drew fact-checks sourced claims whenever a Note wanders into models-and-incidents territory.

Routing is the control. It’s conducting the orchestra rather than hiring one enormous soloist who claims to do everything. One mega-agent that “does everything” is how you end up with confident nonsense and a wide blast radius. Narrow jobs are what make review possible, and review is the security feature that still fits inside one human brain.

Lane card — steal this shape. A lane needs a role stated in one line (“you research and cite; you do not ship”), the inputs it’s working from, where the output lands and what “good” looks like, an explicit list of non-goals, and a clear rule for when to stop and ask instead of inventing. That’s a prompt written as a contract, and it’s the same craft as writing a crisp ticket. Fancy temperature knobs don’t rescue a vague ask — use them if you want, but they only tune a response. They don’t fix a bad prompt.

Control surface versus specialist surface

Cursor is my default kitchen — named roster, files, terminal, MCP, house rules as first-class routing. Antigravity gets the invite when the hour calls for parallel agent-first coding and I don’t need the whole bench in one window. A shinier manager surface doesn’t earn unsupervised access on its own, and it doesn’t automatically win — I choose whichever tool fits the moment I’m actually in.

Job-tracker weeks stay in Cursor on purpose, because the signals markdown, CLI, and chat history all need to live in one place, or routing and review split across windows and I lose the thread. Antigravity gets the bounded coding spikes instead — parallel Jules-and-Andi work on a product slice, on days when life admin isn’t the priority.

Handoff protocols

Handoffs are contracts, not chat scroll folklore. If the “spec” only lives in a transcript, you’re vibe-adjacent no matter how many agents you’ve named. A written intent, even an ugly one, gives you something to point back to. Chat is a workbench, and workbenches get cleaned off — they’re not the system of record.

A useful handoff says who owns which lane (Steve routes Andi a chore; Steve doesn’t ask Andi to invent a sync protocol), where the artifact lands, what “done” means in a sentence, what’s explicitly out of scope, and when to stop and escalate. Escalate instead of inventing — a UI chore doesn’t invent sync protocols, a competitive scan doesn’t invent a product launch, and research doesn’t quietly become a production commit. Blast radius is a habit you build, not a rule you post once.

Parallel work is fine when the rails don’t overlap. It becomes theater the moment three agents all think they own the same file, and “done” just means whoever finished typing last.

Here’s the difference in practice. A bad handoff says “go find jobs and update the board” — no lane, no guardrails, no clear directions, no non-goals, no definition of what constitutes done, and the agent invents its own status while future-me argues with a mystery ledger because I wasn’t controlling the work. A good handoff says: opportunity-scout lane, scan these sources, write the shortlist to a dated markdown file, don’t touch the tracker database, don’t mark anything as applied, escalate if the geography or title is ambiguous, and “done” means a set of pursue/maybe/skip bullets with links. A separate lane, or a human minute, handles the actual database update. Clear verb, clear blast radius.

Artifact handoffs beat “keep going in the same thread until it feels done.” Files can be diffed. CLI verbs can be re-run. Markdown can be argued with. Feelings make terrible regression tests.

Quality checkpoints and gates

I insist on documented gates because agents are excellent at helping me skip steps I should be taking — and that’s not a moral failing on their part, it’s their job description. Momentum is what they’re for. Brakes are what I’m for, and writing the gates down is how I remember to actually use them.

The spine I run on: discovery comes before dig — a sticky note of what needs to change before repo surgery, a fit note before I build an entire application package. Confirm before mutate — anything that rewrites the world outside the editor needs an explicit yes from me. Humans stay on the dangerous verbs: push, publish, secrets, dependency bumps, apply or no-apply, “ship it.” Hygiene comes before features the moment something graduates from a sticky note to a real repo — that’s not a vibe, it’s a door I leave open for the version of me who shows up later. And I hold onto diff literacy: if I can’t explain a change, it doesn’t leave the building. Named agents don’t get me out of reviewing the work — they make review possible, because the work arrives in a shape I can actually inspect.

I also keep parking lots separate from product homes. Ideation lives in a research hub. Anything that’s earned a real thesis gets its own space. Same idea applies to the job search — signals and research notes aren’t the ledger of truth.

And one more, learned the hard way: back up before you scrub. If something sensitive has to leave a file, copy it somewhere that won’t sync, then scrub. Never delete the only copy and hope you remember the console later.

A pretty roster with none of this is just theater. Real security is quieter — and it’s also how I keep my own self-respect on the days Jules lands a thousand lines and my entire job is saying “not yet” to three of them.

Worked example: the job-search tracker

This isn’t one of my public products, so it’s worth explaining why it’s the centerpiece here. It came out of one afternoon where things were slipping through the cracks. I wasn’t keeping a spreadsheet honest, I was relying on memory for where I’d applied and when, and I was doing research by hand in a handful of places when automation could scan several vectors at once — role, company, sentiment, fit — and hand me back a “pursue, maybe, skip” with the reasoning attached.

So I built the thing. I wanted a dashboard I could actually look at, and reports that wouldn’t let me gaslight myself into thinking I hadn’t checked anything all week. The same agents could take a base resume and help draft an honest, job-specific package — one that led with the multi-million-dollar project I shepherded instead of the test framework I happened to build. I wanted scouting on a schedule, and I wanted to know at a glance how long ago I’d applied and whether I’d heard back.

Mapped onto the operating model: scouting and normalizing opportunities is agent-shaped work — research lanes, shortlists, watchlists. Updates go through CLI verbs into a local ledger, not vibes that might make it into a spreadsheet that lies to me by Friday. Fit notes and drafts can be agent-assisted and parked for reference. Deduplication happens before I ever see the list, so I’m not staring at the same posting twice. Apply or no-apply stays mine. That’s a dangerous verb with a career attached to it.

The gates that matter here are mostly emotional ones. I review the scouting report and approve or override the suggested disposition before anything hits the tracker — I might chase something under the compensation bar if it’s close and interesting, or say no to a “pursue” the agents liked. Before we build a resume package, I ask for a deeper dive on the company and role. Before I apply, I verify the package actually matches what we discussed, then run an agentic ATS-and-recruiter-style review so I’m not burning an application on theater. Everything lands in one interface so I can see current state, what’s waiting on me, and what got skipped or parked — daily and weekly reports exist for the same reason.

The gate that matters most, emotionally, is the simplest one: agents can scout and draft, but they don’t get to decide that I applied. The board exists so my own feelings can’t gaslight the week.

Same manners apply on product code. Jules still doesn’t get unsupervised push on Synesis, Phronesis, or LiveBytes drafts. Tesla still owns the edges that bite. This operating model isn’t “agents for code only” — it’s routing and blast-radius awareness, wherever the work happens to live.

Smoke tests and regression

Hiring managers, and future me, deserve a better answer than “the demo still looked fine.” If I change an agent’s brief, a model, a routing rule, or a CLI contract, I need a way to notice if the work quietly got worse. Not frontier-lab eval theater — solo-studio smoke tests, boring on purpose.

Rot on this bench looks like scope creep that quietly turns research language into ship language, gates skipped and dressed up as momentum, diffs I can’t explain in one pass, confident-but-wrong citations (which is exactly why Drew exists), or hygiene skipped because the afternoon felt lucky. On the job tracker specifically, it looks like wrong status verbs, invented applications, ledger drift away from what’s actually on disk, or “pursue” language that somehow became “I applied.”

After any prompt, model, or roster change, I run a short checklist rather than trust the vibe. I re-run the same narrow scout query from last week and check the shortlist still lands in the expected shape, with no stray database writes. I sample one harmless read path and one write path, to make sure nothing has started freelancing with SQL. I confirm apply/no-apply still can’t happen without me. I check that a fit note still calls out gaps instead of cheerleading, and that the size of the output roughly matches the size of the ask — if it rewrote the world, I stop. On the general bench, I re-run a known chore and compare it to a prior good run, deliberately tempt a lane to invent outside its job and see if it stops to ask, and check that sourced claims still survive a Drew-shaped fact check. And I make sure the confirm-before-mutate prompt still interrupts me — if it got quietly easier to blow past, that’s a regression, not a convenience.

I treat model, prompt, and roster changes like dependency bumps: confirm, sample on a small blast radius, then widen. I don’t swap the whole kitchen and find out Friday that the ledger lies and nobody picked up the groceries.

If I had to say the eval posture out loud, in hiring-loop language, it’s this: I know the assistants didn’t quietly degrade the craft, because I re-run a short checklist on the same instruments every time I change the assistants.

Accountability and the author of record

Assistance isn’t abdication. Feeling like a developer — or like the author of my own job search, or my own Notes — was never about typing every line. It’s about owning the change.

Some verbs stay dangerous no matter how good the roster gets: push, publish, secrets handling, dependency bumps, apply or no-apply, “ship it,” and creating a new repo or graduating a parking lot into a real product home. My test for whether I still get to call myself the author of record is simple. I chose the problem and the non-goals. I understand the change before it leaves the building. I own the dangerous verbs. And I could still defend the work if the model vendor vanished tomorrow. If those are true, the roster is a force multiplier. If they’re false, I’m a spectator holding a fancy transcript, and the imposter voice loves that job description.

The ego story matters more than people admit. “The AI built it” and “I built it with help” sound similar right up until the 1am debug session, when one version of that story hands you a mystery and the other hands you a known tradeoff. If you can’t claim the work, you can’t debug it either — and that includes debugging why you applied to a role that was a splinter you should have skipped.

So here’s what I actually tell hiring loops, and myself: yes, I use agentic tooling heavily. No, that doesn’t make the repo, or the job ledger, a mystery to me. Yes, I’ll talk about tradeoffs and failure modes and what I refused to automate. No, I will not pretend my process is vibes and a system prompt. The tracker exists because I wanted instruments when the process hurt, not because I checked out. Same posture in the product repos.

Being the one in charge doesn’t mean doing everything myself. When I’m tempted to anyway, I remind myself the instruments exist to absorb the repeatable, time-consuming work — so I can keep my focus wide, instead of living under every rock, pulling weeds.

What this is not

This isn’t a universal agent OS, and Steve, Jules, and Andi aren’t the One True Org Chart — they’re my mnemonics. Steal the lane idea; don’t cosplay my cast list if different names fit your brain better. This also isn’t “I scaled AI for two hundred engineers.” I didn’t. It’s a solo studio, with one adult on the dangerous verbs. It isn’t unsupervised-autonomy cosplay — a shinier agent manager doesn’t earn unsupervised push. It isn’t “every life-admin task should be agentic” — grocery lists can stay boring; I’m not automating a one-off sixty-second task just because I technically can. And it isn’t vibe coding as a shipping methodology for secrets, sync, money, or anything other people depend on. A throwaway Saturday widget can live dangerously. A career ledger and a mail sync path cannot.

What it is: instruments you can run with one adult in the room. Speed from the roster. Authorship from the verbs I keep for myself.

Closing

I get coverage from a team that doesn’t sleep. I keep authorship by keeping the verbs that matter. The short Notes were the parts — this was the operating model: roles, handoffs, gates, a worked ledger that keeps me honest, a smoke-test checklist for when the assistants change, and a hard line around what I’m not selling.

Instruments, not vibes. That line was originally about products, and about the Organon workbench. It’s also about how the work actually gets there, and who’s allowed to say it can leave.

~ Trish

← All notes