← Rooted, the 60-second version

Rooted, the deep dive

Everything in the order it actually happened, including the detours.

Chapter 1 · Early 2026

The itch

Fair warning

This is the rough version. The other page is the trailer; this is the director’s cut with the bloopers left in. Expect plain text, half-finished thoughts, a few spreadsheet screenshots and things I got wrong. Less polish, more truth.

I read self-help books, listened to podcasts and kept journals. The pattern was always the same: a good idea on Monday, forgotten by Thursday. I tried habit trackers and goal-by-habit setups. They all went quiet within a few weeks.

The one thing that worked was a person who reminded me. When someone checks in, I’m far more likely to act. I also noticed I’m already good at some things and weak at others, and I didn’t know which was which.

So I took personality tests. They were interesting, but they didn’t connect to anything. There was no single place that linked what I lack, what to do about it and which habits help.

The idea: build the test first. It started with four areas (finance, clarity and purpose, growth, human connection) and ended at six. I called them roots: strengthen your roots and nothing shakes you.

  • 1Self-Direction
  • 2Communication
  • 3Financial Clarity
  • 4Growth Openness
  • 5Social Connection
  • 6Purpose Clarity
The six roots the assessment scores.

Chapter 2 · April 2026

Testing it by hand

I wrote a questionnaire and tried versions on friends and family. About 20 people went through it, and each round made it better. I kept the final version at 16 questions.

Then I turned the questionnaire into a prompt and ran it in Claude, chosen because it felt like the more privacy-focused option. Repeating the same prompt turned into a Claude skill. That was my first real work with AI. Chats hit token limits, and a new chat lost all context, so I wrote a context document to carry it across.

People filled in a form and I produced a report, written as a letter. I scored each answer by hand with a rubric, found the top two gaps, quoted their own words back to them and kept it under 200 words. My test before sending: if the letter could apply to anyone, rewrite it.

What I found. Between 25 April and 4 May, 10 people filled in the feedback survey. They rated the experience 4.0 out of 5 on average. Eight said the letter “understood me accurately”, one said it was partly right and one said it felt too generic. All ten agreed the focus area it picked was right. People said it felt like a conversation, not a survey. A couple asked for a plan to improve their lowest scores, which became the next chapter.

A technical lesson. My context file grew past about 900 lines and the agent started skipping parts of it. That sent me to read about RAG and wiki-style documents for LLMs: split the knowledge up so the agent loads only what’s relevant. I still use that today.

Scoring rubric sheet: each dimension scored 1 to 5 with the phrases to listen for and red flags
The scoring rubric I used by hand: a label, key phrases and red flags for each score.
Step-by-step guide and template for writing the profile letter
The letter recipe. The test: if it could apply to anyone, rewrite it.

Chapter 3 · Late April 2026

From insight to daily tasks

The report gave a score of 1 to 5 for each root. A score isn’t useful on its own, so I needed something to do about it. The rule I settled on: one task a day, rotating the root by weakness, with the difficulty rising as the score rises.

I wrote about 20 tasks by hand, researched what makes a good one, then generated around 200. I reviewed every one and tuned many of them to land at 5 to 10 minutes. Energy varies day to day, so every task also got a floor version: a 2-minute fallback for bad days. Each one came with two end-of-day questions, one about what you did and one about how it felt.

I tracked it all in a Google Sheet and scored by hand. The first cohort was six people: me and five friends, starting 28 April. A task in the morning, a follow-up in the evening, for seven days. One person extended to day 20. Almost every day came back completed, but I was the one nudging every morning and evening. So the idea worked with a nudge. That’s why reminders became a core part of the app.

What I wrote down as wrong. Tasks felt repetitive when they only addressed gaps. People sometimes didn’t understand a task well enough to keep it or send it back. The low-effort floor tasks were sometimes so small they weren’t worth doing. And an end-of-day question makes no sense on a task about tomorrow.

Cohort tracker: seven days of tasks per person, mostly marked as fully completed
The first cohort tracker, seven days each. Names cropped out.
Task bank rows with the full task, a 2-minute floor, why it works and two end-of-day questions
The task bank: a full task, a 2-minute floor, why it works, two end-of-day questions.

Chapter 4 · May to July 2026

Privacy, then the web app

Some people hesitated: it’s a Google Form, why would I share my deepest secrets here? In this space, privacy matters most, so I took it seriously.

I built a web app. Names and answers are kept apart by default, on Supabase with row-level security, so even someone with database access can’t read another person’s data. Sign-in is through Clerk. The repo’s first commit was on 1 May and the web app was live by 18 May, about two weeks later. The rest of May and June went into hardening it, with July spent refining.

The reminder problem. Push notifications from a web app (PWA) didn’t work reliably. WhatsApp is how a lot of people in India talk, so I tried that. Meta’s verification and message templates took about a week of struggle, and the reminder still sent people back to a sign-in screen. They stopped returning.

The honest bit. Friends validated it early. Later I realised some of that was kindness, not a real signal. I added three assessment tracks (stuck, struggling or low, thriving) with different task sets for each, and a gentler path for very low scores.

How I built it. Claude Code wrote the PRD, the TRD and the specs, broken into milestones. Spec-driven, not vibe-coded. Then manual testing, then Playwright, and local didn’t always behave like live. Hosting on Vercel. I also set up a wiki and daily logs so the agents kept their context.

The tooling pain. I started on a Windows gaming PC. One editor was unstable, my Claude $20 plan ran out at peak hours, I tried cheaper models for cheap tokens, and some tools didn’t connect to each other. Eventually I moved to a multi-agent terminal setup.

Web app welcome screen

Sign up

Web app reminder preference screen

Reminder choice

The web app, version 2.

Chapter 5 · June to August 2026

The pivot to mobile

India doesn’t pay much for wellbeing services, so I aimed at the US. That means an iOS app.

First I wrapped the web app. It wasn’t smooth enough. I then expected moving from React to React Native to be mostly copy and paste. It wasn’t. The agents rewrote a lot, the design direction changed, and it took about a month to match what the web app already did. I should have started on mobile. The mobile repo’s first commit was 26 June, and the busiest month was August.

API sprawl. Pages made too many calls, repeated the paid-or-trial check and loaded slowly. I went through and trimmed them.

Apple said no four times. All four were fair.

  1. A missing subscription terms link (Guideline 3.1.2).
  2. Sign in with Apple (Guideline 4): I asked for a name that Apple had already given me. Clerk was receiving it and discarding it, so I fixed both the settings and the screen. That was 5 August.
  3. A problem with one of the subscription items, not the app itself.
  4. Two bugs on an iPad (Guideline 2.1): a “restart to update” prompt on launch, and a screen that wouldn’t scroll after registration. The first one was a gap in my setup that I’d never have guessed: the build sat in the review queue for nine days and quietly picked up my live updates.

What helped: proper screenshots, a screen recording of the whole flow, repeated testing, and GitHub CI for automated tests, because builds kept breaking while I was at my day job. I almost quit. Apple approved the app on 20 August.

The web app has since been discontinued, since there was nothing left worth attacking. Android is in closed beta.

Chapter 6 · September 2026

Version 2: one daily list

People liked it, but I didn’t feel done. Something was missing: I needed to know more about myself, and I wanted habits, without turning Rooted into yet another habit tracker.

I ran research with several agents, Kimi included. The verdict held: one coached task a day stays. But people wanted a say in what to work on. So I added goals, each linked to a root and broken into habits.

I also opened the task table and found 192 tasks that only asked someone to notice something, and none that helped them build a system. I rebuilt the task engine around that.

The home screen became one list of the day’s activities: the assessment-based task, goal tasks (daily, weekly or a checklist) and habits. Habits are capped at three, usually one primary and one secondary goal, so nobody gets spread thin.

Apple cleared this version in about 20 minutes. The earlier ones came back rejected after about an hour. Early signals are small but real: two written App Store reviews so far, both five stars, a week after launch.

Version 3 home: a path of days with only day 1 unlocked

v3: a path of days

Version 4 home with tasks and habits in one list

v4: one daily list

The home screen, before and after.

Chapter 7 · Throughout

Running the agents

Claude’s $20 plan was enough at first, with Opus and Sonnet. Then my automation outgrew it, so I mashed plans together. Terminal windows couldn’t talk to each other cleanly. I used BB for about two to three months for the custom agent setup, diffs and merging PRs, then added Codex, and moved tracking to GitHub issues so nothing depends on one tool.

The setup. A chief-of-staff orchestrator takes my brief and creates tickets, then hands each to a sub-orchestrator with its own worker, bug worker, senior and advisor. The sub-orchestrator plans, the senior checks the plan, the worker builds, and the senior reviews the code. After two revision rounds, it escalates to the advisor, a smarter model. I tuned it so small surgical changes skip the plan.

QA. Simulators and tests keep growing, and I read the tests here and there.

Design. I pull UI research through an MCP, then sketch in ASCII, then HTML. After two failed rounds I copy it to Figma, fix it by hand and hand it back. Roughly 80% agent, 20% Figma. I use skills for copy and other jobs, and older skills fall away as models improve.

Hardware. I went from Windows to a 16 GB, 512 GB MacBook. With several simulators, tests and Docker for the local database, that isn’t enough.

Memory. The context file, wiki, PRD and TRD were all fine. What was missing was a way for different models to know what each one had written and saved. I use agentmemory, an open-source tool, through a custom Mac menu-bar plugin that runs all the time. It distils memories into learnings, so before doing something an agent first checks what we did and how we solved it last time. I also used a keep-awake tool so the laptop didn’t sleep while agents ran; BB made that unnecessary.

AgentMemory menu bar showing healthy status, 790 sessions, 21,017 observations and 4,195 memories
6 October 2026: 790 sessions, 21,017 observations, 4,195 memories.

Across the web and iOS repos, as of 6 October 2026:

PRs merged
690+
tickets raised
380+
commits
3,300+

How a change ships

  • Orchestrator · Writes the plan
  • Senior · Approves the planPlan approved
  • Worker · Builds it
  • Senior · Reviews the codeChanges required
  • You · Check it yourself, then shipShipped

Green tests alone never counted as approval.

Chapter 8 · Today

What I’d do differently

Build mobile first. Don’t build the web app and then port it.

System design is still my gap. I don’t know what happens at a million concurrent users, and I’d rather say that than guess.

Directing agents. You don’t need to know every technical detail. You do need to understand systems and how specs work, and you need to read the plans. A spec you skim may be understood in a different direction by the model. Today’s AI is smart, and if you ask it to do something it usually will. But now and then it takes a shortcut or a hack, heads off the wrong way, or does something you never asked for.

Guardrails against that take time to build. You don’t get it right on day one. It’s a process: you learn, and the agent’s context keeps learning too.

That’s the long version. Back to the 60-second read.