01

It looks like a finance app.
It's a decision engine.

Years of tracking my own money, and no app ever fit — they look backward, forget everything older than a year or two, can't run the futures, and flatten a cross-border financial life into the wrong buckets. So I didn't design a budgeting screen. I built an engine: it reads every transaction for what it means, values everything I own across two countries, and simulates thousands of futures before it draws a chart. The interface is the last layer — the reasoning underneath is the product.

Product, design and architecture — end to endFinance Chief · my personal CFOBuilt by directing AI agents · a web app and an iPhone app
What it reads
Years of transactions · every account · two countries
What it runs
Thousands of simulated futures · a full retirement model
Status
My daily money tool · live demo on a fictional family

No app could hold the way I plan.

I plan obsessively: where the money goes now, what we own, when we move, when work becomes optional. For years the app I used tracked the past, and everything else lived in a long note on my phone and in my head.

I'd spent years designing BI and analytics tools at work. Finance Chief became the place I explored everything I couldn't build there, for the most demanding user I know.

This is not just a personal finance app — it's the map of my brain, out here, in real time. Constantly planning, strategizing, ideating, forecasting, now backed with actual logic and real-time numbers.Me, early in the build

Why I didn't just keep making do with the money app I already had:

I have a huge Apple note for a bunch of finance stuff that Rocket Money can't hold.Me, describing what came before

I defined the job
before a single screen.

The first decision was what the app is for. Two jobs: an honest picture of where the money goes today, and a forward answer to the long question — could we stop earning one day, and be fine? Anything that serves neither doesn't get built.

Not stingy. Not careless. Clear.

Today's half of the job, as four questions.

MandatoryWhat do we have to spend every month?
OptionalWhat is optional?
CuttableWhat could we cut?
Too muchWhat feels like too much?

The job the app exists for. A category only earns its place if it helps answer one of these.

When the builder started producing ever-finer spending categories, I stopped it and restated what the categories were for:

I'm trying to optimize my life. Not live stingily, but not overspend either. … So this is our end goal. Not endless categorization.Me, to the AI builder

One engine,
read in layers.

Finance Chief is four layers on one deterministic core: the ledger (what happened), the balance sheet (what we own), the forward model (what could happen), and a thin AI edge for the open questions. The same engine runs the web app, the in-browser what-if sliders and the phone, so a number can never disagree with itself across screens.

Code where it has to be right. AI only on the tail.

DataBank feeds and older archives, spliced into one historycode
MeaningEvery transaction classified by intentcode
ValuationPositions, metals and rates from live sources, freshness-stampedcode
Simulation5,000 simulated futures put odds on the goalcode
Open questionsWhat-ifs and explanations in plain languageAI
InterfaceThe last layer: it shows what the layers above decideddesign

Deterministic, instant and free at every layer that has to be right. AI only on the open-ended tail.

The ledger:
my life, read from transactions.

A transaction history knows more about a life than any budget: what's fixed and what's a choice, what quietly grew, what ended, what changed when income changed, and what each habit costs over years. More of the build went into reading that correctly than into anything else.

I had years of spending data spliced from old apps and live bank feeds. This is when I decided the spending analysis was the heart of the product:

Let's go all in on the spend analysis and insight generation idea. … Things that I can't possibly find — but you can.Me, to the AI builder

What the engine reads from a transaction history.

Where it actually goesEvery transaction read for intent, so wealth-building, transfers and paybacks never pose as spending.
What I can’t get out ofEvery obligation as a stream: active, paused or ended, priced at its last payment, with its full price history.
Habit or billFrequent spending is counted in visits and never called a subscription.
How much is too muchEach category read as a habit, an episode or too thin to judge. Episodes are listed, never averaged.
Who I’m really payingMessy bank descriptors resolved into real merchants. The engine proposes, I confirm.
When life changedTrips found from physical evidence, and what I protected or let go when income changed.
What it costs, long termEvery finding priced against the long-term goal, as a range from conservative to optimistic.
Whether to trust the numbersThe app audits itself: stale sources, shrinking syncs and blind spots are flagged, not hidden.

All of it computed in code from the ledger itself. No AI is involved in finding any of it.

It reads what money means, not just where it went.

One transaction, checked top to bottom.

A tested precedence chain, checked top to bottom, so nothing counts twice and nothing counts as spending that isn't.

An anomaly can't lie. Burn uses the median month, not the mean, so a one-off repair can't inflate it. Where a monthly rate would be a lie, the engine refuses to show one.

Who I'm actually paying. Bank data carries no merchant ID, only raw text that changes all the time. The engine resolved hundreds of raw names into the real merchants behind them. When it isn't sure, it asks me: same merchant, or keep separate?

The first time I looked at the raw ledger laid out month by month, uneven and unpredictable:

That's what real spending looks like. No real symmetry at all. It's wild. Random. Forces you to look at every single transaction.Me, reacting to my own data

Two sources, one history.

Archives from older apps, imported and de-duplicated
Live bank feed

Spliced exactly at the bank feed's earliest date: no overlap, no double count, every row tagged with where it came from.

Every insight is a claim
until it survives the guards.

A finding that sounds specific is easy to believe and easy to get wrong. So nothing reaches me until it's been checked against the ways transaction data lies: a provider switch that looks like a price rise, one merchant under many names, a winter bill compared to a summer one, a one-off mistaken for a habit, a change too small to matter.

Confirmed, reframed, or rejected.

A claim
“This bill went up”
Five guards
Substitution: did a provider just change?
Aggregation: is it one merchant under many names?
Seasonality: compare like months
One-off: an event, not a habit?
Significance: big enough to matter?
Outcome
Confirmed
Reframed
Rejected

An insight is a hypothesis until the evidence holds. A language model may only phrase what survives.

What comes through reads like a sentence, not a chart.

Part of every month is spoken for before I decide anything.Fixed versus elastic spending
I went out more often, and each visit got a little smaller.Frequency versus basket
When income changed: what I cut, what held, and what changed only because an event ended.What I protected and what I let go
This stream is paused, not ended. It has been quiet this long before.Pause, measured against its own history
These one-off payments are listed, not averaged into a monthly rate.Episodes stay episodes
If this habit holds, here is what it costs the long-term goal, as a range.Long-term cost

Illustrative shapes: no real figures

The engine can describe what I did. Whether a purchase was an impulse, or worth it, is a judgment it has no evidence for:

Impulse is mine to mark, you can't derive it.Me, drawing the line

The balance sheet:
everything we own, in two countries.

Bank and brokerage accounts sync live. Retirement accounts, stock grants, gold, homes and land in two countries are valued from live sources, each stamped with how fresh it is. Investments in India get their own rupee sleeve, so they aren't converted to dollars and back. Tax is modeled on published slabs, with sources and dates, not a flat guess.

Grouped by what I can reach, not by asset type.

You can reach thisCash, brokerage, anything you could use this month without a penalty.Spendable
Locked until each of you is old enoughRetirement accounts, unlocked per owner on a date the app derives, not one I typed.Real, later
Real, but you can't spend itHomes, land, gold. Valued live, counted in net worth, never mistaken for cash.Wealth, not money

The balance sheet is grouped by what I can actually reach, not by asset type.

Stock grants kept showing up as if they were salary. I set the rule that they are wealth to track, not money to spend:

Don't calculate vesting in the monthly expenses or income. … That is just for you to have knowledge about and to build in the background.Me, to the AI builder

Provenance on every figure.

HERSI told the app
LEDGERRead from the transactions
MEASUREDComputed from the data
SOURCEDFrom a cited outside source
MINEThe engine’s own assumption

Every figure says where it came from, so a guess can never pass as a fact.

The forward model:
life paths, ranked.

I keep several futures in play: different cities, a move home to India, a two-country detour abroad for my career, a job offer to weigh. Each is a life path with its own move, income and costs. The engine ranks them against one goal, runs thousands of simulated futures for the odds, and back-solves what each path needs. A signed offer replaces the forecast with real numbers on the path it belongs to.

The decision, not a dashboard.

1
Path Abest path now
~72
2
Path B
~61
3
Path C
~48

Illustrative: the mechanism, not real figures

Why the paths are something to re-point, not a prediction to defend:

Life is unplanned. What we're doing here is planning for the future as accurately as we can with the knowledge and plans and ideas we currently possess.Me, on what the model is for

A floor and a ceiling, never a single line.

Ceiling: if things go wellFloor: what you can count on

Money that might not come, like rent from a tenant who might not appear, rides in the ceiling only.

Could we stop earning one day,
and be fine?

The retirement model answers one question across decades. It splits life into phases computed from our dates of birth, runs every combination of floor and ceiling for income and markets, tests crashes at different moments, and moves one factor at a time to see which ones actually decide the outcome. The page is ordered as an argument: the answer first, then what it's made of, then what could break it.

Three phases, each a money event.

Phase oneThe earning yearsSalaries, saving, investing, a move. Starts today, not at some future milestone.
Phase twoWork becomes optionalThe portfolio carries the household. Any salary is a bonus, never the plan.
Phase threeLocked money unlocksEach partner's retirement accounts open on their own date.

Phase boundaries are computed from dates of birth, not chosen. Each one is a money event.

When the builder kept treating retirement as the day the salary stops, I explained what the goal actually is:

The whole point of [the retirement year] is that work is no longer necessary. … If I want to work after that, bonus.Me, to the AI builder

All four corners, honestly paired.

Floor growth
Ceiling growth
Floor income
Misses
Clears
Ceiling income
Misses
Clears

Four honest corners. Never one variable's floor paired with another's ceiling. Illustrative outcomes.

Two walks that disagree, shown together.

The income walk“It doesn't yield enough.”Asks what the portfolio pays out each year, and finds a gap.
The withdrawal walk“It lasts, and keeps growing.”Takes out only what we'd use, and lets the rest compound.

Same money, opposite verdicts. Both are shown, side by side. Showing only the comforting one would be flattery.

The model kept measuring how much the portfolio pays out. That wasn't the question:

It doesn't matter how much it's paying me. What matters is how much I'm using it. … I only want to take out what I need to use — the rest can stay and grow.Me, to the AI builder

Findings I couldn't have run in my head.

Timing beats size.

Crash in the first year of drawdown
Same crash, a decade into drawdown
A lost decade with no recovery

Same shock, very different cost, purely from timing. Illustrative proportions.

Returns decide the goal, not career. Holding returns fixed, the whole range of career outcomes moves the result about a third as much as the range of market returns. Both low-return corners miss; both high-return corners clear; career flips neither.

A shortfall that wasn't. Measured as yield, the plan showed a big gap. Measured as withdrawals, the money lasted and grew. The gap was an artifact of the framing, which is why both walks stay on the page.

Small rules with big effects. Chasing the best advertised deposit rate is worth almost nothing once deposit insurance caps the amount. Gold is the one sleeve that rose in a crash year, and it earns its place as diversification.

One finding I called before the model measured it: that over the earning years, what we keep putting in matters as much as what the market adds.

[Years] of income and contribution matter a lot more than just money sitting and compounding.Me, before the numbers came back

Why the retirement page reads top to bottom as a story, instead of a grid of cards:

This is about storytelling. … And this story is about the next [few decades] of my life.Me, on the page order

AI on the edge.
The phone in my hand.

There is one Ask. It answers only from a factual snapshot of my data, picks a chart view while the app supplies the real numbers, and can draft a new life path that stays a draft until I add it. What a conversation produced is kept, stamped with the date of the data it used. The chat itself isn't.

I'd been running my planning conversations in long AI chat sessions, and rereading them was useless:

I want only one Ask feature … artifacts and results from the chat saved. Not the chat itself.Me, to the AI builder

What the AI may and may not do.

The AI may
  • Answer from a factual snapshot of my data
  • Pick a chart view; the app fills in real numbers
  • Draft a new life path as a change to an existing one
  • Phrase an insight that already passed the guards
The AI may not
  • Compute an outcome
  • Invent an income curve or a figure
  • Change anything without my approval
  • Spend past a monthly cost cap

The model proposes a shape. The engine computes the worth. I decide.

One Ask, and nothing changes until I say so.

Ask“What if the move happens two years later?”
Draft path

A fourth path, drafted as a change to an existing one. Nothing is saved yet.

Add as a pathDiscard

The AI drafts; I decide. What the conversation produced is kept. The chat itself isn't.

The phone that replaced my old money app.

The phone is about the month: what I spent, against my own target and my own baseline. Deep planning stays on the big screen.

This month
Dining and groceries
Against my own target · illustrative
Takeout is having a big month. Your kitchen misses you.Illustrative nudge: wry about the spending, never about my circumstances
Tap a transaction to tag itEvery screen says how old its data is

The phone holds no keys and no ledger. It reads from my own machine, alerts only outside quiet hours, and never twice for the same thing.

What I asked the phone to be, and the one thing I wanted it to have that money apps never do:

The phone version is to basically mimic my use of Rocket Money — seeing my transactions, categories, spend trends, spend insights. … I want some humour, you know?Me, to the AI builder

Then I taught the builder
to stop guessing.

Finance Chief was built by directing AI coding agents. Getting code written was the easy part. The hard part was stopping a confident system from inventing things about my money and my future. Every correction went into a teachings file the same session, more than a hundred of them, each tied to the rule and the code that enforce it.

Partway through, I realized the corrections themselves were the real spec, and asked for every one to be written down:

With every prompt … I'm teaching you how my brain works.Me, to the AI builder

The ledger: every correction became part of the model.

What the engine did
My correction
The model now
Read a switch of phone carriers as a cancellation and a price rise
One bill ended and another began: same service, new provider
Check for a substitution first
Paired unrelated spending as if one replaced the other
“not every expense has a trailing effect”
Substitutions only between continuing bills
Called a gym "stopped" during a break
Some habits happen in phases
Paused, measured against its own longest silence
Showed a price I no longer pay
Plans change; the last payment is the price
Current price = last payment, history kept apart
Spread one-off payments into a monthly rate
“don’t spread it to make it monthly”
A monthly figure is a claim. None where it would lie
Called restaurants a subscription
“It’s not a recurring expenses. It’s a frequent expense.”
Frequent isn’t recurring
Built a "trip" out of restaurant visits at home
“noticing a pattern that is REAL. Not invented.”
Trips need physical evidence
Split one gym into two rows on the cut list
“If I can notice it - the engine needs to learn to notice it as well.”
Group by the real counterparty
Reported a number that was true but told a false story
A card credit was paying part of that bill
Net the credits; show the latest change

Each correction went into the project's teachings file the same session, and into the model as a rule with a test.

The forward model: every assumption had to be mine, or labeled.

What the model did
My correction
The rule now
Quoted one growth rate as if it were the answer
A base case and a best case, for every investment in stocks
Every growth figure shows both ends
Called it a "ceiling" while pairing the worst income with the best markets
Floor with floor, ceiling with ceiling: a real matrix
Never mix one variable’s floor with another’s ceiling
Used a central bank’s target as the inflation forecast
A target is a hope, not a forecast
Use the rate I plan for, and say whose number it is
Set a yield the portfolio had to hit
“Yield doesn’t need a target”
Take what we need; let the rest compound
Put stock vests into monthly income
“don’t calculate vesting in the monthly expenses or income”
Vesting is wealth, not take-home
Counted unvested stock twice
Once as a balance, once as future vests
Count each thing once
Let a missing price key quietly erase an asset from the forecast
A missing key is a reason to stop refreshing, not to forget
Never drop a value just because a source is down

Ask, don't presume.

The first spending findings the engine surfaced were confidently wrong: a switch of phone carriers read as a price rise, a winter energy bill compared to a summer one. So the work went back to the data before building more insights, and the engine learned to ask when the data can't settle something.

The moment it became the rule: the builder flagged one month's spending as unusually high, and I knew before it finished explaining that it had misread the data.

The moment you said [that number] was too high for the month, I figured you were wrong. I knew, I knew it that moment. … You should always ask me. That's our number one rule. Remember, don't presume anything.Me, to the AI builder

And the same idea, applied to the app's own sources of truth: the ledger outranks my memory.

Go by the real data you see. My numbers can be off — they're always an approximation based on memory.Me, to the AI builder

“The projections are for me — because I know the reality of it and I built it. An app didn't build the projection for me.”

Why every assumption in the model is mine, or labeled as someone else's

Built in weeks.
Used every day.

The first version was live a week after the first commit. It has replaced the money app I used for years. It runs on my own machine with my real accounts, and the phone reads from it. The public demo runs on a fully fictional family, behind a scrub, a leak sweep and an audit of what the site actually serves, so no real figure ever ships.

Built over~16 weeks
Commits400+
Engine logic~27k lines
Automated tests, all passing1,100+
Screens30 web pages + an iPhone app
The software's own scale, from its git history. No personal figure appears anywhere on this page.

How it came together.

The foundationWeeks 1–2

Years of archives imported, live accounts and assets, the first classifier, honest burn and runway, three life paths, simulated futures, the design system, the demo and its privacy gate.

The shell and the phoneWeeks 3–7

A personal always-on instance, the navigation, the first iPhone build, the cost governor for AI.

The data-layer resetWeeks 8–9

After the first findings came back wrong: income derived from transactions, merchant identity, verification guards, the obligation ledger, the teachings file.

Stop presumingWeeks 10–11

Frequent versus recurring, trips, one Ask with kept results, long-term cost on every finding.

The long viewWeek 12

The retirement model: phases, four corners, the two walks, crash tests, tax, gold by weight.

The phone takes overWeeks 13–14

The phone replaces my old money app: the month, tags, a spending target, wry nudges, alerts.

Money that comes backWeeks 15–16

Card credits, refunds by shape, the categorizer asking through Warden, a home screen that is the review.

From the project's git history. The ledger work clusters in the middle: most of its commits landed in three weeks.

Honest by design,
all the way down.

Burn uses the median month, so a spike can't lie. Every figure says where it came from and how fresh it is. Every insight survives five guards. Uncertain money rides in the ceiling only. Two models that disagree are both shown. And the app audits itself: stale sources, shrinking syncs and blind spots are flagged on a page of their own, instead of hidden.

5guards every insight must survive
4corners every long-term answer is shown at
0outcomes the AI is allowed to compute

From a calculator to a chief.

What I'd build nextThe move, and the judgment layer

Every data feed today is American, so the app would go blind the day we move. India's account-aggregator feeds are next, along with putting the cross-border tax and residency research I've already done into the app itself, as orientation for a real professional, never as advice. And a way to mark what was worth it, in my words.

What it left behindA way of building

Define the job before the screens. Fix the data before the insights. Treat every finding as a claim. Show the floor and the ceiling. Let code answer what must be exact, keep AI for the open questions, and ask instead of presuming.

No one else was going to build this for me —
so I did.

Not a budgeting app: a decision engine that reads what my money means, values everything we own, runs the futures we're weighing, and checks every finding before it believes it — with AI only on the tail. The BI tool I couldn't build at work, architected and shipped end to end, for the most demanding user I know.

How I describe it now:

It's a personal software I built to offload my thinking — the hundreds of thousands of thoughts and numbers and calculations that I run in my mind for my future. There was no app in the market that could do that for me.Me