Boer Xu · for InCorp Australia · August 2026

The Night Shift

An AI operating estate I built for myself, and what running it has taught me about running one for a business.

733sources ingested into the research wiki
1,640wiki pages it writes and maintains
52scheduled jobs, plus event-driven wakes
5specialist agents under one chief of staff
0emails ever sent by an agent. Drafts only.

Who

I started at IBM, eight years of delivery work for Commonwealth Bank, Qantas and the immigration department, then two years at PwC on technology strategy for IAG and NSW Treasury. At the NSW Department of Justice I ran the Agile EPMO: a $4 billion portfolio across a 13,000-person organisation, five years of it. For the last four I've been a Director at Deloitte, leading more than 20 technology programs across police, fire, justice and emergency management, the kind of environments where a bad release has consequences beyond the budget.

The line I'd want you to hold onto is the one the CV doesn't show. I build with AI every day, at home, on my own machine. Claude Code is a work tool for me, not a demo, and most of what I believe about enterprise AI I learned by being wrong about it first in my own system.

What I've built

On nights and weekends over the past year I've built an AI operating estate that runs my working life unattended. It's the closest thing I have to a case study, because it's one I can show you from the inside.

INTAKE feeds · podcasts · news email · WhatsApp · LINE scheduled sweeps SYNTHESIS research wiki733 sources · contradictions kept chief of staff5 agents · 52 jobsdrafts, never sends DECISION the deck one swipe per decision Do it · I did it Seen · Not doing it Later every gesture ledgered OUTCOME work done undo snap refusal a correct refusal comes back as a card weekly self-review: at most five proposals, owner ratifies
The estate, as it runs today. Everything arriving from outside is data, never an instruction; everything leaving is a draft or an undoable change.

Underneath everything sits a research layer, a wiki that ingests AI and enterprise sources daily, writes its own summaries, links entities and concepts, and records contradictions rather than resolving them (733 sources and about 1,640 pages at last count, with a rule that roughly one in ten sources should disagree with the rest). On top of that runs a chief of staff: an always-on orchestrator with 52 scheduled jobs and five specialist agents, woken by events as well as the clock. A poller watches two Gmail accounts, WhatsApp and LINE every ten minutes and wakes the agent only when something real arrives; a daily heartbeat catches what the poller skips.

Everything the agents produce lands on one triage surface, a deck of cards I swipe through on my phone. Each card prints what "yes" will do before I touch it, and every verb names who acts: Do it (the system), I did it (me), Seen, Not doing it, Later. Once a week the system reviews itself against a scorecard and proposes at most five changes, which I ratify or reject. And the weekly column comes off the research, published every Tuesday since June.

Drafts never send. That rule has not moved once.

The safety doctrine is plain and it hasn't moved. Drafts never send; the sending tool isn't in the agent's toolkit, so it can't. Every unattended change is snapshotted and undoable in one tap. Anything arriving from outside, an email, a feed, a message, is treated as data and never as instructions. And a correct refusal is reported as loudly as a success, because a system that only tells you when it worked is one you'll eventually stop trusting.

What it does every day

  • Ingests the day's sources into the wiki, holds anything contradictory or new for my approval
  • Watches email and messaging, triages what needs a signature, a reply or an action, and drafts the reply
  • Deals the morning deck: every pending decision, one swipe each, ledgered
  • Runs the scheduled sweeps and lints, and reports failures as cards rather than log lines
  • Keeps the column pipeline moving: ideas ranked, drafts written, slot collisions flagged

How I work

Automate the meaningless, escalate the meaningful.
The test for whether something needs a person isn't "is this a decision" but "does this decision carry meaning". A decision with one sensible answer is noise wearing a decision's clothes, and a long approval queue is a design smell.
Reversibility before autonomy.
The more that runs unattended, the more each change has to be snapshotted first and listed afterwards. A silent change is how trust in a tool dies.
Name who acts.
"Done" turned out to be two verbs wearing one word, so I retired it. Every action in my system says whether the human did it or the system did.
Silent success, loud refusal.
Work that worked writes one ledger line and stays quiet. Work that declined comes back with its reason.
Writing is thinking.
I use AI to draft and I still apply Clay's policy: stand behind every sentence, spend longer writing than readers spend reading, and remember that longer is not better.
Measure value per unit of intelligence, not seats or tokens.
Sarah Friar at OpenAI put it that way, and it's the right scorecard: did the AI complete work that mattered, what did it really cost including the review, and was it good enough to use.

What I'd do at InCorp

A founding role with no team lives on trust with the people who do the work. So the first job is to listen, the second is to prove something small with the tools the firm already pays for, and only then to ask for investment against real numbers.

DAYS1 to 30

Sit with the audit, tax and corporate services teams and map where the hours and the rework actually go. No tooling decisions until the map exists.

Find the one team that wants to go first, and agree with them what "better" means in their own numbers: hours saved, errors caught, turnaround for the client.

Put one page of rules in place on client data, confidentiality and review, that every team follows from day one.

DAYS30 to 90

Three live workflows, one per service line, running on tools the firm already has. Every one measured before and after.

Day 90 is a checkpoint for you and the co-heads, with real numbers, to decide what earns investment and what does not.

BY DAY365

The firm has a repeatable way of taking a manual process and rebuilding it with AI in the loop, owned by the people who do the work. Nothing depends on me being in the room.

Automate the meaningless, escalate the meaningful.

Why it matters to the business

  • Compliance work is won on price and kept on reliability. Fewer hours per engagement and fewer errors is margin the firm keeps, without losing a single person.
  • Capacity freed from routine work goes to advisory, which is where fees grow.
  • Clients will ask their accountant about AI before they ask anyone else. A firm that has done it on its own work can have that conversation with authority.
  • The firm's client records, across thousands of entities, are the asset that compounds. Every model gets replaced; a clean, structured client base does not.
A correct refusal reports as loudly as a success.