← Back to blog

Why I Don't Want an AI That Just Does Tasks

· 3 min read

My view on why 'AI as a tool' undersells what's actually possible, and why I'm building Nova as something closer to a colleague than a command line.

<p>There's a framing of AI that I think is quietly wrong, even though it's everywhere: AI as a tool. You give it a task, it does the task, you check the output, you move on. That's how most people use these systems, and it's fine as far as it goes. But it caps what you can build. A tool doesn't get better at your specific problems over time. A tool doesn't carry context from last week into today. A tool doesn't notice patterns across your projects that you haven't noticed yourself.</p><p>My view is that the real leverage isn't in AI answering questions faster. It's in AI accumulating context about a specific domain — your business, your codebase, your clients — and compounding on that over time. That's a completely different design problem than 'build a chatbot that's good at coding.' It means the system needs memory that persists, specialization that deepens, and enough autonomy to act without you re-explaining the same context every session.</p><p>This is why I keep coming back to multi-agent architecture rather than a single do-everything model. A general-purpose assistant is good at being average across everything. What I actually want is something closer to a small team: an agent that knows AIREP's multi-tenant branch structure cold, one that knows Find a Sign's marketplace logic, one that just handles orchestration between the others. Specialization isn't a nice-to-have in this model — it's the entire point. A generalist model answering ERP questions with generic best practices is worse than useless if it doesn't understand that branch-scoped data isolation is a hard constraint, not a suggestion.</p><p>There's also a temptation, once you have an AI system that can write code, to have it write all the code and step back. I don't think that's right either, at least not yet, and maybe not ever for the parts that matter. The value I bring isn't typing speed — it never was. It's judgment about what's worth building, what corners are safe to cut, and what constraints are non-negotiable. An AI system that removes my typing but keeps my judgment in the loop is useful. One that tries to remove the judgment too is dangerous, because judgment is exactly the thing that's hardest to verify from the outside until something's already gone wrong.</p><p>This shapes how I think about the self-improvement work on Nova specifically. Autonomous agents reviewing and refactoring their own code sounds impressive as a headline, but the actual design question is much less exciting and much more important: what's the boundary of what they're allowed to change without a human looking at it first? Self-improvement that skips that boundary isn't self-improvement, it's just unsupervised drift. Getting the boundary right matters more than getting the agents clever.</p><p>None of this is about AI replacing the parts of the job I actually like. It's about AI absorbing the parts that were always just overhead — the repetitive setup, the context-switching between projects, the manual status-checking across five things running in parallel. That's the actual case for the orchestration dashboard work: not 'AI does everything,' but 'I stop losing time to bookkeeping about my own work.'</p><p>I don't think this view is controversial exactly, but it's easy to lose sight of when the tooling around you is optimized for quick wins and demo-able chat interfaces. The boring version — memory, specialization, clear boundaries on autonomy — is less impressive to show off, and it's also the only version I think actually compounds.</p>

Comments

No comments yet — be the first!

Leave a comment

Comments are held for moderation before appearing.