- Noah Jacobs Blog
- Posts
- On Engineering Away Errors
On Engineering Away Errors
Make errors impossible so you don't have to handle them.

2026.10.11
CLXXIII
[Rewriting the Whale; Deleting Unnecessary Abstractions; Kill Your Darlings; Telling Your User to do Things that Won't Work; Vacuums with Errors; LLMs to Promote Design Level Thinking; Plateau Smashing pt Infinity]
Thesis: Making errors impossible is preferable to handling them.
[Rewriting the Whale]
We're once again going to do a quite big rewrite of WhiteWhale.
This is funny, as we decided this within 10 days of crossing a quite significant revenue milestone and adding 25% MRR in about 2 months, slaughtering a plateau we had been stuck at.
It's also funny because we just pushed one of our biggest features to date, adding contact data in for free, and obviously haven't been able to see the real impact of that because it was so recent.
The rewrite we're setting off on will accomplish a few things:
Greatly decrease the complexity of the UI / UX / system configuration
Refocus the value of the product to the output of ours that drives ROI for our customers
Decrease the complexity of our backend and make it a lot more simple to debug
Give us a quite nice, structural margin improvement
This is a lot of birds with one stone, and it might not work. Who knows.
Funnily enough, I think this is exactly what I said about the contact data when I started working on it a few weeks ago.
And, of course, just like then, there is no free lunch; I gave a lot of pros above, but there is a lot of disruption, too, some of which could be construed as pros but could also be cons given a poor execution:
It will be an expensive, complex rewrite
It will take some of the control away from the users
It will impact our pricing structure
But, these bets never really are all pure positive, even if they are ostensibly positive Expected Value.
The goal really is to rewrite the product in such a way that whole classes of errors and issues become impossible or exceedingly unlikely.
That's the theme we'll be exploring today: engineering away failure states.
I'll be somewhat vague about what we're actually changing until we release it, but I'll give enough to be useful.
[Deleting Unnecessary Abstractions]

this is from a youtube video by Casey Muratori on c++ but I think the sentiment applies everywhere: the default state makes sense and is accepted everywhere.
The notion of engineering away failure points works both for systems people see and for systems they don't; both the user's journey and the underlying code base, as much as possible, should be engineered to render failure modes impossible.
This implies something of a law, a sort of corollary that can be eeked out of it:
Failures and errors are an artifact of system design
As a comically simple example, an electric vehicle cannot have a problem with the combustion engine it doesn't have, and a rotary phone cannot have a problem with the touch screen it doesn't have.
Of course, both the electric car and a rotary phone have failure modes that the combustion vehicle and the smart phone don't have; this is not a value judgement on either, just an illustration that many of the problems something can have are downstream of it's design.
Then, obviously, somethings can have fewer failure points than other things while still accomplishing the same goal. That said, those things often times come with trade offs and reduced functionality.
A book, unlike an e reader, can't run out of battery and can have no hardware or software failures. But, a book needs a light to be read in the dark, while an e reader does not. Moreover, the books simplicity comes at the price of the extensibility of the e reader - with the e reader, you can read any book there is a digital copy of. That's not the case with a book - you'd need to buy another copy!
Given that ‘more’ adds more failure points, the super super interesting question for any task then becomes what is the simplest possible version of this thing with the fewest possible failure points that can still achieve my goal?
[Kill Your Darlings]
I think it's super hard to jump down to the least complex version of your product that does the things you need it to do, especially when you're maybe at the start not exactly sure what the things you need your product to do are.
It's a lot easier to give people a thing and watch them to see what features are getting in the way of them achieving the outcome they want.
From our own example, we've removed a lot of complexities that were not needed overtime. Two big ones were multiple signal types and children icp lists. I've written about both decisions in the past, so I won't belabor the point here, but, in short, both features made the product substantially more extensible without making it easier for our clients to do what they actually cared about: showing ROI.
Now, both features were cool and in theory let them do way more with the product, but we found users fixating on the configuration of these things at the expense of actually winning business with our tool!
Now, we've identified some more darlings that don't really add value to the user experience; a lot of these darlings have to do with account set up and how we actually price.
High level, we charge on the number of accounts that a user is 'tracking.' In theory, this is good, because it is tied to our internal cost, promotes up-sells, and let's the user tell us what companies they care most about.
This is at odds with our platform being as powerful as it can be, because really, they want US to tell them which accounts they and their team should focus on, and don't give a fuck about what firms they're 'tracking.'
This is an unnecessary abstraction that, at best, is a weak proxy for ROI. We think we'll fair a lot better if we remove it.
[Telling Your User to do Things that Won't Work]
We have this very powerful search feature in our software that let's a user enter a list of accounts or parameters and find the best ones on it.
There are a few problems with this. For one, this is still just more work in between them and the thing they actually want, which is to show roi.
Secondly, this invites quite unrealistic expectations: if we give you a system that allows you to check for all sub 50 employee law firms in the state of Arkansas that have a new CFO in the last 2 weeks... well, you might ask that question, even if it's incredibly difficult to answer and likely has very few actual positive examples!
In theory, these features are powerful, and we'll definitely keep an iteration of them, but it will be severely prioritized and more of a question of upfront configuration rather than a down the road addition to an instance.
And, we've already seen that even if we warn users that this will have few results, they still might attempt something like this. In other words, we set ourselves up for failure!
This would be important to figure out if this was the way that people actually got ROI; we're convinced that we can preserve the functionality that drives return, though, while cutting much of the extra nonsense.
[Vacuums with Errors]
I had to use a vacuum this weekend to help clean up a broken glass. The suction power on the thing was non existent.
The little monitor on the back had a warning - there was a nice graphic of the vacuum and how to open it up, saying airway blocked explicitly.
This is awesome. It told me what the issue was and how to fix it.
All products, especially software, should be like this. If something is wrong, it needs to tell you very clearly and also tell you how to fix it.
We've added some level of better handling to WhiteWhale to make these error states more clear, but if fixing it is still an art, rather than a science, we're fucked - no one has time for learning an art when they are trying to get roi.
Well, some people do, but not enough. Great software should make simple things easy to do and complex things possible to do.
Fixing an issue with an implementation should be a simple thing, and it should be easy to do it.
But, the more interesting question is how do you rebuild the system so that 99 out of 100 implementations don't even have anything that go wrong.
I think it's fantasy to say that you can build a useful software product with 0 errors possible, but it's certainly not fantasy to engineer as if that were the goal, as long as it doesn't blind you to being robust to the ones that do occur.
[LLMs to Promote Design Level Thinking]
My temperature on using LLMs to develop has recently changed slightly more favorably towards using them when I realized just how much more thorough I am myself. I think this can be good in that it enables you to do more system design level thinking, but it also does contain danger in that you can miss mistakes it makes.
To be clear, I'm still not advocating for 'vibe coding' or letting the LLM steer. I'm reflecting on how I'm starting to use LLMs to do the things I should be doing anyway but don't have time to do.
I noticed this first when I was doing work on a postgres query and the claude web browser downloaded and hosted a postgres in a sandbox environment, recreated a small part of my schema, and tested the query against it.
The point being, you really can get these things to do work for you that would take too much time on your own, and more so than just speeding up the happy path of development.
I had a query I needed optimized this week, so I gave claude the query, the plan, the relevant tables from the schema. Then, I had it write a script that tested 3 different improvement hypothesis and saved the performance for each to a text output file. Then, I had it merge the 2 successful hypothesis into one new query, and tried it against baseline and a 4th hypothesis. Finally, I had it check on 20 test cases to confirm the output was byte for byte the same against the original and that the query plan was consistently preferable.
Pleasantly, it was.
That said, they're still not perfect; I had it run some similarly robust queries to analyze where a 4 table queuing system was going wrong* and having a lot of unexpected failures. The analysis was great, but there is still the heinous occasional mistake, like throwing in a "drop table" command for a temp table that it just assumed had a name unique in my database.
Even with the sort of caution this necessitates, overall, I can see the allure of using these as tools to engineer & improve systems a lot more robustly and thoroughly than I, as a solo dev, may be able to afford to do on my own.
This is very different to me than being told to just vibe code and trust the system and the code or some such nonsense. And, it's more useful than how I was using it on the backend before, which was just oftentimes writing specific functions while I was dancing around the code base.
I still see myself using it like that, but in addition, levering it to more robustly do hard things seems like quite a good move.
*perhaps where it is going wrong is that it’s 4 tables!
If you enjoyed this post, I’m here every week about Bootstrapping SaaS, engineering, ai, philosophy, the West, etc.
[Plateau Smashing pt Infinity]
All that to say, it feels clear that we're about to hit another plateau with WhiteWhale after just jumping up over the last one.
Perhaps the contact data addition will help us grow faster, but I think we're still in a spot where we need it to be significantly easier for a user to not ask the system for impossible queries, and just really for the user to not feel like they need to play around with random settings to get the result they want.
At the end of the day, every single customer who is not getting clear roi will churn, sooner or later. Some already have emotionally churned but just haven't gotten around to cancelling the subscription yet.
Even cutting in half the number of customers who are focused on the wrong thing in our product would get us over the next impending plateau.
And delivering clear, consistent ROI for 9/10 customers rather than only some of them will get us over more than that.
Live Deeply,
