Home/Thoughts
Thoughts

$1.57 Told Him Nothing

When your API cost per interaction rounds to zero, you're overengineering your pricing model.

We were mid-call. A builder I'm coaching is walking me through the billing logic he'd been architecting for weeks. He wanted to charge users based on token consumption - input tokens, output tokens, different rates for each, a variable hard limit tied to the user's own configuration, and a monthly reconciliation at the end. He'd even pulled comps from Gemini trying to figure out the pricing.

And I'm sitting there watching him describe this complicated system, thinking: does this guy know what his per-interaction cost is right now?

So I asked him. He didn't know exactly. He pulled up his OpenAI usage dashboard. Total spend for the day: $1.57.

I sent a message to my own assistant while we were talking. Refreshed the dashboard. Still $1.57. The cost of a single interaction didn't even move the number.

That $1.57 showed me his costs were a rounding error - and it should have shown him the same.

The Token Pricing Trap

Here's what happens when you're building a product on top of OpenAI's Assistants API. You've got input tokens, output tokens, context window sizes, knowledge base population, and the fact that different users will send wildly different message lengths. You can run all the scenarios you want - long prompts, short prompts, heavy context, light context - but until you have real usage data at scale, every number you put in a spreadsheet is a guess dressed up as a model.

The guy I was coaching had identified this problem himself. He pointed out, correctly, that charging per "AI interaction" was already flawed because one interaction might consume 1,000 tokens and another might consume 10,000. If you charge a flat rate per call, you're penalizing light users and subsidizing heavy ones. And he wasn't wrong about that.

But then he went the other direction - toward a system so granular that it would require users to understand tokenomics, set their own hard limits, manage their own billing thresholds, and essentially co-own a pricing problem that should be invisible to them.

That's equally broken - just in a different direction.

What .57 Means

When your OpenAI bill for the day is $1.57 and a single message doesn't even register as a change on the dashboard, you're operating in a cost range where the unit economics almost don't matter - as long as you charge a sane flat rate with healthy margin built in.

GPT-4 Turbo runs at roughly $5-$10 per million input tokens and $15-$30 per million output tokens depending on the provider. At those rates, a typical back-and-forth conversation - say, a few hundred words in, a few hundred words out - is costing you fractions of a cent. Maybe half a cent. Maybe less.

If you're charging users $7 per thousand messages the way some competitors in the chatbot space do, and your cost per message is $0.005 or lower, those margins are fine. They are the business.

Ask yourself two things: what does this tool cost to deliver, and what is it worth to the user? Price from the value side, build in enough margin to cover the variance in usage, and stop treating your cost basis like it's a precision instrument when you're working in sub-cent territory.

The Overage System He Built - And Why It Almost Worked

To be fair, the guy wasn't completely off base. He'd built something reasonable for handling usage spikes: auto-bill users every $10 as they consume AI credits, let them set a hard limit on how much they're willing to spend, and if they hit the hard limit, shut the AI off until they raise it. That's a smart pattern. That's essentially how Stripe handles metered billing, and it's a proven model.

The underlying unit was the problem. He was trying to make the billing unit "tokens" when it should have been something simpler: a message credit, a session, or a monthly flat rate with an overage toggle.

I made him one request: don't make users put their card in twice. If someone gave you their payment information during the free trial, you already have it. Store it. Use it. The user experience of coming back to an app to pay a surprise bill you didn't understand is how you kill retention.

Auto-bill and notify them it happened. Give them controls. But don't make them re-enter a card they already gave you. That's where you lose customers.

Free Download: 7-Figure Offer Builder

Drop your email and get instant access.

By entering your email you agree to receive daily emails from Alex Berman and can unsubscribe at any time.

You're in! Here's your download:

Access Now →

Flat Rate with a Ceiling Beats Variable with Complexity

There's a competitor they were referencing - a chatbot platform that charges $7 for 1,000 message credits per month, and $99 for 10,000. Is that company making money? Probably. If they're not, the product doesn't exist at scale.

But more importantly: users understand it. One thousand messages for seven bucks. Ten thousand for ninety-nine. You don't have to explain what a token is. A pricing calculator becomes unnecessary. You don't have to explain why their bill was $12.47 one month and $18.23 the next.

Flat rate, with an overage cap, billed automatically - that's the model. It's not the most precise model. You will have some users who cost you more than you collect from them in a given month. You will have many more who cost you almost nothing. That's how SaaS works. The value is the outcome your product delivers.

If you want to get the pricing right, run this exercise: What does a power user cost you per month in API fees if they really hammer the product? Double that. Is your flat rate still above that number? If yes, ship it. If no, raise the rate or add a usage tier. You can make that call in ten minutes. You don't need a token calculator.

Pricing Should Create Confidence, Not Anxiety

The other piece that was buried in this conversation is user psychology. When you charge variable rates tied to usage the user can't easily predict, you create anxiety. People start rationing their usage of your product. They don't send the follow-up question because they're not sure what it's going to cost. They don't add more documents to their knowledge base because they're worried it'll inflate their bill.

That anxiety is the enemy of adoption. And every dollar of revenue you think you're protecting by charging the heavy users more - you're losing multiples of that in usage, retention, and word-of-mouth from users who never hit their stride because your pricing made them nervous.

A flat rate, or a simple credit system with a predictable ceiling, removes that anxiety. Users know what they're getting. They use the product freely, get value from it, and they stick around long enough to upgrade. That's the flywheel.

This is the same principle I keep coming back to with agency owners who are undercharging for their services: the number on your invoice shapes how seriously people take the engagement. A $350 engagement gets treated like a $350 engagement. A $3,000 engagement gets treated like a $3,000 engagement. The pricing itself is part of the product experience. If your pricing feels cheap and confusing, your product feels cheap and confusing.

Build the Simple Thing First

By the end of the call, we knew what to ship:

That's a product. You can ship that and explain it to a user in one sentence. Once you have usage data showing you what each active user costs, you can adjust it.

The elaborate token-based pricing model? That's a v3 problem, assuming v3 ever needs to exist. Right now, your cost is $1.57 a day. Build accordingly.

Need Targeted Leads?

Search unlimited B2B contacts by title, industry, location, and company size. Export to CSV instantly. $149/month, free to try.

Try the Lead Database →

The Lesson That Applies Beyond SaaS Billing

Founders spend weeks on a precision problem that doesn't need precision yet, while the work that blocks revenue sits untouched.

If you're early-stage and your OpenAI bill is $1.57 a day, your pricing model should take you an afternoon to design, not weeks. Every hour you spend tweaking per-token margins is an hour not spent building the feature that gets your first hundred users, writing the cold outreach that books your first ten demos, or closing the deal that tells you whether anyone wants what you're building.

Speaking of which - if you're still figuring out how to build the outbound system that fills your pipeline while your product is being built, grab the cold email scripts here. Five templates that have booked meetings. Tested and ready to send.

And if you're at the stage where you need more than scripts - where you need to look at your whole go-to-market motion and figure out what's broken - Galadon Gold is where that conversation happens with me and a community of people who are in it right now, not people who were in it ten years ago.

Back to the guy I was coaching: he's going to have the core of this built out and tested within a week. The pricing model went from a multi-week engineering problem to a one-afternoon design decision. Not because I'm a genius. $1.57 showed him his costs were negligible - sometimes you just need someone to point at the number.

Your cost is almost nothing. Charge accordingly - meaning charge enough, with enough margin, that you can weather variance and still make money. Then go sell it. Stop optimizing and start selling.

If you need a better system for finding the right people to sell to - whether you're doing outbound for your SaaS or for an agency - check out ScraperCity's B2B lead database and its suite of scrapers alongside tools like Apollo and Clay. Build your list, send the emails, and go close. The pricing conversation is secondary to the pipeline conversation - always.

Ready to Book More Meetings?

Get the exact scripts, templates, and frameworks Alex uses across all his companies.

By entering your email you agree to receive daily emails from Alex Berman and can unsubscribe at any time.

You're in! Here's your download:

Access Now →