Future of GTM + Product

Read time
7 min read
Published
October 6, 2026

The Cheapest Intelligence That Does the Job

Most GTM automation sends every step through the biggest model on the market, then wonders why the bill outgrows the pipeline. How we route each step to the cheapest thing that does it well: code first, small open-weight and fine-tuned models for language, a frontier model only for real judgment, open-source tools where they are safe, and a person for every customer-facing promise.

Outcome / Lower cost per finished task

Mindlyft / Service as a Software. The work, done. Delivered as software.

01

The bill nobody shows in the demo

A demo runs one call. A revenue team runs thousands. Every transcript, every inbound email, every ticket and every CRM sync is a model call, and a workflow that looks free on Tuesday afternoon costs real money by the end of the quarter.

The price spread is enormous. On public list prices this month, the top tier from Anthropic and OpenAI costs $10 per million input tokens and $50 per million output tokens. A small open-weight model on fast hosted inference, gpt-oss-20b on Groq, costs $0.075 and $0.30. For the same token, that is roughly 130 to 170 times cheaper.

That gap is the whole design problem. If every step goes to the most expensive model, the automation costs more than the work it replaced, and the first budget review switches it off.

02

Most GTM work is not a reasoning problem

Look at what actually happens after a sales call. Match the account. Find the open opportunity. Parse "end of next week" into a date. Check the field is not already set. Write the update. File the ticket in the right project. Draft the recap.

Only two of those steps need language understanding: reading what was promised, and drafting the recap. Everything else is lookup, matching and arithmetic, and code does those perfectly, instantly and for free.

So the first rule of every build is boring on purpose: if a step can be written as code, it is code. Matching on a phone number, a domain or a house number is a rule, not a prompt. In one build we pin street numbers in code so that 412 Elm can never be matched to 413 Elm, whatever a model thinks.

03

The ladder: the cheapest rung that clears the bar

Every step in a workflow gets routed to the cheapest thing that does it well enough, and escalates only when it has to:

Rules and code for lookup, matching, dates and dedupe.

A small or open-weight model for reading and classifying language: which kind of request is this, who promised what, by when.

A fine-tuned specialist when the same narrow job repeats thousands of times in your words.

A frontier model for the genuinely hard cases: an ambiguous email, a messy negotiation, a long thread.

A person for anything that commits your company to a customer.

The routing is not a guess. Each decision carries a confidence score. Below the floor, the item climbs a rung: a 0.7 confidence floor is a reasonable starting point, and anything under it goes to a stronger model or straight to a human. Most of the volume never leaves the bottom two rungs, which is where the cost goes down.

04

Fine-tuning: teach a small model your one job

A general model knows a little about everything. Your workflow needs it to know one thing very well: what a commitment looks like in your calls, how your team names products, which tickets are really billing tickets.

The best recent evidence comes from finance, not sales, but the shape transfers. In June 2026, Bridgewater's AIA Labs published work with Thinking Machines on six expert filtering tasks. Frontier models given a plain prompt averaged around 50 percent accuracy; with carefully engineered expert prompts, the best one reached 78.2 percent. An open-weight model they trained on the task, Qwen3-235B, reached 84.7 percent, which they describe as 29.8 percent fewer mistakes, at 13.8 times lower inference cost per task. Two honest caveats: that trained model is not small, only smaller than the frontier models it beat, and it is the authors' own evaluation in one domain.

The older evidence points the same way. Predibase's LoRA Land study in 2024 fine-tuned 310 small models on 31 narrow tasks and found them beating the GPT-4 of the day by 10 points on average, with 25 of those fine-tuned models served from a single GPU. The frontier has moved since, so read it as direction, not a current scoreboard.

The lesson is not that big models are bad. It is that a narrow, repeated job is exactly where a smaller model trained on your own examples can win on both accuracy and cost. And the training data writes itself: every time a rep approves, edits or rejects a drafted action, that decision is a labelled example of what good looks like for your team.

05

Open-weight models and fast inference

We run our call extraction on an open-weight model served on fast hosted inference, with a second provider as the fallback when the first one is down. Open weights mean the model cannot be deprecated out from under a customer, the cost is predictable, and if a customer needs it, the same model can run in their own cloud.

The cost difference is not subtle. The model we use for extraction, gpt-oss-120b, lists on Groq at $0.15 per million input tokens and $0.60 per million output tokens this month, a small fraction of a frontier model's price. Connecting models to the tools is getting cheaper too: the Model Context Protocol, the open standard Anthropic introduced in November 2024 for connecting AI tools to data sources, means one integration can serve many models instead of one.

06

Open-source tools, where they are safe

A lot of the GTM stack now has strong open-source options, and for many teams, especially small ones, they are the right call: you own the data, there is no per-seat bill, and nothing disappears when a vendor changes its pricing.

Twenty is an open-source CRM, mostly under AGPLv3. Chatwoot is a support desk under the MIT licence. Mautic does marketing automation under GPL v3. Listmonk runs newsletters and mailing lists under AGPL-3.0. n8n is a strong workflow engine, but its licence is source-available, not open source, and it limits commercial use. Cal.com's open repository, now Cal.diy, is MIT-licensed, but its own README recommends it for personal, non-production use, so for a business we use the hosted product or a different tool.

Free is not free, though. Self-hosting means someone patches it, backs it up and watches it. Some projects keep their AI features in a paid tier: Chatwoot's AI assistant is not in its free Community Edition. And some licences are source-available rather than open source. We pick per customer: open source where the team can own it, the vendor they already pay for where they cannot.

07

Evals or it did not happen

Every model in a workflow is tested against a set of real, hand-labelled examples from the customer's own data before it goes live, and again every time anything changes: the model, the prompt, the provider.

We are strict about what we say out loud. A score measured on the same examples you tuned the prompt against is an upper bound, not a promise, and we label it that way. When we switched our own extraction to a different model this year, we stopped quoting the old accuracy number, because it no longer described what was running. A number you cannot reproduce is marketing.

08

Built to run on its own

Cheap only matters if it keeps working when nobody is watching. The parts that make a workflow sustainable are the unglamorous ones:

Budgets and stop conditions on every agent, so a loop cannot run up a bill overnight.

Retries with backoff when a provider rate-limits, and a fallback provider when one goes down.

Idempotent writes, so a retry never creates a second deal or a second ticket.

Monitoring for silent failure: the sync that stopped, the queue that grew, the confidence that drifted.

Approvals that teach: the corrections a human makes become the examples the next model version is trained and tested on.

09

Where cheap goes wrong

Small models fail on edge cases a big one would catch. Fine-tuned models drift when the business changes and nobody retrains them. Open-source tools cost time instead of money. Each of those is real, and each is why the ladder escalates instead of guessing, why evals run on every change, and why a person still owns every customer-facing promise.

The goal is not the cheapest possible system. It is the lowest cost per finished task that a customer can trust.

10

What this means if you are buying

Ask anyone selling you GTM automation three questions. What does one finished task cost to run, at your volume? Which steps use which model, and why? What happens when the model is wrong? If the answers are "it's all GPT", "trust us" and silence, you are paying for a demo.

If you want to see it done the other way, we engineer your first workflow free, in your stack, and show you the cost per task before you pay for anything.

FAQ

How do you reduce the AI cost of GTM automation?

Route each step to the cheapest thing that does it well: code for lookup, matching and dates; small or open-weight models for reading and classifying language; fine-tuned models for narrow jobs that repeat; frontier models only for hard judgment calls; and a person for customer-facing commitments. Escalate on low confidence instead of sending everything to the biggest model.

Do you need a frontier model like GPT or Claude for sales automation?

For a minority of steps, yes: ambiguous emails, long threads, messy negotiations. Most GTM work, such as deduplication, field mapping, routing and date parsing, is better done by code or small models, which are faster, cheaper and easier to test.

When is fine-tuning worth it for GTM workflows?

When the same narrow job repeats at volume in your own language, such as extracting commitments from calls or classifying tickets. Approval decisions from your team become labelled training examples, and a small fine-tuned model can then beat a general model on accuracy and cost for that one job.

Are open-source tools a good fit for a GTM stack?

Often, especially for small teams: you own the data and avoid per-seat pricing. But self-hosting means maintenance, some projects keep AI features in paid tiers, and some licences are source-available rather than open source. Choose per team.

How do you know an AI workflow is accurate?

Test it against hand-labelled examples from your own data before launch and after every change, and treat a score measured on the examples you tuned against as an upper bound. Only quote numbers you can reproduce.

Want the GTM engineer without the headcount?

Start with one workflow engineered free, then get unlimited GTM engineering requests handled at a fixed rate per 4-week cycle.

Get your first workflow free
Start with one workflow

Tell us the call that keeps leaking.

We engineer the follow-through inside the tools your team already runs: the CRM update, the ticket, the recap, the handoff. Nothing customer-facing ships without your yes, and every write leaves a receipt you can reverse.

or book a 45-minute call$5,995 per 4-week cycle4-week cyclesFirst workflow free, you keep it