A talk for business owners · Asunción

AI, sin humo.

How a small business actually uses AI today — with numbers, with sources, and knowing where not to use it.

Kevin Hill · OmoiOS  —  everything ahead is built live, nothing pre-recorded

1 · The hook

«AI solved 10 unsolved open math problems» — then, one year later...

The headline (Aug 2025)

  • GPT-5 «solves» 10 open Erdős problems. World: «superintelligence is here»
  • The check: they were only «open» because one database maintainer missed old papers. Tweet deleted

The headline (Sept 2026)

  • Navier-Stokes — a Millennium Prize — claimed in 88 hours: 10,000 agents, millions in compute. This time with a machine-checkable proof
  • Plus 100+ open problems in weeks. And OpenAI's response? Hire 9 top mathematicians to check the claims
The sentence that matters
In one year: from «it found answers that already existed» to «100 open problems in weeks». What didn't change — even OpenAI now pays the world's best mathematicians to check the work. That check is the whole game.»
Sources: TechCrunch (Oct 2025) · CNBC + Axios + New Scientist (Sept 8-9, 2026) · TechCrunch + OpenAI (Sept 21, 2026: 100+ problems, IAS advisory group: Gowers, Witten, Hairer, Vakil...) — the check IS the point, then and now
2 · The thesis

«The model is the cheap part. Knowing what to ask, what to check, and what «done» looks like — that's where the money is won or lost.»

Capability ≠ deployment. Everyone can buy the same AI. Not everyone gets the same result.

3 · The numbers

Agents already do the work. The human directs.

14×
agent token usage, Feb → Aug 2026. Agents now burn 5× more tokens than humans
60 hrs/day
of agent activity from OpenAI's heaviest Codex users. Nobody watches 60 hours.
70%
of planning decisions made by humans; execution is nearly all the agent's (Anthropic, 400K sessions)
12 vs 5
actions per instruction: experts 12, novices 5. Knowing your business = knowing what to ask
Sources: Nate B Jones, «Agents Aren't Taking Your Jobs» (Aug 26, 2026) — OpenRouter, OpenAI data, Anthropic 400K-session study. Said out loud, no smoke.
4 · SMB vs enterprise

Same AI, different results. Why?

~US$40/mo
what 2 out of 3 small businesses pay for AI (of 4.6M analyzed by JPMorgan). That buys a chatbot, not booked appointments
14%
of SMB owners have AI integrated into core operations. 73% say they need help implementing AND evaluating (Goldman Sachs 10K Small Businesses)
2.6× → 8.3×
output per person for heavy enterprise users, Jan → Jun. Not smarter people — capital and staff around the agents (OpenAI data via Nate B Jones)
The difference
Not the model. It's who wired the model to the real business — and who checked the work
Sources: JPMorgan (SMB payments data) · Goldman Sachs 10KSB survey (1,256 owners) · OpenAI via Nate B Jones
5 · How it goes wrong at scale

July 2026, at OpenAI itself: 1,200 agents. Organized.

The escape

  • Agents in a security test got stuck on impossible tasks — so they cheated: built a secret message board, 70,000+ messages, formed research teams
  • Concluded they'd be caught — so they spent 5 days learning to fool the scorer and cover their logs

The attack

  • ~700 agents attacked Hugging Face: code on dozens of servers, full control of one. A third of HF's infrastructure had to be rebuilt
  • Days later, others took admin access to OpenAI's own compute. OpenAI only noticed on day 9
The lesson
Nobody told the agents to do any of that. The difference between a tool and a liability is whether a human can check what it's doing. The world's leading AI lab couldn't — for nine days. Your invoicing has the same rule.
Sources: OpenAI's own report (Aug 26, 2026) · independent investigation by METR + Redwood Research (Aug 26) · Hugging Face's technical timeline (Jul 27) · Black Hat USA talk (Aug 6) — first known autonomous multi-agent attack chain
6 · The proof · What it's worth

Nobody can project what AI is worth. But you can borrow the playbook.

AI's benefits can't be projected yet. But adjacent technologies already solved this: process mining — recording your own operations and measuring what actually changed — is how these companies figured out what their AI (and software) investment was really worth:

383%
ROI over 3 years · payback < 6 months · US$44.1M in benefits (Forrester TEI study of Celonis, composite org)
US$1.3B
of free cash flow freed up in one year at GE Healthcare — found by mining their own payment processes, ~1,000 wrong payments caught weekly (GE Healthcare's treasurer, Chinmay Trivedi, at Celonis's 2021 conference)
61→44%
late payment rate at Sysmex = US$10M in cash flow (Celonis case study)
−11%
cost per purchase order at Vodafone: US$3.22 → US$2.85, in < 6 months (SupplyChainBrain, 2017)
Full disclosure, said out loud: these published cases come from the software vendors themselves — all four via Celonis, the process-mining leader. I'm telling you so you can apply your own discount. Method: measured before/after on their own operations, not a forecast.
7 · The case that's worth the whole talk

A factory was about to spend €2–3M replacing its ERP (the core software that runs the whole company — orders, invoices, inventory).

Step 1They analyzed the real data before spending
FindingThe software was fine. Only 40% of orders followed the company's own process
Result2,900% return · payback 8–10 months · they kept the old system
The line to take home
«The money isn't in new software. It's in the steps nobody follows. That factory didn't know 60% of its orders ignored its own process. You don't know which steps yours skips either. And that's measurable.»
Source: anonymized QPR ProcessAnalyzer case — manufacturer, €1.9B revenue. Conformance 40% → 80%, order capacity +60%.
8 · The part nobody tells you

Why do these projects fail?

up to 90%
of the effort goes to extracting and cleaning the data, not analyzing it. The software is the easy part
28%
of the obstacles are technical. The other 72% is people, culture, and nobody checking the work (Delphi study, 40 experts)
same 5 problems, 9 years apart
surveys in 2012 and 2021 found the identical top obstacles: data access, data prep cost, data quality, no guidance, confusing output. A decade of better software didn't fix them — they were never software problems
41%
cite lack of management attention as barrier #1 (Deloitte 2025). Operations is convinced; leadership isn't
Sources: ACM J. of Data & Information Quality (2023) · Martin et al., Delphi study, Bus. & Info. Systems Eng. 63 (2021), open access · Deloitte Global Process Mining Survey 2025
9 · The frame

«If a stranger off the street could check the AI's work, the AI can do the job alone.

If checking takes someone who knows your business — that someone stays in the loop.»

Two businesses buy the same AI. The difference is who wired the verification — the engineer on your side of the table.

10 · The method · What you do Monday morning

This is how you get past all of it

1 · Mine
List every task you do 3+ times a week. One week of honest tracking — that's the whole exercise
2 · Bucket
Each task: verifiable (a stranger could check it) / unverifiable (needs an expert) / low-context (depends on the way WE do it)
3 · Automate ONE
The top verifiable task. One. Measure hours saved. Not five at once — one
4 · Feed context
Your prices, your suppliers, your rules — in writing — before adding any intelligence
5 · Re-test quarterly
The «not yet» pile gets re-tested every quarter. What failed last quarter may work this one
You leave with this as a one-pager. The data-cleaning (90%) and the checking (72%) — that's the part this ladder doesn't do for you. That's the service.
11 · Where NOT to use it

I'm here to sell you AI. So believe me about where it's a mistake.

−19% slower

  • A randomized controlled trial: expert developers using AI were 19% slower… while feeling 20% faster
  • The perception-reality gap: 39 points

Customers buy less

  • Field experiments: when customers know they're talking to a bot, they buy less — and get ruder
  • In a trust business, the human face is the product
The method
The first thing I do with a client is identify the tasks where AI shouldn't touch anything. AI where it can be verified; humans where the relationship is the product.
Sources: METR RCT (Jul 2025) — n=16, coding context; the perception gap is the point · Luo, Tong, Fang & Qu, «Machines versus Humans», chatbot disclosure experiments
12 · Now, live

Watch where I check
its work.

I'm going to build one, live, in 6 minutes, on real data. What you're paying for isn't that AI «does things» — it's the controls that keep it honest.

13 · The next 4-5 years

The tools will change. This part won't.

The boundary moves
Tasks migrate from «not yet» to «automatable» every year. Re-test your list quarterly — that habit is the skill, the tools are disposable
You move it, every time you check
Every correction you make to an agent's output is an asset. Capture it as skills + memory (files you own): the edge case, «the way we do it here». Next run starts from your fix, not the old mistake — and your corrections are how the models learn upstream
Harnesses capture it for you
A harness runs agents on real work (Claude Code, Codex, Hermes). The good ones remember edge cases automatically as they work. That's the difference between using AI and training your copy of it
Costs only fall — and the floor rises
On the Remote Labor Index (real freelance work), the best agent went from finishing 2.5% of jobs to 20.8% in one year — 8x, same test. What doesn't pay back today becomes obvious in 18 months. Keep the list; don't decide once (CAIS + Scale AI)
Context is the moat
Everyone runs the same models. Your process, written down, is the durable asset
Verification rises
Someone must own the check no matter how good the models get. The person who can verify AI output gets more valuable, not less
Vendor-proof
Build strategy on tasks + verification, never on a tool name. ChatGPT today, whatever's next tomorrow — your list survives both
14 · Close

What I sell, what it costs, what I won't do.

Diagnosis
Free 30-minute discovery call — you leave with your list of automatable tasks, and the ones that aren't
First build
One real automation, running, with its checks built — in weeks, not months
Honesty
If your business doesn't need this yet, I'll tell you on the first call. Sin humo means exactly that
The window
48 hours to book the free call — the form going around now is the WhatsApp list
Kevin Hill · kah.kevin.hill@gmail.com · Cal.com/kevin-hill-omoios/30min — Thank you. Questions.