Podcast/Episode 18

Why Your AI Gives You Confident, Wrong Answers | Ahmed Elsamadisi

Fifteen years on the layer between data and the questions asked of it. He created the Activity Schema, which is open source and which other companies now consult on. And he still cannot fact-check what an AI tells him.

That last part is not modesty. It is the argument.

Ahmed Elsamadisi is the founder of Polyform, where the work is building data an AI agent can actually be trusted with. His starting claim is that most people making decisions on AI-generated numbers have no mechanism for checking them, and that the systems producing those numbers were never tuned to be right in the first place.

In the latest episode of Velocity: Performance Marketing Podcast by Hawky, host Surender (Co-Founder, Hawky.ai) asked him whether a data agent can be engineered to say "I am not sure." The answer ran through why his own AI is forbidden from writing SQL, why pointing an agent at GA4 returns fluent nonsense in two different ways, and where 10 to 20 percent of most ad budgets quietly goes.

About Ahmed Elsamadisi: The Man Who Made Data Legible to Machines

Ahmed created the Activity Schema, an open-source model that turns everything happening inside a company into events. Not tables, not reports. Events. Someone signed a contract. Someone made a payment. Someone viewed a page.

The point is not elegance, it is decomposability. A metric built on events can be taken apart. You see that total sales means the event "signed contract," you see what signed contract means, and you can pull up a single customer and watch what they actually did.

That chain is what makes an answer checkable: "That kind of ability to go from a metric to its components, to its how the components are assembled, to what the components mean to example customers, is the same way that a human being would do it, but it's how the AI figures that out as well."

Without the chain, a number is a number that appeared.

Helpful, Not Right: Why You Cannot Fact-Check Your AI

To verify an AI's answer you need to read SQL and know how your company models its data. Neither is a marketer's job.

SQL is genuinely awkward, he points out, because it is not linear. It is stack-based, running left and right before it runs up and down, so reading someone else's query means holding the whole structure in your head. On top of that sits undocumented business logic that lives with a handful of people.

His own verdict is the one worth sitting with: "I can't do it, and I've been doing this for a long time."

He cites research on AI-generated data analysis in which the strongest available model answered roughly a quarter of questions correctly against raw company data. Supplying it with 30 to 40 metric definitions as worked examples lifted accuracy to around 95 percent. That sounds like a solution until you do the division. One answer in twenty is still wrong, after everything has been done properly.

The damage is not in the error rate. It is in the gap between the error rate and what users assume: "The people building it using the most expensive models tell you it's 25%. You using it, you're probably using a much cheaper model, think it's a hundred percent, and that gap is where things go wrong."

And the failure mode is not refusal. It is fluency. "If your problem is I need to know how many sales I got so I can do my job, what's helpful is me giving you a number. Like I can say 36, and you'll say, cool, now I can put it in a deck and move on."

Why it behaves that way is a design decision, not an accident. The prompts behind these assistants are written to be supportive and guiding. "There's no reliability on accuracy."

Building Blocks, Not Queries: Why His AI Is Banned From Writing SQL

This is the counterintuitive part, coming from a company selling AI for data work.

"We never let our AI write SQL. Like it's literally you cannot do that."

Instead the agent assembles pre-validated building blocks, each one already checked, so the assembly is correct because every piece was correct before it arrived. Nothing enters the system unvalidated: "We don't put data in the system until it's validated and reliable."

The person doing that validating is part of the product rather than an implementation cost. Polyform puts someone on the account weekly, confirming each new building block and metric before an agent goes near it.

He is equally firm about what does not work. Asked how you get an organisation to agree on what a metric means, his answer is that you do not: "No, no, no, no. We never do that. That's how you spend like six months in consulting."

Build the shared concepts first, since a signed contract means the same thing in every department. Then, when someone asks what total sales means, refuse the premise of a single answer: "What are the seven ways you want to define it? Let's create seven different sales metrics."

Seven maintained definitions is cheaper than one agreed definition, because the agreement never arrives.

Why GA4 Breaks AI Agents in Two Opposite Ways

For anyone running paid media, this is the most immediately usable stretch of the conversation. Asked directly whether an agent can simply query GA4, Ahmed does not hedge: "GA4 with agents is pretty much garbage. It gives you such little useless information." On the MCP server route specifically, "it's utterly trash."

The trap is that the raw export is not the escape either. Google will stream every event into BigQuery free of charge, which he calls a genuine money saver, and then you open the table: "If you try to select star and see what the table looks like, it's not even usable."

There is a reason for the shape. A different Google team built the format for a different product, that product was deprecated, and the export carried the original structure across. It was never designed around the questions marketers ask.

So the reporting layer hands an agent aggregates with the definitions already applied and no rows underneath to audit, including the attribution model decision that assigned the credit. The raw layer hands it everything, unlabelled, with nothing that means visit or order until somebody decides what those words map to.

What makes the verdict pointed rather than dismissive is that GA4 is still the first thing he connects. It carries click-level tracking nothing else has. The condition he attaches holds the whole argument in one clause: "not aggregated GA4, but raw data."

Alongside it he pulls email engagement, because that is where identity resolution starts. When someone opens an email and lands on a URL, the address ties to the cookie, and two sessions that looked like two strangers resolve into one person. Without that stitch, the buyer who clicked on a laptop and purchased on a phone appears twice, as a click that failed and a sale nobody can attribute.

His fix is unglamorous and, he says, standard on every account: "Bring it to warehouse, warehouse, model it, clean it, use it."

Then the line most vendors would never say. Under one to five thousand dollars a month in spend, do not build any of it: "Don't invest in the data work, just kind of gamble, know that you're gambling and it's fine."

The Click Taxonomy: Where Ad Budget Actually Dies

Ahmed's read on paid media is that most of what gets called performance is not performance yet, because nobody has removed the noise first.

On attribution his default is blunt, with one condition attached. Use first touch 99 percent of the time, unless analysis shows your audience genuinely spans a media mix across platforms.

Then the clicks themselves, which are not one thing. Some people click and close immediately, having clicked by mistake. Some open, see it is not what they wanted, and leave, which is a creative and targeting failure you have already paid for. Some click and explore, which is the signal that targeting worked. Some click and buy. And some click, return days later, and buy on a different device under a different cookie, which sends the conversion somewhere else entirely.

Then the category most marketers have felt without ever sizing: "How many ad clicks are you spending money on for somebody who already uses your product? It's another 10 to 20 percent of your marketing budget."

He is similarly direct about platform-reported returns, recalling advertisers told they had earned four times their spend: "I spent the ten thousand dollars and I didn't get forty thousand dollars. And they're like, yeah, but you will, because of proprietary algorithm."

Strip the junk out and paid media behaves the way theory says it should: "Performance marketing done right, it should be relatively linear until saturation."

Two Tests That Need No Data Team

Before investing in any of this, two checks will tell you whether your current reporting is real.

The first is linearity. Take last month's spend and the revenue that actually reached the bank, then do the same for this month. "The easiest way to tell if your system is modeled correctly is that is it linear? If I double my spending, does it double my results?" If spend doubled and revenue did not move with it, the problem sits upstream of your campaigns.

The second is addition. Work out what share of traffic is organic and subtract it, then total up the sales every platform claims. If the combined claims exceed what you actually made, at least one of them is wrong. "A lot of times it will. So you know that everyone's lying to you."

He offers a noise threshold too. If weekly conversion rate swings by more than 10 percent, stop reading the movement as signal: "At that point, just flip a coin." A ROAS figure read the same way, in isolation and week to week, tells you about as much.

Key Takeaways from the Episode

  • AI is built to be helpful, not correct. The prompts behind most assistants optimise for supportive and guiding, and nothing in the output distinguishes a verified number from a confident guess.
  • An answer is only checkable if you can walk the chain. Metric, to components, to a real customer whose journey you can inspect. Without it nobody can verify the number, including the people who built the system.
  • The accuracy gap is the danger, not the accuracy. Roughly one answer in twenty stays wrong even after metric definitions are supplied, while most users assume the number is certain.
  • GA4 fails agents twice, in opposite directions. The reports give aggregates with definitions already applied and no rows to audit. The raw export gives everything, unlabelled, in a shape built for a product Google deprecated.
  • Model before an agent queries. Warehouse, model, clean, then use. Ahmed calls this standard on every account rather than a special project.
  • Under five thousand dollars a month, do not build any of it. Accept that you are guessing and put the money into media.
  • Two tests need no data team. Does spend scale linearly into revenue, and do the platforms' claimed sales add up to more than you actually made.

Apply these insights to your campaigns.

Hawky AI applies creative intelligence automatically across your ad library.