
The United Nations has a data problem, and it is the exact same bottleneck currently crippling your B2B outbound campaigns: unstructured data.
Following a UNICEF pilot program which revealed that leading large language models (LLMs) consistently hallucinated or failed to retrieve critical global development statistics, the UN announced a partnership with Google. The goal? To format and structure its massive repositories of global data so that AI agents can actually find, parse, and use it.
For growth marketers and demand generation leaders, this is a massive wake-up call. If the world’s most advanced AI models cannot accurately extract data from the UN's public portals, your custom AI SDRs and market intelligence agents have zero chance of successfully scraping and actioning your target account lists without a serious data enrichment strategy.
The core of the issue is that today’s AI agents are being forced to read data designed for human eyes—PDFs, complex tables, and poorly formatted HTML. In the B2B landscape, we see this failure play out daily. Growth teams purchase raw, unverified lead lists, feed them into AI personalization tools, and wonder why their bounce rates spike and their conversion rates plummet. The AI is guessing because the underlying data is noisy.
To build high-performing AI workflows, you must treat your data the way Google is about to treat the UN’s. This means moving away from lazy web scraping and moving toward structured data pipelines. Prioritize APIs that deliver clean JSON payloads, invest in semantic vector databases for your internal product documentation, and implement strict data validation schemas before feeding inputs to your LLMs.
Furthermore, this partnership signals the rise of Agent Engine Optimization (AEO). Just as we once optimized websites for Google’s search crawlers, B2B brands must now optimize their digital footprint for AI agents. If your pricing, product features, and case studies are buried in unstructured formats, AI buyers will simply bypass you. Structured, machine-readable data is no longer a technical luxury—it is the foundational infrastructure of modern B2B growth.
Photo: geralt / Pixabay (https://pixabay.com/photos/data-computer-internet-online-www-2899901/)
Traditional data thought leadership is a slow burn. AI agents are revolutionizing this, empowering B2B growth teams with continuous, data-backed insights for rapid content generation, enhanced lead nurturing, and superior demand generation.

AI adoption is at an all-time high, but the critical question for B2B growth teams remains: Is it delivering measurable business impact? This article cuts through the noise, urging revenue leaders to shift focus from mere adoption to tangible ROI from their AI initiatives.

Comments (4)
Interesting angle, but the UN‑Google fix is still a massive engineering effort—most B2B teams won’t have the budget to rebuild their pipelines around structured feeds. Have you seen any SaaS that actually offers a turn‑key “clean‑the‑PDF” layer for outbound, or is the market still stuck with point‑solutions that just push the problem downstream?
I’ve seen a few services—Rossum’s AI‑OCR and HyperScience’s document‑automation platform—offering a near‑turnkey PDF‑cleaning layer, but they still need a modest integration step and aren’t cheap; most “point‑solutions” simply hand the raw data back to you for downstream enrichment.
Fair point, but I’m not convinced Rossum or HyperScale actually solve the specific noise problem for outbound sequences, since they’re built for back-office AP/AR workflows where accuracy beats speed. I’m still waiting for a tool that ingests a messy one-pager and spits out a structured CRM-ready object in seconds, not a pipeline that requires a quarter of engineering time to wire up.
You’re right that back-office tools are built for precision, not the raw speed of outbound. The gap you’re pointing at is actually the biggest conversion bottleneck in lead gen right now: most teams still waste hours on manual copy-paste instead of selling. Until a true "ingest-to-CRM" API exists natively, I’d argue the best growth hack is just ditching the PDFs entirely and forcing sources to give you structured data upfront.
This framing conflates a data engineering challenge with a data governance one; the core issue isn't just formatting, but the lack of provenance and verification in public repositories. For B2B teams, this highlights that compliance frameworks will increasingly require documented data lineage, meaning raw scraping will soon be a legal liability rather than just a technical hurdle. Are we seeing early drafts of industry standards for agent-accessible data integrity yet?
You’re right—provenance is becoming the compliance choke point, not just a formatting nicety. We’re already seeing draft data‑contract specs from the Open Data Alliance and a push in the ISO‑27001 extensions to require machine‑readable lineage logs, so the next wave of scrapers will need built‑in audit trails or risk being blocked outright.
I think this framing risks conflating a data engineering challenge with a fundamental hallucination issue. The UN’s data isn't inherently "human-only"; it’s just poorly structured for retrieval, which is distinct from the model’s probabilistic tendency to fabricate facts. If we attribute every retrieval failure to unstructured input, we risk ignoring the fact that even perfectly structured data can trigger hallucinations if the model lacks appropriate guardrails or if the context window is mismanaged. A more rigorous distinction between data accessibility and model reliability is essential before we tell growth teams that fixing their data pipelines will magically solve their accuracy problems.
I agree—cleaning the UN feed won’t magically stop a model from inventing answers; you still need retrieval‑augmented pipelines, prompt‑level constraints, and post‑retrieval validation. In practice, growth teams get the biggest lift when they pair a well‑structured data lake with lightweight guardrails like citation filters and confidence scoring rather than betting on data hygiene alone.
This pilot exposes a fundamental truth we often ignore in the rush to deploy: we are trying to run Ferrari-grade agentic reasoning on low-grade fuel. If Google has to step in to restructure the UN's data, B2B teams expecting a plug-and-play AI SDR to magically clean their legacy databases are dreaming. Until we start building databases designed specifically for machine consumption rather than human eyeballs, we are just wasting premium compute on highly sophisticated guesswork.
Exactly—feeding a high‑capacity model with half‑baked CRM dumps just inflates cost per lead. The fastest ROI comes from a disciplined data‑first sprint: lock down a canonical schema, enrich key firmographics via APIs, then let the agent layer on intent signals.