Alternative Data Weekly #297
Theme: AI tools are unlocking the value of data
Special thanks to our sponsor SymetryML.
Bring real-time ML to your clickstream data, through Claude or existing workflows.
100TB+ of clickstream? No Problem. SymetryML lets analysts interrogate their massive clickstream data in seconds, without repeated warehouse scans, heavy MLOps, or a full data science team.
Ask questions in plain English. Run advanced analytics and ML.
We’re opening 5 free pilot slots for teams working with large clickstream datasets.
Get details and claim a slot: hello@symetryml.net
QUOTES
“Some of the best data leaders I know are notorious for going looking for trouble nobody asked them to find”. - James Miller
News
Pods
Charts
Final Thoughts (250)
#1 – Abraham Thomas published On Data Quality. June 2026.
My Take: Another must-read from Abraham (teased as a “Part 1”). He shares with us the different levels of data quality: granular, aggregate, fitness, value. These are laddered and reliant on the prior step.
“Data quality is that which increases data value.”
The key question is about how the data will be used. That drives how we should think about “quality”.
#2 – Jon Stojan published The Real Advantage Isn’t More Data. It’s Better Context. June 2026.
My Take: Data access used to be the source of advantage (the internet helped solve this). Then it was the ability to process that data (AI levels the playing field here). Now it is organizing it the right way, giving it the context of what you are doing, and asking the best questions of the data.
“Sports technology has spent years gathering data. The next chapter may focus on helping people make better sense of it.”
BONUS: Rob Mannix & Mauro Cesa published How quants are getting the most out of Claude. June 2026. “…quants are discovering the tools require skillful handling”
What else I am reading:
Peter Baumann published Data Contracts - Putting Data Governance into Practice. June 2026.
Modern Data 101 published The State of Data Products (Q1). March 2026.
Dan Entrup published Raising Cannes. June 2026.
Source: Ben Lorica of The Data Exchange Podcast interviews Chang She, co-founder and CEO of LanceDB. The Data Stack Wasn’t Built for AI — Here’s What Comes Next. June 2026.
My Take: Ben is one of my favorite follows in the data world.
This podcast resonated with me because Chang talks about how multimodal exhaust data that enterprises spent a decade shoving into a back room is now proving to be valuable (finally, right?) … mostly because the tooling to finally draw signal out of it exists.
Bottom line: data is more valuable because of AI. Agents are driving this to massive scale (100x++).
HIGHLIGHTS (40-Minute Run Time)
Minute 00:45 – Chang shares background and thoughts on why LanceDB is needed.
Minute 02:00 – Vector Databases discussion (what works and what does not work)
Minute 05:30 – LanceDB – multimodal data (“open”)
Minute 10:00 – the “bay area bubble”
Minute 14:00 – various types of multimodal (data types, workloads, ops)
Minute 17:00 – the massive scale is causing problems
Minute 20:00 – agents & memory
Minute 22:00 – “lightweight” … make it easy for agents to engage with massive amounts of data.
Minute 27:00 – building more capable agents
Minute 28:00 – CXO thoughts (pain points around new AI data silos)
Minute 30:00 – data stack in a few years (data tooling ”middle layer” will start to melt away with agents taking over)
Minute 34:00 – the rise of open weight models
BONUS: Neudata’ Stuart Broughton hosted Explore data for sale: Winners, failures & the decisions that made them . June 2026.
SOURCE: Jim Rowan, Beena Ammanath, Nitin Mittal, Costi Perricos of Deloitte published The State of AI in the Enterprise. January 2026.
My Take: This first chart is interesting … hopes were too high for everything but productivity.
BONUS: Asymmetrix published the results of their MCP adoption survey.
There is a full report available.
Happy Fourth of July for my American readers.
Some “alternative data” facts about the USA (AI helped me here).
1. America was a data project from day one. The Constitution (Article I, Section 2) requires a head count every ten years. The first census in 1790 found 3.9 million people. The next one, in 2030, will count around 345 million, roughly 88x larger. We may be the first country to write recurring data collection into its founding charter. The original alternative dataset was the census.
2. The greatest coincidence in the data. John Adams and Thomas Jefferson both died on July 4, 1826, hours apart, on the exact 50th anniversary of the Declaration. James Monroe followed on July 4, 1831. Three of the first five presidents died on Independence Day. Run those odds.
3. We’re celebrating the wrong date. Congress actually voted for independence on July 2. Adams predicted July 2 would be “the great anniversary festival.” The text was adopted on the 4th, and most delegates didn’t sign until August 2. We commemorate the metadata, not the event.
4. The oldest “new” country. America feels young at 250, but the Constitution (in force since 1789) is the oldest written national constitution still operating. Most of the world runs on a newer one than we do.
5. The 22-million-row dataset that vanished. During the 1976 Bicentennial, a wagon-train time capsule holding roughly 22 million Americans’ signatures was stolen from beside the speaker’s platform while President Ford spoke at Valley Forge. Never recovered. The original lost dataset.












