Alternative Data Weekly #301
Theme: Connection is cheap. Context is expensive.
ADW is free, but I appreciate the support of those who are paying subscribers.
Special thanks to our sponsor Sidekick Lab!
Sidekick is the Agentic AI layer that gives your enterprise a living, governed understanding of your data estate.
We provide the ontology for your data to finally speak the same language as your business.
QUOTE
“MCP makes connecting data very easy. But what happens when you have 10,000 datasets to connect? MCP is going to eat up all your context. It’s not practical.” Ying Hua (~35:40)
News
Pods
Charts
Final Thoughts (MCP = Margin Compression Protocol)
#1 – Maribeth Martorana & Andy Graham published Enterprise Data Is Becoming the Next Competitive Advantage. July 2026.
My Take: Once your internal data is organized (i.e., governed, structured, compliant, etc.), then the data’s value compounds when using AI tools.
The right question from the executive level is, “how do we increase the value of our enterprise data assets?”. I’ve said for years the value of just getting your own data house in order is worth the effort, whether you formally commercialize the data externally or not.
Every investment in governance, metadata, quality, lineage, and stewardship is increasing the value of a strategic enterprise asset.
#2 – Srinivasa Mathkur published The Context Gap: Why Data Products Anchor AI Success. July 2026.
My Take: The models are very good, but they need very specific context to be great. This is where humans add a ton of value. “Knowing a lot in general” and “knowing enough about this specific situation, right now” are very different things. Models are very good at the former and need a lot of help to be very good at the latter.
When the “customer” of the output is an agent, not a human, this amplifies errors early in the process.
Srinivasa takes us through the journey of creating an environment where you get the best output from the tools. “What does a model have when it answers you”? From semantics to governance to data quality, and provenance/lineage … these are all essential to generating trustworthy output.
#3 – Adrian Krebs published How Hedge Funds Source Web Data in 2026. July 2026.
My Take: Web scraping is like a free puppy, not free beer. The maintenance is what gets you.
When funds are thinking about “build vs buy”, they should think about how the collected data will be used, how the pipes will be maintained, and if the “what is collected” or the “how it’s collected” is proprietary to their workflow or a commodity input.
BONUS: Matt Ober’s The Cost of Information is Expensive. July 2026. “The game of data is costly when it’s exclusive data.”
What else I am reading:
Simon Burton published Discussion with The Terminalist. July 2026.
Matthew Bernath published The State of Data Monetisation in African Financial Services. July 2026.
Benn Stancil published Is “jobs” a four-letter word? July 2026.
Joe Reis published To Every Agent Its Own Database. July 2026.
Sarah McKenna commented on Public Search Results and Responsible Data Access. July 2026.
Bilal Jafar published Point72 hires AI expert to predict weather for traders. July 2026.
James Evelegh published How to unlock the full value of your data. June 2026.
SOURCE: Ethan Kho’s Odds on Open podcast published Ex-Balyasny PM: Quant is like Blackjack, Fundamental is like Poker. July 2026.
My Take: Ying Hua, founder of Implied, breaks down how top multi-manager hedge funds synthesize quantitative discipline with discretionary analysis to extract repeatable market alpha.
I really enjoyed this interview. The investment world is changing quickly (isn’t it always) … Ying shares some valuable perspective.
Quant-amental ... what does that even mean? Using quant and quant data to inform fundamental analysis. Historical pattern matching = quant. Things that are not pattern matching increase in value.
AI tools shift focus to “art” … things that are well studied can be outsourced. What does analyst do?
Gather data
Process data
Analyze data (hard for LLMs to do)
Q: How quickly can you update models? A: Not fast enough. (Full disclosure: this is what we do at SymetryML.)
“We changed our process so that every model can be updated within two minutes.” (01:30)
Forecasting a loss in insurance is “art” (interesting example of alt data usage for insurance use case). Another example shared of using credit card data (19:00).
Another person who thinks these tools will expand the demand for investing talent.
A differentiated view today pays more.
“AI is just another wave of quant, but for word data instead of number data.” (21:30)
HIGHLIGHTS (70-Minute Run Time)
Minute 01:00 – interview starts.
Minute 05:30 – the importance of sizing.
Minute 08:00 – insurance example (“quantamental”).
Minute 12:00 – can fundamental investing be automated?
Minute 17:00 – does demand for investing talent expand or contract as a result of AI tools?
Minute 21:00 – why can’t I just tell Claude to do all this for me?
Minute 30:00 – how do you make sure the models don’t commodify your business?
Minute 35:00 – MCP connecting data … chaos ensues
Minute 38:00 – AI picks up regime changes?
Minute 39:45 – First thing you’d advise a fundamental PM to do with AI?
Minute 42:00 – what moats expand (quality of talent)? Ask good questions. Reading people.
Minute 44:45 – willingness to learn the technical side.
Minute 49:00 – fundamental is poker; quant is blackjack.
Minute 53:00 – choosing the right sector is key.
Minute 58:00 – philosophy for building wealth.
Minute 65:00 – traits of the good and the bad investors.
Minute 69:00 – importance of thinking from first principles.
SOURCE: Andre Retterath published The Analyst Role Is Dead, No? July 2026.
My Take: Data quality is key. The job of the analyst is changing. Really cool review of use cases across the board.




BONUS: LSEG published AI’s Role in Wealth Management. July 2026.
My Take: Good data is key. Wonderful series of use cases for a very relationship-driven industry. Many would think this to be counterintuitive.


MCP = Margin Compression Protocol
The question of whether data providers will plug into AI agents is answered.
Asymmetrix surveyed the industry: 56% have launched an MCP server, 40% are considering one. 96% in. In the same month, S&P launched Adaptive Retrieval, LSEG shipped its MCP connector, and Bloomberg explained why it built MCP and then pointed it at itself. Wildly interesting timing, as in the same week The Terminalist explained what all this means for the industry.
Everyone has pipes. How does it get priced? The same Asymmetrix survey shows answers ranging from a free add-on to a 50% subscription bump. That wide of a range indicates no one knows.
My take: Once an agent can easily query five vendors through the same connection, “our data is special” becomes a testable claim, checked in seconds, thousands of times a day. Bloomberg sees this, which is why their MCP faces inward. Keep the agent, keep the customer. S&P and LSEG are betting the opposite: open the pipes, win on trust, citations, and token efficiency.
Both can’t be right over the long run.
My thought is that some version of Jevons paradox will happen and the amount of data being “used” in various ways will explode. Data-hungry non-human agents can consume data at a pace no human can comprehend. Price per query might collapse, but volume goes vertical. The fight is over who captures it.
For data owners, reachable by agents is now table stakes.
The edge is being the source the agent goes to when the answer has to be right.







Thank for the call out in your weekly newsletter to myself and Maribeth Martorana.