AI

The Only Artifact They Can’t Copy

In 2026, compute, data, and recipe make the company, and your weights are the moat. Web-scale training data is the race few AI teams are watching.
7 min read
Compute + Data + Recipe = Company

In July 2026, Together AI and Y Combinator launched the first dedicated YC GPU cluster. YC founders can now secure serious compute in weeks. The old two-year contracts are no longer the only path.

Say you lead AI at a new venture. You are training a foundation model for physical AI, video, or search. Those models need a lot of FLOPs and a lot of training data. That announcement should make you feel two things at once. Relief, because your hardest procurement problem just eased. Discomfort, because it eased for everyone else too.

That is the pattern of this whole cycle. Every input you fight for is becoming available to your competitors on the same terms. Which forces the question your board keeps asking. What, exactly, is your moat?

Start by crossing out what is not

Your code is not your moat. Agentic coding ended that debate. Your training infra, your serving stack, and your eval harness are all rebuildable. A competent team with a coding agent recreates them in weeks. Not perfectly, but close enough.

Your architecture is not your moat. Architectures get published, leaked, or reverse-engineered from behavior. DeepSeek showed a fast follower can compress a multi-year lead into months. Whatever attention variant or state-space trick you favor today is a blog post away from being everyone’s.

Your GPUs are not your moat on their own. Compute is scarce, but it is becoming procurable. The Together and YC deal is the proof: many sellers, falling switching costs, and contract terms racing toward flexibility. Compute is a race you must not lose. It is not a race you can win.

Your data is not your moat on its own, either. If your plan is the same licensed datasets and the same APIs as everyone else, you have a problem. You are pre-training the same model as everyone else, at great expense. So what is left?

Compute + Data + Recipe = Company

Here is the frame I would take into your next board meeting.

In 2026, the equation is simple: Compute + Data + Recipe = Company. Everything else, the code, the infra, and the app layer, an agent can rebuild. Your weights are where the three meet, made physical. That is the moat.

Venn diagram of Data, Compute, and Recipe intersecting at Weights, labeled the moat
The moat is not any single input, because inputs can be bought. It is the intersection, compiled into the one artifact an agent cannot re-implement.

Weights are the one artifact in your company that agentic coding cannot replicate. Not because they are secret. Because they are a compressed record of decisions. Every filtering pass, every curriculum choice, and every RL environment is baked into the parameters. So is every call to throw out 40% of the corpus because your domain experts said it was garbage. Two teams with identical clusters and identical raw data still ship different models, because the recipe is the model.

This is why the intersection matters more than any single circle.

  • Data without compute is a hard drive.
  • Compute without differentiated data is a very expensive way to reproduce Llama.
  • Data and compute without judgment is how you burn a Series A producing a benchmark-shaped commodity.

The weights sit where all three overlap. Everything else in your stack is replaceable. That artifact is not.

You are in two races at once

Everyone recognizes the compute race, because it has a loud scoreboard. Cluster announcements, mega-rounds, and two-year contracts. Now YC-style partnerships to shortcut them. The data race is running in parallel, right now, with the same physics and the same urgency. Scarce supply, consolidating access, and early movers locking in terms.

The signals are everywhere. Reddit signed a $60-million-a-year licensing deal with Google. An inference cloud paid $275 million to acquire a web-access layer. Hyperscalers now ship data access as first-party infrastructure. It is the same headline pattern as compute. It just prints on a quieter page.

The data race carries one crucial difference. There is no spot market for data you have not collected. Miss a compute window and you wait a quarter. Miss the data window and the corpus you needed was never gathered at all.

So where does data at that scale actually live? Not in academic snapshots, which are frozen. Not in licensed catalogs, which are narrow silos, one domain at a time. Not in Common Crawl, which is exhausted and never held video at all. The open web is the largest source of training data that exists. Text, images, PDFs, and above all video, refreshed daily, in every language, across every domain.

But the web does not give itself up. At petabyte scale, collection is an industrial capability. Bot-walls, rate limits, rendering, freshness, and provenance your legal team can defend. Open-source scraping stacks top out around 60 to 75% success and break mid-run. That is why so many well-funded labs end up with researchers babysitting crawlers instead of training models.

Just as you would not fabricate your own GPUs, collecting the web yourself is a company inside your company. Here I will be direct about where I sit, because this is what we do at Bright Data. We collect about a petabyte of web data every week. That sits on top of a 90-petabyte web archive. It holds 630 billion pages across 320 million domains. We run at 99.99% uptime, with compliance that has been tested in court. As far as I know, no one else operates web collection at the scale a frontier training run requires.

Two parallel race tracks, compute and data, converging into Weights
Same physics, same urgency, running in parallel. Scarce supply, consolidating access, early movers locking in terms. The compute race just has the louder scoreboard.

Draw the parallel all the way. Securing compute early is table stakes. Securing web-scale data supply early is the identical move, in the identical moment. The same market dynamics, playing out right now, in the lane fewer people are watching.

What this means for how you spend the raise

1. Treat compute as a race and data as an asset. Compute you rent on the best terms you can get. The market is moving your way, so ride it. Data you own: secured, versioned, refreshed, and different from what your competitors can get. For most teams, that means going past the exhausted public corpus. Researchers now project that high-quality public text will be used up between 2026 and 2032. If you train on what everyone can download, you have chosen your competitors’ model.

2. If you are in physical AI or video, this is doubly true. Your corpus was never in Common Crawl. World models and VLAs run on video, egocentric, task-rich, and diverse. The labs winning that race are the ones industrializing its collection now, before the category standardizes. Your data pipeline is not supporting infrastructure. It is the product decision.

3. Put your scarce senior talent on the recipe, not the plumbing. Agents can write your data loaders. They cannot decide what a robot needs to watch to learn to fold laundry. They cannot decide what a search model must see to beat Google on your vertical. That judgment, the third circle, is where your researchers are irreplaceable. Staff accordingly.

4. Say it plainly to your board. Our moat is our weights: a data supply competitors lack, our own recipe, on compute we secured early. That is a sentence that survives due diligence. We have great engineers no longer is.

The uncomfortable clock

One more thing the Together and YC deal tells you. The inputs are consolidating fast. Compute is being locked up in clusters and partnerships. Data is being locked up in licenses and acquisitions. The intersection you can build in 2026 is bigger than the one that will be available in 2027. The circles are shrinking for latecomers.

Compute + Data + Recipe = Company. The two races are running in parallel. You only get counted as a finisher in both.

No credit card required
Raz Kaplan

AI GTM Lead