Differential Labs

We are hurtling into a future we do not properly understand — and this may speed up further.

One of our best hopes for getting good outcomes is differential technological development — achieving stabilizing capabilities before destabilizing ones.

At Differential Labs, we build scaffolds/infrastructure/evals to help speed up macrostrategy research: thinking about the position the world is in, and where it is going.

Our work aims to help sensemaking on the most important questions keep up.

ifferential Tech

The strategy of differential technological development has received significant attention over the last quarter-century. Our belief is that it is especially with the rise of powerful LLMs that this is becoming a high-leverage strategy.

For more of our thinking on this, see:

Our focus — macrostrategy research automation

Research is a major driver of what happens in the world, and research automation is going to be tremendously important. Through the intelligence explosion / industrial explosion, we will probably move through different eras.

  1. we are roughly here

    Augmentation

    perhaps −20% – +50% uplift

    Research largely human-executed, but with AI assistance for some tasks providing uplift.

    Where we are right now for many types of research.

  2. Centaurs

    perhaps 1.5× – 30× uplift

    Most fine-grained research tasks performed by AI, with significant human integration and steering.

    Like Claude Code, but for research.

  3. Full automation

    uplift as a multiplier on humans stops being the right metric

    A lot of research can meaningfully be automated end-to-end, but it still benefits from human direction — at least from the best humans.

    There’s a highly uneven distribution of compute across human researchers, on efficiency grounds.

  4. Obsoletion

    Humans have little if anything to add to the research process.

There are a lot of details here which are not determined — which research domains reach which stage when, what it looks like in a fine-grained way, and so on. This means that working out how to make research automation go well — what we might call research-automation research — is potentially a high-leverage activity.

What could making it go well look like?

There are a few different high-level strategies.

Differential development by application domain
Some kinds of research seem generally better to accelerate early than others: macrostrategy and alignment, and more broadly wise decision-making and conceptual research, ahead of broad AI progress or weapons research.
Differential paradigm development
Using paradigm in a sense loose enough to cover quite fine-grained things: some ways of automating research to a certain level may be more broadly desirable — methods that give more transparency, say, or more robustness.
Differential deployment
Making valuable tools available sooner to actors that it seems good to accelerate.

The best strategies in practice may employ a mix of these approaches. Since some types of research automation are fraught, it’s worth being somewhat careful about how it proceeds. But developing expertise and kickstarting a field of research-automation research among thoughtful, well-intentioned actors seems good.

We are focused on macrostrategy research — thinking about the position the world is in, and where it is going. We think this is an important part of helping to position us to handle challenges well, and it may become more important over time if the pace of change accelerates.

About Differential Labs

Differential Labs is a remote-first startup, majority-owned by a nonprofit, Differential Initiatives. Our goal is to try to help the AI transition go well, by building beneficial forms of research automation.

For now we are effectively functioning as a nonprofit, and we may do so forever. We went for the slightly unusual corporate structure because we think there’s a chance that at some point it will make sense — from a make-things-go-well point of view — to commercialize something we’ve developed, and that this may be easier if we’re set up this way from the start.

Team

  • Owen Cotton-Barratt

    Founder

    Owen is an ex-mathematician. He has been working on AI futures strategy since 2013, and focused since 2024 on beneficial AI applications for epistemics and coordination. He is a research associate at Forethought.

  • Lawrence Phillips

    Founder

    Lawrence has been building LLM-based epistemic tools since the days of GPT-3.5. Before Differential Labs, he was cofounder and CTO at FutureSearch; before that, he led AI-forecasting research at Metaculus.

Beyond our core team, we work with a number of research collaborators and consultants. The governing board for the controlling nonprofit is Owen, Lawrence, and Nicole Ross.

Our work so far

Notes on where our scaffolds have got to. We’ll add to this as things change.

August 2026

We’ve been iterating over different scaffolds and testing things. To give a rough sense of where things are at:

  • The cleanest win to date is in brainstorming — we have a scaffold that seems systematically better than asking an LLM chatbot or agent to brainstorm ideas.
  • We have scaffolds which write complete research articles.
    • There is something to the articles — they often contain at least some novel-to-us insights, much more so than asking LLMs one-shot.
    • However, overall they feel pretty second- or third-rate. They do not reliably focus on the most important dimensions, and they typically produce text which is a slog to read.
      • It’s a bit like having a grad student who is competent in some ways but whose taste is off.
  • We have some ad-hoc investments in evals to tell how we’re doing; but so far we’ve been working in the domain where you can manually inspect outputs and roughly tell what’s good.
  • We’re experimenting with other ways to integrate research tools into a Centaur-like setup.

Overall it’s feeling like we’re moving into the foothills of the Centaur era. It feels like there’s valuable stuff to be had from systems at the moment, but it’s also kind of annoying as an experience, and this could stand to be improved. Maybe this is similar to where coding agents were in early 2025?

We do think there continue to be a good number of low-hanging fruit. The repeated experience is something like: see a way that things are dubious come up with ideas for a scaffold that would address that problem implement, and perhaps with some tweaking it generally works. We are capacity-constrained on trying more of these things.

Our vibe-based assessment is that improvements in these scaffolds are helping advance automation of this research at a significantly faster rate than background improvements to LLMs. Although this is a bit complicated by the fact that recent updates to LLMs have not always seemed to be improvements in our domain — in particular, we have tasks, which are not particular to scaffolds we have built up, where Opus 4.6 seems to outperform later Opus and Fable models.