August 2026
We’ve been iterating over different scaffolds and testing things. To give a rough sense of where things are at:
- The cleanest win to date is in brainstorming — we have a scaffold that seems systematically better than asking an LLM chatbot or agent to brainstorm ideas.
-
We have scaffolds which write complete research articles.
- There is something to the articles — they often contain at least some novel-to-us insights, much more so than asking LLMs one-shot.
-
However, overall they feel pretty second- or third-rate.
They do not reliably focus on the most important
dimensions, and they typically produce text which is a
slog to read.
- It’s a bit like having a grad student who is competent in some ways but whose taste is off.
- We have some ad-hoc investments in evals to tell how we’re doing; but so far we’ve been working in the domain where you can manually inspect outputs and roughly tell what’s good.
- We’re experimenting with other ways to integrate research tools into a Centaur-like setup.
Overall it’s feeling like we’re moving into the foothills of the Centaur era. It feels like there’s valuable stuff to be had from systems at the moment, but it’s also kind of annoying as an experience, and this could stand to be improved. Maybe this is similar to where coding agents were in early 2025?
We do think there continue to be a good number of low-hanging fruit. The repeated experience is something like: see a way that things are dubious come up with ideas for a scaffold that would address that problem implement, and perhaps with some tweaking it generally works. We are capacity-constrained on trying more of these things.
Our vibe-based assessment is that improvements in these scaffolds are helping advance automation of this research at a significantly faster rate than background improvements to LLMs. Although this is a bit complicated by the fact that recent updates to LLMs have not always seemed to be improvements in our domain — in particular, we have tasks, which are not particular to scaffolds we have built up, where Opus 4.6 seems to outperform later Opus and Fable models.