Blogging my mind

Do We Really Need More Data, or Just a Smarter Way to Handle It?

Lately, I've been thinking a lot about the direction we're heading with foundation models in biology.

If you look at the current landscape, the playbook seems to be "bigger is always better." In Spatial Transcriptomics (ST), the status quo is to train massive models from scratch on millions of cells. But if we're being honest, a lot of these "revolutionary new models" feel less like breakthroughs in AI engineering and more like an excuse for big players to flex their proprietary datasets. The transformer architectures stay mostly the same; the real moat is just having the deep pockets to buy, scrape, or gatekeep massive amounts of data and burn endless compute.

It turns the field into a game of who has the biggest budget, rather than who is making the smartest engineering choices. It completely prices out smaller, independent labs.

That's exactly why we wanted to take a different approach. What if the real breakthrough isn't accumulating bigger data, but simply being smarter with what we already have?

The Idea Behind DRIFT

I'm incredibly proud to share DRIFT (Diffusion-based Representation Integration for Foundation Models), our new framework just published as an ISMB 2026 Proceeding in Bioinformatics!

Instead of playing the game of data-monopoly and reinventing the wheel, we built DRIFT as a plug-and-play module. It incorporates spatial context and drastically denoises data using mathematically principled heat diffusion through graphs.

The part I'm most excited about? It requires absolutely zero retraining.

It plugs right into existing, frozen scRNA-seq foundation models. Instead of spending weeks and thousands of dollars training a model from scratch, DRIFT gives you deep spatial awareness instantly. We wanted to focus on the how (the elegant math) rather than the how much (throwing raw computing power at the problem).

Why This Matters to Me (and Our Field)

When we benchmarked DRIFT, the results genuinely blew me away:

At the end of the day, more isn't always better. Real innovation shouldn't belong exclusively to whoever has the biggest server farm. Sometimes, it's just about making the open data you already have work a whole lot harder for you.

Check Out the Paper

If you want to dive into the math, look at the benchmarks, or see how we pulled it off, you can read our full paper here:

Diffusion-based representation integration for foundation models improves spatial transcriptomics analysis