Rather: Someone should make the millennium problems for human alignment so the top researchers will actually try to solve that problem. :-)
It's a good list, but "alignment" is much more targetted. I attended a workshop with researchers from OpenAI and Anthropic also present, where the objective was to hash out what the problems in alignment even are. This, it turned out, was extraordinarily difficult.
The problems with alignment are to do with how we define and search for this vague notion when it is mathematically ill-posed at present.
I guess to me the grandiosity here feels fake and unearned - like they vibe slopped it together without much input in the way of deep thought or expertise.
For example, synthesize arbitrary dna sequences > 3000 nt in length with error rate < 0.001 at > 95% purity. Easily verifiable, beyond the frontier, and generally useful.
I think some of the items on the list are also way less likely to be worked on (like cryogenics) than others, because they really require some heavy infrastructure to iterate. My money for which of these gets cracked with help of an AI would be rubisco, that can be quite effectively worked on in a closed loop lab fashion with a standard lab or CRO.
Biology is a really really big field. Many problems in biology are that of math and coding. Just as a simple example look at the golden ratio in living organisms. Biology follows a lot of different algorithms because they are energy efficient and come with massive secondary benefits.
What is this example meant to illustrate?
>The golden ratio (≈ 1.618) and its related golden angle (137.5°) show up in nature because they maximize space, exposure, and flow.
The golden ratio is an algorithm, but it is also a natural law that systems of many different scales follow. Wherever we look we find examples of it. Without this algorithm the world is a much harder place to explain.
What many scientists starting to look for is these algorithms we have not identified in natural systems as an explanatory means of their workings. Also because they are algorithms they are things we can put into computer systems to increase efficiency, or may naturally emerge from evolutionary learning systems like AI in what we call emegence.
SMEs please correct the record if I'm mistaken.
There’s also this study [0] I read recently, along a different path of using known growth factors and proteins to kick off regeneration.
https://www.biorxiv.org/content/10.64898/2026.08.05.742898v1...
"Extracellular matrix particle treatment induces digit regeneration in soft-tissue preserved amputation (SPA) model of adult mice"
Also his work on algorithmic intelligence in cells and the mathematical foundation of intelligence will change many fields, including AI, if it can be empirically tested.
None of these is likely the root to the wide array of chronic illnesses we can't treat and those seem like the next step to really put a lot of research into given how many people suffer from them and where we are today they seem achievable with the right investment.
They all use DNA.
But I agree that nobody has shown in the lab how any of this leads to a genetic system that can self-reproduce reliably and assemble biomolecules/metabolites. This is a missing link. Just showing how RNA or amino acids arise, says nothing about any genetic system. They can not even answer whether RNA or DNA viruses existed first; and whether these existed before organisms/cells did.
Seems a bit of an ambitious goal for a language model though! There are presumably dozens of steps before you get to an RNA-based full-blown cell, and it's only been very recently that humans have managed to make an artificial self-replicating "cell" (container) of any type in the lab.
I really really recommend looking up Michael Levins bio work on intelligence at the cellular and organ level.
His most recent work is starting to touch on some pretty far out concepts. And I should note, that he does not state they are true or not, but they are attempting to make empirical methods on testing them.
One of the concepts they are looking at is "math as an actual thing" and that some simple algorithms have new deeper behaviors that we're just finding now. The hypothesis is that over the eons life has probed mathematics at scales humanity cannot even begin to imagine and is exploiting this deep algorithmic efficiency in accomplishing any number of tasks. These algorithms could help explain the missing links.
I'd suggest starting with some of his older work to see they are a credentialed research scientist and not making up 'woo' whole cloth.
He seems to be essentially the only biologist that hackernews knows and he always (always!) comes up in biology discussion. I’m wondering how this particular situation arose.
Dogma in science tends to blind people, and when a field starts slowing down and running into things we cannot address then we have to start looking at what we should be questioning.
Wait, isn’t this an already solved problem? I remember in the 90s school books (material written in the 80s-70s likely) it was presented as “so and so did electric discharges in an atmospheric gas mixtures and they got organic molecules” therefore the origins of life is solved. I remember thinking how cool that was. I guess the “therefore” step was sort of a lie if it’s still an open problem…
We don't even know if Miller or Urey themselves would have (or did) frame their results that way. The honest framing is 'hey this is one way to form amino acids, and maybe that pathway is reflective of the origin of life, but that supposition remains conjecture on our part.'
still not all the way to something alive-like
I recently watched this video on the topic https://www.youtube.com/watch?v=OsXGgrXwllc
Arguably, the origin of life is a question for all time -- no way that a VC webpage is going to change whether someone takes that problem on, and it will be commercialized with big-name capitalists instead of 'FutureHouse' should it ever be developed in the lab.
> Cryopreservation + Demonstrate the ability to cryopreserve and recover live wild-type mice with high viability. Specifically, demonstrate the reversible cryopreservation of live, intact, wild-type adult mice in a whole-body frozen or vitrified state. The mice must remain frozen or vitrified for at least 24 hours, must be recovered with >99% viability, and must not suffer any permanent organ damage or bodily harm. Somatic genetic engineering is discouraged but permitted. All experiments must be conducted with ethics approval.
This has been done in the 50ies, freezing and thawing mice with microwaves. IIRC the recovery rate was about 74% with no observed side effects. The research was given up on because in this specific case a mouse is a bad biological model. Thawing agents must scale cubically with size. The author says that at approximately the size of a cat you can’t quickly enough evaporate all the agent because the energy required will simply burn the tissue.
TLDR: it’s possible, it’s been done, it could be perfected if necessary, it does NOT scale beyond mice
Look at what's happening in the math community- rather than excitedly embracing the power of AI's ability to generate new proofs, they are screaming for it to slow down (and quite a few want it simply stopped). Many scientists, even in ostensibly more real-world fields are broadly cut from the same cloth.
A much more practical reason we're not yet seeing a lot of headline scientific mathematical breakthroughs is just that it is ungodly expensive! e.g. The Navier-Stokes result cost around $20M at API prices, and academics just don't have that kind of money to spend. You'll see more mathematical and scientific results from practitioners when either the cost of compute needed for these sort of brute force results is more in line with the size of academic grants, and/or the AI companies donate more compute to the scientific community.
There are different reactions from different mathematicians of course - Terrance Tao vs Cedric Villani, and no doubt a lot of shock at the speed of advance, but it seems the reasoned complaint why they don't want the AI companies themselves working on these problems is because the outcome is not the same - you get a result that in of itself may have been suspected or useless (Navier Stokes), but no write up of any new math or insights that were developed along the way, which is the real reason mathematics and people like Erdos pushed these famous problems in the first place - because they were expected to yield interesting mathematics, just as years of work on FLT had done. Imagine if instead of Wiles's work, and all that had gone before him, all we had was a $20M compute bill, hundreds of pages of impenetrable math, and a billion lines of Lean proving it was true?!
There is a difference between what's easy/hard for a human vs computer, and LLMs haven't changed that. You might expect a computer to be good at tasks requiring prodigious memory and compute, and it turns out that some of these long-standing math problems are of that nature - not requiring new breakthroughs but rather just massive exploration of what is already known and what they were trained on.
There will no doubt be more math results like this, but presumably also ones that are "hard for a human, easy for a computer", requiring massive search (e.g. find an example/counter-example cf Navier-Stokes & Jacobian conjecture) rather than creativity.
Demonstrate the ability to cryopreserve and recover live wild-type mice with high viability.
> 06 - Somatic limb regeneration
Demonstrate the ability to regenerate lost limbs in adult wild-type mice.
Interesting, but looks like these problems are proposed in September 2026, unlike original 7 Millennium Prize Problems of Maths
I think GP was saying the proof of the pudding is in the tasting: the longer a problem has provably resisted resolution the more difficult it is considered...
bombastically decorating a problem as equivalently difficult does not make it so.
This list of "Millenium Problems for Biology" contains such brainfart level "analogies" that there the list will be ridiculed, for the question / challenge itself displays a lack of understanding of the subject in question. Science is also asking the right questions.
Consider for example:
> 10. Protein Amplification Chain Reaction:
> Demonstrate exponential amplification of arbitrary peptide substrates.
> Specifically, demonstrate input-protein-dependent synthesis of new, full-length, sequence-faithful covalent polypeptide copies from amino-acid monomers without a nucleic-acid template or preformed cognate scaffold, in a single pot reaction. For the challenge to be considered complete, at least 100 random peptide sequences of at least 50 amino acids each must be preregistered, synthesized, and pooled. It must then be shown that the abundance of these peptides in solution can be amplified at least 1000x with at least 90% sequence accuracy on a per-residue basis. Reasonable modifications may be added to the peptide sequences to facilitate post-amplification analysis if necessary, provided they are not active in the amplification. Methods that rely on explicit sequencing of the peptide are not permitted. Methods that rely on reverse translation to generate a nucleic acid intermediate are not permitted, because they are duplicative with a separate Millennium Problem.
The analogy is very clear: to amplify DNA or RNA one uses PCR, basically throw the desired product in a cauldron with monomer building blocks, then by repeated heating and cooling the lone monomers find their permitted locations on a complementary pre-existing strand, and form the new polymer strand.
So it seems natural to ask for a generalization to protein polymers, except every biologist or chemist knows its nonsense: proteins don't have a complementary strand! You can't demand chemistry or physics to magically copy without a complementary template!
You may ask "but if that were true, how can we already have PCR for RNA?"
Well pretty simple: while this is done routinely, its only possible indirectly: convert the RNA to double-strand DNA, use PCR on this DNA and then convert the amplified DNA back to RNA!
The demand to not involve sequencing or the hypothetical reverse translatase from one of the other problem statements turns this one into a non-existence theorem, but the challenge doesn't describe a winner for demonstrating its impossibility!
I assure you that any chemist or biologist being asked why we dont have PCR for protein, will understand your lack of knowledge, and explain how PCR works, so that you understand that PCR was only possible because of the complementary strand!
This list will be ridiculed for being not even wrong.
At least those things cannot be solved by just burning GPT tokens.
I thought we already managed to successfully cryopreserve and recover small rodents like hamsters in the 50s.
Edit: See https://en.wikipedia.org/wiki/Cryopreservation#History
Also if you think humans would ever become a space faring species, would it really make sense to stick to our carbon based biology or should we invest in transforming into silicon based beings
Edit: A true expert would never define a problem in such a sloppy way: "Specifically, the protein must convert N₂ to ammonia at rates that are at least of a similar order of magnitude to the rates of naturally occurring proteins, and must fall well below the sequence- and structure-similarity thresholds relative to all known nitrogenase and nitrogenase-like proteins. The protein may be designed de novo, discovered in nature, or engineered or evolved from naturally occurring starting points."
It seems that understanding biology could be characterized as trying to figure out "how did nature engineer this".
For that matter isn't all of science this way?
You keep on making assertions - perhaps you could back it up with some explanation of why you regard this as a more fundamental discovery (or however you would like to characterize it)?
Reverse translatase and protein amplification in particular.
How the fuck do you plan on selectively priming protein amplification. If you know ANY protein chemistry, you will know "the juice is not worth the squeeze" -- how would I exponentially amplify a protein? I'd do mass spec proteomics, synthesize the DNA, and express it.
Simply amazing that electrofixation is not on the list.
Cryopreservation, even though I don't care much for it.
The rubisco one is sort of not dumb, but if you actually care about carbon fixation you'd just not bother using rubisco at all instead.
Programmable Proteases is fine.
Somatic regeneration is fine.
Now you can enjoy one man-week of rage across the remainder of your life as you won't be able to unsee it. You're welcome
That there are no more efficient variations of it in nature just tells us that the local minimum is really deep and that natural evolution, as it is, can’t produce anything better, not even with a billion years and 10^30 organisms serving as a “brute force lab”. It’s also the kind of problem an AGI system would tackle for purely ideological reasons, i.e. to prove that it is superior to nature.
> Specifically, demonstrate the unassisted emergence of self replicating RNA- and protein-based cells from a plausible primordial soup with a plausible energy source. A “cell” may be any compartment with a defined boundary.
So, right now, nobody can explain how life originated. Proving that aminoacids, DNA or RNA form, does NOT mean that this is equal to life. This is a problem that the whole field has - it still can not explain how life originated. Showing that individual components can arise spontaneously, is not the same as showing e. g. how a cell formed and so forth. Where does the encoding problem fit into any of that, for instance? You need to prove how you can assemble systems. Just having a working ribosome does not connect it to DNA as a genetic backup system; and RNA tends to be unstable. Even when you have assumptions how this works together, you need to prove that this is how things originate(d). Nobody has done so since decades. It is an unsolved problem.