Article
AlphaFold 3 and GPT-4o Solve Different AI Problems
AlphaFold 3 helps researchers form structural biology hypotheses. GPT-4o reduces interface friction. Each needs a different standard of evidence.
AlphaFold 3 and GPT-4o Solve Different Kinds of Friction
AlphaFold 3 and GPT-4o were announced within days of each other in May 2024, but they should not be treated as two examples of the same breakthrough.
AlphaFold 3 reduces friction in forming hypotheses about the structures and interactions of biological molecules. GPT-4o was introduced as a model that could work across text, vision, and audio with a more direct conversational interface. One belongs inside a scientific process that still requires experiments and qualified interpretation. The other belongs inside a human workflow that must earn trust through reliability, privacy, consent, and useful recovery when it is wrong.
The shared lesson is not that AI has replaced an old way of working. It is that a model becomes valuable when its output fits a real process and survives the right test.
AlphaFold 3 widened the structure-prediction problem
The AlphaFold 3 paper describes a model for predicting the joint structures of complexes that can include proteins, nucleic acids, small molecules, ions, and modified residues.
That scope matters. Earlier protein-structure work helped researchers reason about the shape of a protein. AlphaFold 3 was designed to model a broader set of molecular interactions, including some of the relationships that matter when researchers investigate how a candidate molecule might bind to a target.
The paper reported higher accuracy than several specialized methods across most of the tested categories. Google DeepMind summarized one group of results as an improvement of at least 50 percent for interactions between proteins and other molecule types. That is a vendor summary of specified benchmarks, not a universal improvement across biology.
The narrower statement is also the stronger one. AlphaFold 3 produced competitive or better results on defined structure-prediction tasks. It did not prove that machine learning had replaced physics, molecular dynamics, or laboratory research.
A structural prediction is a hypothesis, not a completed experiment
The recovered episode outline says AlphaFold 3 can predict structures and dynamics in hours and could reduce drug-discovery costs by 50 to 70 percent. The reviewed primary sources do not establish those claims.
AlphaFold 3 predicts structures. A static structural prediction is not the same thing as simulating molecular motion over time. It can help a scientist decide what to investigate, but it cannot establish biological function, safety, efficacy, or clinical value on its own.
Google DeepMind's own launch material describes AlphaFold Server as a way for scientists to generate predictions and form hypotheses that can be tested in the laboratory. That sequence is important. Prediction can make the search more informed. Experimental validation remains part of the work.
The same boundary applies to drug discovery. A better model may improve one stage of a long process, but the episode's estimates about years saved, costs removed, and billions of dollars created were not tied to a reproducible study. They should remain in the raw source as recording-era claims, not move into the canonical article as facts.
GPT-4o changed the shape of the interface
OpenAI introduced GPT-4o on May 13, 2024. The "o" stood for omni, reflecting a model designed to accept combinations of text, audio, image, and video and generate text, audio, and image output.
The architectural point was as important as the demo. OpenAI said the model was trained end to end across those modalities instead of routing voice through separate speech-recognition and speech-generation systems. The company reported audio response times as low as 232 milliseconds and an average of 320 milliseconds in its testing.
Those figures describe a controlled vendor result. They do not guarantee the latency of every device, network, product tier, or conversation.
The launch boundary was easy to miss. OpenAI initially released text and image input with text output in ChatGPT and the API. The richer voice and video experiences shown in the demonstration were scheduled for staged release. A launch demo can show the intended direction before the full experience is generally available.
A natural interface creates a new class of mistakes
Voice and vision can reduce the work required to explain a problem. A user can show the model a screen, point a camera at an object, or speak without translating the situation into a carefully written prompt.
That convenience can also make an answer feel more grounded than it is. A fluent voice can still misidentify an object. A model can infer more than the user intended from a face, voice, room, or screen. Background audio can affect performance. Stored media and conversation history can create privacy questions that do not exist in the same form in a short text exchange.
OpenAI's GPT-4o system card discusses risks involving unauthorized voice generation, speaker identification, accent performance, sensitive-trait inference, audio robustness, and misinformation. Those are not side issues. They are part of the product test.
Scientific AI and assistant AI need different scorecards
AlphaFold 3 should be evaluated against a defined scientific task. That means identifying the target, comparator, dataset, benchmark, confidence measure, known failure modes, and experimental validation that follows.
GPT-4o should be evaluated in the workflow where someone intends to use it. That means testing task accuracy, latency, privacy, consent, accessibility, escalation, and what happens after an incorrect answer.
Both products can produce an impressive demonstration. Neither demonstration establishes complete value.
This is the viewpoint I would carry forward from the episode: match the evidence to the job. A molecular-structure model is useful when it improves a scientific search without disguising uncertainty. A multimodal assistant is useful when it makes interaction easier without turning fluency into false confidence.
Continue the conversation
The episode captures the excitement around two May 2024 launches and the different futures they suggested. Listen on Spotify or watch on YouTube.
Sources and editorial notes
This article uses the preserved [[E16 - Transcript - Google Drive recovered|raw production outline]], the AlphaFold 3 paper in Nature, Google DeepMind's launch explanation, OpenAI's GPT-4o launch, and the GPT-4o system card. It is an editorial explainer, not scientific, medical, clinical, drug-development, privacy, procurement, or investment advice.
Sources
Follow the evidence.
- Google DeepMind about pagedeepmind.google
- NIST AI RMF Measure guidanceairc.nist.gov
- Google DeepMind AlphaFold 3 launchblog.google
- Google DeepMind AlphaFold pagedeepmind.google
- Isomorphic Labs company siteisomorphiclabs.com
- OpenAI about pageopenai.com
- Current GPT-4o API documentationdevelopers.openai.com
- GPT-4o system cardcdn.openai.com
- FDA machine-learning transparency principlesfda.gov
- AlphaFold 3 papernature.com
- GPT-4o ChatGPT retirementopenai.com
- GPT-4o launchopenai.com