Categories
Uncategorized

Architectural Gaps in Physical AI: When’s the Big Moment?

In some specific robotics use cases, the magic moment has clearly already happened. Autonomous driving had its moment in 2025. That’s why in 2026 you can ride a wheeled robot, a Waymo, all over SF with minimum hassle.

In Ukraine, autonomous drones now rule the sky (and the battlefield). At Amazon, more than 1 million robots help humans to move, package, sort, and pick, and we take it for granted that the stuff we ordered earlier today shows up at our doorstep the next morning. In China, humanoids now outrun Usain Bolt, the fastest-sprinting human.

But for many other use cases, it seems robots aren’t quite there yet. Is this just a question of needing to patiently wait until robots outperform humans, task by task, or are there deeper structural gaps? Based on growing real world deployments of our robots in SF and NYC, we see two gaps over and over again.

1/ We are using the wrong brain architecture

Engineering is about making choices. Remember the MIT saying – sleep, friends, grades – choose two? The same applies to robots. If you operate in the real world, you have a finite energy, constrained mass, and finite ability to dissipate heat. And if you want people to buy your robots, they need to be reasonably priced (unlike, say, anything ever manufactured by Boston Dynamics).

What this means for robot brains is that trying to build one all-capable model, the one to rule them all, is a fool’s errand. Sure you can prepend “universal”, “all powerful”, or “grand unified” to your approach, but that’s just a labeling exercise. The solution is to operate a smart router (with context) as the core of the brain, and then dynamically (re)optimize which data go to which models, as a function of the robot’s task and environment.

Need to balance on your toe? Focus all your resources on that stability task. Listening to a lonely elderly parent in the park? Well then use all your compute for empathetic personalized conversation and gesture generation. Helping a lost child find their parents at SFO? Now you care about path planning and scanning human (parent) faces for expression of concern.

As a capable robot, each of those tasks deserves your focus, and that’s why, as a robot, you should be constantly adjusting which models (and compute and actuators and sensors) you should be using in this very instant for optimal performance. At OpenMind, we are placing our bet on intelligent routing/attention as the foundational brain function of a robot, the goal being to use the best task-optimized model for the immediate task, whatever it may be.

2/ We left out the physics

As the name implies, physical AI (and robotics) is about physics, not just bits. Let’s say you want a machine to perform a specific task in the physical world, such as fly. Presumably, you would not start by collecting videos of humans flapping their arms. A more promising strategy is to focus on what actually constrains and enables that capability – Navier-Stokes allows you to compute what an efficient wing looks like. For human readers, don’t take this (too) personally, but the way in which you drive cars, fold t-shirts, do patient intake in the hospital ED, or pick up apples, may not be generally optimal given the underlying physics. The relevant question is – given a specific desired physical task/action, and given your current physical form (One arm? Two wings? Three feet? Two wheels?), and your current battery level, what is the optimal sequence of movements to accomplish that task? That’s a nonlinear optimization problem, not a “replay memorized motion sequences” task. The framing as nonlinear optimization gracefully accommodates changes in your current physical form and capabilities. If one of your arms falls off, no problem – you just change how you move, without more than a moment’s recalculation.

OpenMind’s definition of the Physical Turing Test

The physical Turing test (by Nvidia’s Jim Fan) builds on Alan Turing’s test and refers to the ability of a robot to operate in the real world without it being apparent to observing humans that they are seeing a robot rather than one of their own. This is one good way to define a discrete watershed, but it’s still very much centered on our self-reflection.

At OpenMind, we have a more general definition with three basic ingredients. We think that the “ChatGPT” moment for physical AI will be when robots have these three attributes:

1/ [Rapid learning of extreme dexterity] Given only the laws of physics and knowledge of its internal architecture and force generating capabilities, a robot must learn to move, run, jump, and manipulate diverse physical objects within 24 hours of being allowed to freely explore the real world. Further, the robot will gracefully and quickly overcome loss of one major sensor (e.g. a camera) or actuator (e.g. a finger, wheel, or foot) without external help.

2/ [Success with diverse cognitive, planning, and physical tasks] The robot will complete a succession of diverse tasks, encompassing verbal, emotional, reasoning, coordination, spatial navigation, manual/physical, and learning, without prior knowledge of which specific tasks it will face, and in which order.

3/ [Safety/Constitutional Robotics]. By virtue of its core design, the robot will reconcile the above capabilities with an internal constitution that minimizes risks of the technology for humans. You might well wonder how this can be possible – TLDR it’s probably not possible with our current approach to AI since we have very limited understanding of how frontier reasoning models actually work. It’s way too easy to hide behaviors inside of these models and they almost certainly will try to solve problems in ways that harm humans.

Categories
Uncategorized

Geography of Physical AI and Advanced Manufacturing

Many people associate AI with coding, or helping students with their homework, or making sudden progress on the Riemann Hypothesis. It’s less well appreciated that all those capabilities map directly to physical AI. The tools that we use daily, such as Fable and GPT-5.6 Sol, can also help a robot integrate information, build a contiguous sense of its environment, break hard tasks into manageable chunks, develop a plan of action, and invent new routes from A to B. There are some obvious implications (robots are waking up and can dance) but also second-order effects. Most of those are about maps and physics.

Manufacturing everywhere for everyone

Most obviously, if machines get smarter and more nimble, it’s easier for them to make stuff. Goods get cheaper and easier to find. You don’t have to haul everything across oceans. A factory no longer has to make one product a million times — it can be reconfigured continuously. Making small custom batches finally makes economic sense. Economies of scale still apply but accrue mostly on the digital side. One AI, deployed in a million places around the world, can bake fresh croissants in all of them. And every croissant teaches it something. Sense, decide, and act until the croissants are VERY flaky but still have chew. It doesn’t so much matter where it happens to be baking. Digital intelligence is not tied to a place and learns from every deployment. Of course, the butter, wheat, and electricity are all somewhere in particular.

Physical limits overtake labor costs

What happens next? The hourly cost of labor matters less and less. If the labor in a pair of jeans drops to a dollar, suddenly the cotton and the zipper matter. Countries whose high labor costs have long priced them out of making things will gradually become competitive again. Once you take labor out of the equation, what’s left are the laws of physics and the facts of geography. How big is the country? Can a robot arm move faster than the speed of sound? (Yes, but…) What’s in the ground, and where? Is there an ocean nearby? What’s the average temperature, and is there clean water? Are the roads any good? Are there fast trains? Where’s the nearest yttrium mine? Physics, chemistry, and geography, not wages, will shape who prospers. Iceland and Quebec have cheap electricity, Mountain Pass in California has rare earths, and high-labor-cost countries like Germany and Japan suddenly have a path back.

Energy, Information, and Atoms

Why does physics matter so much? Because you can’t beat it, regardless of how smart you are. However the future plays out, energy is conserved, entropy always increases, a robot arm cannot move faster than the speed of sound in steel (or carbon fiber), and information can only propagate at the speed of light. What are the basics? Cognition, compute, and physical action convert energy, typically electricity, to useful ideas and actions. So everything involved in concentrating, converting, and storing energy, and then moving it via power transmission grids becomes a decisive factor. Energy sources such as uranium, wind, solar, and gas contribute to energy independence and unlock powerful compute. You can have all the best chips in the world, but if you can’t power them, then they are all paperweights. Oh right – now you need the pure silicon to actually make those chips. And the lasers and masks and lenses and solvents. Critical materials are critical precisely because they are needed for batteries and chips and superconductors. Getting thirsty and hungry with all this reading? Yes, let’s add the biology too. All biological life needs clean water; we need the seeds that contain the blueprints to grow our food, and that in turn requires fertilizers. And so on.

A Cambrian explosion of robot form factors

If things are easier to make and you have all the atoms, then robots are easier to make and easier to customize. The shapes are going to run wild. Humanoids make sense for only one reason: our houses, cities, and workplaces were built for creatures shaped like us. But think about what an economy actually makes, the billions of items from bicycle spokes to pencils to raspberry Pop-Tarts. What fraction of those really requires two arms and ten fingers? Today there are maybe 300 distinct robot shapes — a four-wheeled robot that carries people (a car), a two-legged robot on wheels with no arms at all, big quadrupeds the Canadian interns ride to the park. Three hundred is nothing. Expect thousands of robot forms. When making things is easy and scaling is digital, the shape gets optimized for the job for everything except the edge cases. Yes sure, I have a Swiss Army knife at home, but it’s only one of hundreds of other tools. You can see this in the war in Ukraine – just in terms of quadcopters, there are at least 6 major design patterns, based on the drone’s specific purpose in a team or swarm of drones.

Curiosity and Nimbleness

There’s one more thing that decides who benefits and it’s the most geographic fact of all. It’s whether the people in a place are willing to actively engage with a technology whose innovation cycle is measured in weeks. Countries that are nimble and let these systems get fielded, tested, deployed, and folded into a fast-changing economy stand to gain enormously over the countries where everything stalls on bad information. Take robot cars. Waymos are clearly safer than human drivers. So why isn’t there a national effort to get them to everybody, when more than 35,000 people a year are killed in traffic accidents in the US alone? I’ve asked people this and the answers range from “it’s too different” to “what about fair wages for NYC taxi drivers”. Sure, those are real concerns. But I’d like to think that most people would put kids getting to school safely ahead of “that’s how we’ve always done things here.” Our habits won’t change the facts (or the physics). Surely, a nimble and creative society can find a path forward that maximizes (and shares) the benefits.

Categories
Uncategorized

Medtech Opportunities for 2024-2028

Given rapid advances in AI, it’s interesting to think about what they mean for US healthcare. Healthcare is a combination of (1) compute tasks (some version of “given their symptoms, how can we help this patient”), (2) simple interventions (“take drug X once a day”), (3) basic patient care in a hospital or other care setting – take vitals, start a line, take blood sample, provide food, and (4) highly specialized procedures e.g. trauma surgery or image guided cardiovascular procedures.

1. Compute/diagnostics/monitoring tasks All text/image/audio compute tasks are ripe for automation and there will be increasing pressure to automate them (higher hospital profits, faster, potential liability for failure to use best methods/tools, improved patient outcomes). Lobbying by professional societies will slow the rate of adoption, but that’s a loosing battle and hopefully the affected fields will rethink themselves rather than focus on delaying the inevitable.

2. Simple interventions can be scaled/automated by connecting an online pharmacy with the triage or diagnostic AI. This trend is already underway, with a human doctor in the loop largely for legal/regulatory reasons. Expect lobbying for transition to a fully-automated integration of #1 and #2 based on super-human performance of the triage/diagnostic AI – if the computer-only system is better than the one with the human doctor in the loop, presumably the FDA and other regulators will cave in at some point (although this may take many years).

3. Basic patient care is here to stay in some form – it consists of many different tasks that can be hard to automate, despite e.g. Japan’s long standing robotics R&D for supporting their aging population through care robots. The main trend will be to replace tasks requiring specialized skills (such as currently provided by registered nurses (RNs) and Physician Assistants (PAs)) with tasks that can be performed by lower paid staff (such as CNAs, certified nurse aides). For example, CNAs do not normally draw blood but hospital procedural changes could normalize that, after additional training for the CNAs. Per #1, any monitoring/diagnostic tasks currently provided by RNs and PAs will be increasingly automated, based on cost/speed/scaling/liability. In parallel, medtech tools and devices that radically simplify existing patient care procedures – such as placing IV lines, taking vitals, or drawing blood – will be championed by hospital CFOs.

4. Highly specialized procedures Upon first thought, it’s hard to imagine things like heart and brain surgery being massively impacted by AI and automation. Sure, minimally invasive procedures are growing and surgeons now frequently use robots, but the basics seem solid (highly trained humans use tools to help patients). The real threat to surgeons (and opportunity for medtech investors and innovators) are tools that allow complex procedures to be completely avoided or replaced by simple procedures. A great example is the replacement of amniocentesis or CVS by non-invasive prenatal testing (NIPT). Rather than first needing to manually collect and biopsy placental cells with a long needle or catheter, equivalent genetic information can be obtained through a simple blood draw and subsequent characterization of the circulating fetal DNA. NIPT is a win for (almost) everyone, since it reduces miscarriage risk to the mom and replaces a highly specialized procedure (done by an experienced doctor with ultrasound guidance) with a vastly simpler procedure (a basic blood draw) that can be performed by a phlebotomy technician. Presumably, startups focusing on down-skilling procedures/interventions currently requiring highly trained doctors, to services that can be performed by aides or technicians, will receive much investment.

TLDR

Trends and opportunities to look out for:

  1. AI decision support tech that reduces costs and improves patient outcomes. AI decision support is the precursor to replacing humans due to the need to first collect human vs. computer performance data for regulatory filings, scientific publications, and marketing materials; decision support tools are a natural entry point and necessary stepping stone to full automation.
  2. Medtech tools/devices that allow basic patient care to be primarily provided by aides and technicians rather than RNs and PAs, since monitoring/diagnostics will be increasingly provided by computers rather than humans.
  3. Medtech tools/devices that dramatically down-skill (or bypass) the need for procedures/interventions currently provided by highly trained professionals. For example, tests using circulating tumor DNA reduce the need for tumor biopsies and redirect payments from doctors and hospitals to genomics/diagnostics tech companies.

Categories
Uncategorized

The End of Human Radiology in the US

Based on a slew of papers over the last few years, as summarized in a Stanford seminar by Bram van Ginneken (“Why AI Should Replace Radiologists”, Nov. 15, 2023), AIs now consistently outperform the best human radiologists across most image/diagnostic tasks. His hope is that human radiologists will lead a transition to AI-based radiology, to improve health outcomes and accessibility while reducing costs. The writing is on the wall and some radiology residents at Stanford are dropping out to start AI-enabled radiology companies, presumably reflecting their agreement with van Ginneken. However, the US healthcare system is a complex for-profit system, unlike European outcomes-focused health systems, and it’s interesting to wonder what the endgame for human radiology will look like in the US. Let’s consider some of the stakeholders:

1/ Patients and Patient Advocacy Groups

Patients rightfully assume they are getting the best care. It will be hard for patients to understand why their images are being read by humans, despite clear scientific evidence that replacing humans with computers would benefit their health, such as by reducing false positives, reducing wait times, and reducing false negatives. At some point, national advocacy groups like the National Breast Cancer Foundation and the National Breast Cancer Coalition will start to ask hard questions about patient benefit and which methods – people or computers – should be used to screen mammograms, just to give one example.

2/ Malpractice Insurance and Trial Lawyers

For an insurer, it’s presumably hard to justify providing reasonably-priced malpractice coverage when a field persists, against scientific evidence, in using antiquated procedures, such as human-based radiology. This is not yet an issue because doctors in the US can only be sued for failing to provide the ‘standard of care’, which is still based on humans. So, as strange as this sounds for a patient, from a liability perspective, it doesn’t matter that there are better technologies out there, since (legally speaking) radiologists do not promise to provide the best care; rather they promise to (and are held accountable for) providing the ‘standard of care’. However, at some point, a smart trial lawyer will connect the dots, see an opportunity, and work with national advocacy groups and affected patients to drive change.

3/ Human Radiologists

Just to be clear, the radiologists I know are awesome people and doctors – sharp, dedicated, passionate, and wanting the best for their patients. What does AI do to their jobs and their job satisfaction? The key issue is liability – imagine the hospital introduces an AI radiology assistant to provide decision support and imagine further that the AI assistant has been demonstrated to outperform even the best humans. Currently, a human doctor must review computer generated findings/suggestions, and can then either (1) accept the computer’s suggestion and sign the note, or (2) disagree with the AI and manually enter an alternative (which, on average, will be worse than what the computer concluded). Very soon, choosing to disagree with a computer known to outperform humans will prompt a call by the hospital’s office of risk management, who are trying to protect the hospital from lawsuits. So then, playing this forward, the human radiologist can either: (1) agree with the computer and click “concur and sign”, or (2) disagree with the computer, write a 4 page memo to risk management to justify the ‘deviation’, and hope they were right. On average, the human radiologist will be wrong, so that strategy is a losing one for all stakeholders (doctors, patients, risk management lawyers, hospital CFO, insurance companies). Rather, the optimal long-term strategy will be to always agree with the quantifiably better computer. At that point, the human radiologist will wonder if the (minimally) 8 years of training, the residency, and the nights were worth it, if their job consists of clicking “concur and sign” 38+ times an hour while sitting at their PACS station.

4/ Tech enabled healthcare competitors

Technology companies with long term strategic interest in healthcare and extensive capabilities in AI, computer vision, and healthcare backends, most notably Amazon, have no need to protect traditional workflows or professions. Presumably, they will seek to optimize overall profit. This is an “all upside” scenario for them – they can offer a better product at lower cost to tens of millions of patients. Certainly, they are regulatory and political barriers to replacing humans across radiology, but these barriers will gradually fall in the face of tech industry lobbying, patient advocates demanding the best care (not just the ‘standard of care’), the difficulty of convincing medical students to chose radiology, and the (increasing) cost of malpractice insurance for radiology.

Outlook

Based on the confluence of the above, my hope is that US radiology will embrace AI not as an existential threat, but as the foundation of modern, reliable, and scalable healthcare. Who are the dedicated, passionate, and smart doctors with excellent quantitative and computer skills who will help to build a healthcare system that provides better care at lower cost? If human radiologists accept that challenge, their profession is secure and they will continue to be at the center of figuring out what’s wrong with people and helping them for a long time to come.