
Jisoo · November 19, 2025
How the age of automation comes about, and what humans become
"The danger is not that robots disobey orders. It is that they obey them too perfectly."
New research pouring out of the frontier every day, billions of dollars in investment, and intensifying technological competition between the United States and China. Where all of these converge lies a vast dream: to remake a new physical world.
After LLMs, robotics is being pointed to as the "next explosive growth field." In particular, as Chinese manufacturing is perceived as an existential threat to the United States, competition in robotics is heating up more fiercely.
Robotics is one of the hardest areas in AI, but ironically, several new AI strategies that have recently emerged are beginning to show a concrete and clear path toward EGI*.
*EGI (Embodied General Intelligence): an "embodied general-intelligence robot" that can perform a wide range of tasks in the real world like a human. **General: a broad range of capabilities that is not fixed to a single specific task, but can handle multiple situations and multiple kinds of work on its own.
I have observed this shift up close. Conversations with frontier researchers, firsthand exposure to robot companies such as Tesla's Optimus and Dyna, and my own technical intuition all point me to one conclusion: by 2045, "robots that reason and act autonomously" will account for half of global GDP. This essay is a step-by-step scenario for how that future becomes real.
2023-2025: The Dawn
Throughout Android Dreams, I use fictional names that stand in for specific types of companies rather than real companies. For example, OpenBrain represents "an American AI lab," while Unioak represents "a Chinese humanoid company."
OpenBrain's LLM shakes the world. Riding on OpenBrain's success, several robotics companies go on to raise rounds of $300 million (about KRW 400 billion). One of them is Waytek, a new American robotics startup. They set out to build an "LLM for robots" in hopes of recreating OpenBrain's success in robotics.
Why did LLMs succeed? Pre-training* succeeded because the scaling hypothesis was right. * Pre-training: the foundational stage in which an AI model first learns language, knowledge, and patterns of the world from large-scale data. This is what the "P" in GPT (Pre-trained) means. In other words, the hypothesis is that as model size and data volume increase, human-like capabilities emerge. Models such as GPT-4 gained near-human language ability and reasoning power by training on trillions of parameters and terabytes of data. This enormous scale was possible because three things came together at once: parallel computing infrastructure (NVIDIA GPUs and CUDA), an efficient architecture (Transformer), and vast internet-scale text data. The reason post-training* of reasoning-specialized models such as o3 could succeed was that OpenAI scaled reinforcement learning** to very large batch sizes by running tens of thousands of GPUs at the same time. Parallel computing was central here as well, and it was also possible because reinforcement-learning algorithms such as GRPO*** were structurally simple yet stable even at massive scale. * Post-training: the additional training stage after pre-training that refines and strengthens a model's abilities in a desired direction. ** Reinforcement learning: a learning method in which a model tries actions on its own and improves performance through rewards based on the results (success or failure). *** GRPO (Group Relative Policy Optimization): a simple, stable algorithm designed so reinforcement learning still works well when run simultaneously across an enormous number of GPUs. For robotics to succeed, data cannot be the bottleneck. Even if collecting data is difficult, it is worth spending more compute to train on data that is more diverse and information-dense. (...This will return later as an important foreshadowing.)
OpenBrain's pre-training worked because it had an enormous amount of text data, effectively scraped from the entire internet. But robotics has no equivalent "robot internet." The action datasets that robots need for training are, even at their largest, less than 0.01% the size of OpenBrain's LLM datasets.
OpenBrain's post-training was possible thanks to reinforcement learning. The model learns by attempting tasks such as complex math problems billions of times and learning from success and failure. But this approach does not work for robots. The real world is too slow, and physical and temporal constraints are too large for robots to interact at that scale.
So the core of all robotics research now converges on a single question: "How can we scale pre-training and post-training for robots too?"
Waytek's Teleoperation Works
From 2023 to 2025, Waytek chooses a strategy of collecting teleoperation* data at massive scale to solve the shortage of pre-training data. It records the process of humans directly controlling robots, then trains a VLA** model to imitate it. The structure is very similar to an LLM.
* Teleoperation: direct remote control of a robot by a human. ** VLA (Vision-Language-Action): an AI architecture that processes vision, language, and action together so a robot can judge and move on its own. For example, it understands "the screen seen by the camera" (visual input) and "a human instruction" (language), then decides "how to act" (action).
The early experimental results are far better than expected. Demos show robots reliably doing tasks such as sorting laundry, making sandwiches, folding shirts, and sorting packages. Some companies go all in on teleoperation, just like Waytek. But industry opinion splits cleanly in two.
One side believes that "teleoperation will ultimately lead to generalized robotics," while the other counters that "teleoperation cannot scale and is not a fundamental solution."
But both sides are missing the point. The tasks AI robots must perform fall broadly into two categories - "narrow tasks" and "general tasks" - and the pre-training and post-training methods required for the two are completely different.
Narrow tasks: simple tasks with some degree of variability, such as sorting packages or folding clothes. General tasks: tasks that require human-level complexity and broad judgment, such as service work, construction, healthcare, education, and a wide range of household tasks.
Teleoperation is highly effective for automating narrow tasks, but it has structural limits when expanded to general tasks and eventually runs into a dead end.
Across the Pacific, China's strength had already grown large enough to pressure the United States.
In Shenzhen, the 100th manufacturing plant is converted into a fully autonomous dark factory*. China already has twice the United States' energy production and ten times its manufacturing capacity, and it is charging relentlessly toward full automation.
* Dark factory: a fully autonomous factory that operates around the clock with little or no human presence. Since people are not needed, there is no reason to keep the lights on, hence "dark factory."
In 2024, Unioak, a Chinese humanoid company best known for robot dogs, begins seriously bringing humanoid robots to market. By 2025, these robots are dancing, presenting flashy demos, and overwhelming the world.
American investors discuss China's dark factories with fear and anxiety, producing a flood of reports and memos titled "America Must Reshore Manufacturing."
People worry about China and fear that AI will take their jobs. But... no one yet truly sees what is coming.
2026-2030: The Vertical Era
Vertical: an approach that goes deep into one industry or one task.
In 2026, the first moment arrives when an AI-controlled robot replaces a human job. Waytek, America's leading vertical robotics company, combines cheap Chinese hardware with teleoperation-based data collection and reaches 80% of human performance on simple but repetitive tasks such as package sorting.
Instead of "selling" these robots, Waytek deploys them to package sorting centers as a form of labor. Center operators, eager to accept the coming wave of automation, willingly approve the robots.
The Rise of Waytek Clones
Pessimists point out that Waytek's robots do not generalize, cannot reason deeply, and are far from "real intelligence." But one fact matters: for the first time in the history of AI robotics, "actually usable robots" have appeared.
Waytek raises another roughly $400 million (about KRW 560 billion), hires operators, engineers, and data collectors in large numbers, and aggressively expands robot deployment. Meanwhile, other startups each raise around $50 million (about KRW 70 billion) and rush to copy Waytek's approach across different vertical industries.
Waytek does not sell robots. Instead, it uses a Robots-as-a-Service (RaaS) model. In other words, rather than "selling" a robot, it charges labor fees for each hour the robot works. This model means customers barely have to worry about the technical and operational complexity of adoption.
As Waytek expands into multiple package sorting facilities, it grows annual recurring revenue (ARR) to $100 million (about KRW 140 billion), even though it has automated only a single task. Watching Waytek's success, American humanoid robotics companies such as Noumena realize something important.
Expensive humanoid robots can never compete with cheap, mass-producible Chinese hardware on simple tasks. To succeed, they must therefore target areas Waytek cannot even reach: complex "general tasks" where humanoids can justify being four times more expensive.
Waytek Scales to $10 Billion in Annual Revenue
As it grows, Waytek gradually shifts its data collection method from teleoperation toward exoskeletons*. Teleoperation equipment is expensive and difficult to operate, while exoskeletons are much cheaper and can record human-level dexterous movement as-is, allowing the company to secure higher-quality data more reliably.
* Exoskeleton: an external skeletal device worn outside the body, used to assist the user's strength or capture fine movements directly.
The tradeoff between teleoperation and exoskeletons. Physical Intelligence's Pi0 and 0.5 studies showed that combining a VLM* backbone with teleoperation data enables fairly strong generalization within the same domain and allows robots to perform basic tasks reliably. * VLM (Vision-Language Model): a model that understands visual input and language at the same time. For example, if someone says, "Please pick up the yellow cup at the back right of the desk," it analyzes the camera image - identifying the desk, right side, back, and the yellow cup's location - understands the language - interpreting what the user wants and how - then combines the two to judge, "Where should the robot arm move to pick up the yellow cup?" The concept extended to decide actions and actually move is the VLA (Vision-Language-Action) described above.
Dyna Robotics' Dynamism-1 adds several technical advances and proves that a teleoperation-based model alone can fold clothes with 99.99% task reliability, for 24 hours, at more than 50% of human speed.
The key insight here is very simple:
You do not need AGI to automate extremely reliable, narrowly scoped tasks.
FigureAI's Helix also showed that teleoperation alone can achieve human-level dexterity and high reliability in logistics.