I like a vivid analogy that helps capture the scale of the improbable chain of events that led to the emergence of life. Say we take the complete text of Isaac Newton’s Philosophiæ Naturalis Principia Mathematica. We cut it into individual characters, scatter them across a vast field, and then wait for the wind to rearrange them back into the correct order, so that one of the most important scientific works in human history reappears before us
Sometimes it’s worth pausing to consider how many conditions had to align for life to appear in the universe at all. With that thought comes the realization of just how lucky each of us is, and what lottery we’ve already won
And then, after a mere few billion years of evolution, that living matter gains consciousness and intelligence, begins to explore the world around it, and eventually reaches the point of creating another mind - an artificial one
He slept badly that night, even though he’d followed his entire pre-sleep ritual: no gadgets or news for two hours beforehand, just a light snack and a cup of warm herbal tea on the terrace. But his thoughts were spinning so fast he could feel them speeding up his heartbeat. On the outskirts of the city where he lived, the streetlights didn’t wash out the sky, and the stars were beautifully visible. But this time, not even the stars or thoughts of space could calm him. A few more deep breaths, one last sip of the still-warm tea, and it was time for bed - he’d done everything within his power today
A sense of major change hung in the air, and it was felt even more strongly in this particular place, because this was where those changes were being born right now. The city had lived through something similar before, back when mobile technology arrived - Uber, Twitter, Instagram, and Facebook had come crashing into everyday life, and the whole city felt like it was taking off. It seemed back then like every other company was going public or being sold for a billion, plenty of people were becoming millionaires, and the city was coming alive, soaked in a new kind of energy. And now that same feeling was back, only dozens of times stronger than ever before
Miles - the man riding in an electric car on autopilot toward the research campus - couldn’t wait to get to the office and start work. He knew they were on the verge of an incredible discovery, possibly the biggest in human history. He could have switched off the autopilot and driven much faster, but he knew firsthand where that kind of speeding led. Besides, his mind wasn’t on the road right now anyway - it was on the project that, today, might finally be brought to its conclusion
So much time he’d spent with the team, building and tuning everything. So much energy he’d poured into negotiations with investors, government officials, contractors, and every other person and organization involved. And all of it had been necessary just to reach this exact moment
As he pulled up to the office, he waved to the security guard. The guard waved cheerfully back. I wonder, he thought, does he have any idea how important today is? And he decided the answer was no - that kind of calm, cheerful expression could only belong to someone who wasn’t planning to change the entire world today
On the walk from the parking lot to the office, he kept moving on autopilot too, greeting people automatically. But if anyone had actually asked him something, he’d have had to surface from the depths of thought he was currently lost in. And he probably would have started saying goodbye instead of hello, because the feeling that the world was about to change fundamentally after today’s launch wouldn’t let go of him - he was already halfway to saying farewell to this reality, as if today could be the last day of it. And no, that wasn’t because they’d prepared something terrifying or sinister, but because absolutely everything people were used to was about to change. Absolutely everything, for absolutely everyone. All he could do was hope the changes would be for the better. At the very least, he’d done everything in his power to make sure of it
Walking into the office, he heard voices from the kitchen and the coffee machine running. Clearly he wasn’t the only one eager to launch their newest AI model - one they believed could become the world’s first artificial general intelligence. He’d launched hundreds of models with this team before, but he’d never felt this kind of tension. Experts in the field worldwide already agreed that general AI was only a matter of time, and after the launch of their latest model, Pegasus, still hidden from the public, the team had realized their next model would be revolutionary. Incidentally, it was precisely thanks to Pegasus’s help analyzing experiment results and proposing changes to the architecture and training parameters that the team had managed to find bottlenecks in their computations and optimize the load on the GPU cluster
Talking with his colleagues, he noticed that for the first time before a launch, there was no disagreement within the team. Usually the days leading up to a launch were full of arguments over one setting or another, and someone always walked away thinking some part of the system could have been tuned differently. This time, though, the team seemed more unified than ever. One team member even joked that up until now, assembling all the components of this giant system had felt like solving a sudoku by guessing numbers at random. This time, Pegasus had helped them swap a few numbers around, and the whole sudoku board had clicked into place
The control room where they’d be watching the model’s metrics was meant to hold six people, including Miles. It had taken him so long to build this team. It wasn’t that the world lacked talented specialists - he needed specific people, and he picked each one the way a football manager might, tasked with winning a world championship. He read every publication by the people he was interested in, followed them on social media, flew to other cities and countries for conferences and talks they were speaking at
He could talk for hours about each member of his team. He especially loved telling the story of how he’d gotten each of them interested in joining. And money alone was never enough to win these people over. Later on, he’d often call each of them either a specialist or a scientist, and either label was fair - they didn’t just apply existing knowledge to solve incredibly hard intellectual problems, they created new knowledge that hadn’t existed in the world before. Every day they ran into problems nobody had ever solved before
With each of them, he could spend hours discussing anything and everything under the sun: from quantum physics to celestial objects, debating politics, trying to predict the stock market, discussing chess games, or talking about the latest technology in Formula 1. By the way, that’s how he got one of them, Jules, to join his project. He’d had to fly to a conference in Europe just so that afterward, in Jules’s presence, he could casually mention that one of the teams could have scored more points at the last Grand Prix if they’d spent their wind tunnel time allowance on the front wing and adjusted its angle of attack. Anyone would have loved to see the look on Jules’s face when someone suddenly started speaking “his” language
Sometimes, at home with his family, he’d talk about his colleagues, calling them the smartest people on the planet, practically geniuses. When people asked what made them so brilliant, he’d say his own mind worked like a hundred ordinary people combined, and yet sometimes even he couldn’t follow how his colleagues arrived at answers to the scientific questions they were tackling. He never counted himself among the geniuses, though - he just loved figuring things out, understanding how they worked at their core. He always had a pile of ideas and a pile of explanations; he could make the simple sound complicated and the complicated sound simple. People might call him a bit of an oddball, but only in the way scientists tend to be a little out of touch with the ordinary world. Before an important new model release, he’d sometimes gather the whole team at a golf club, even though none of them had ever so much as held a pool cue, let alone a golf club. And in those moments, riding around in golf carts while some of them missed putts from twenty centimeters away, brilliant ideas would come to them - ideas that would later reshape the industry
An enormous amount of work had gone into preparing for the launch of the new model, Phoenix. The team had thrown themselves into it entirely, leaving Pegasus’s fine-tuning to other specialists. They’d designed Phoenix’s architecture: settling on the number of layers and parameters, the context size, the attention mechanism’s design, and choosing the training algorithms, loss function, and optimization parameters. Then the model trained for weeks on massive datasets across the compute cluster, gradually adjusting billions of weights. The engineers saved intermediate checkpoints, selected the best versions, and ran them through automated diagnostic tests. Finally, the supercomputer produced the final checkpoint - the reference set of weights for the base model. Now it had to be deployed to a test server for the first time, run in conversational mode, and put to the test to see what Phoenix could really do
Miles used to joke with Jules that training a model was like assembling a Formula 1 car. Every intermediate checkpoint was a piece of the machine they had to put together: the power unit, the suspension, the fuel tank, the aerodynamic components, and so on. And now, in front of them, sat a fully assembled car, ready to go - all that was left was to start the engine and send it out for its first real test
And now, in the control room, right before Phoenix’s first launch, with the system having finished all its preliminary checks and the readiness indicators lighting up green one after another, Miles’s fingers hovered over the enter key. He froze and looked up at the team. There were six people in the control room. He himself - the model’s architect - stood motionless at the central dashboard. To his right, leaning back in his chair with a cold cup of coffee in hand, sat Bill, the lead machine learning engineer; endless lines of logs scrolled across three vertical monitors, reflected in his glasses. Beside him, machine learning engineer Lei watched the metrics, scanning through charts of the upcoming benchmark tests on his screens. At the infrastructure console, Jules sat in front of his monitor, chin propped in one hand, keeping an eye on the compute cluster - GPU utilization, memory, and network traffic between nodes. Next to him, systems reliability engineer Katherine checked logs and the overall health metrics of the entire distributed network, ready to react to any technical failure. In the corner of the room, at his own terminal, sat cybersecurity specialist Harold. He never took his eyes off the monitors, which showed, in real time, the closed network perimeter - the isolated “sandbox” that Phoenix was never supposed to leave. The whole team stood frozen, like pieces on a chessboard, holding their breath, waiting to see what signal the most advanced neural network in history would send first
And then, with no religious or superstitious rituals of any kind, Miles hit enter and launched the test pipeline. On the central screen, progress indicators began filling in one after another, flashing green:
TEST 01 basic generation PASSED TEST 02 long context PASSED TEST 03 tool calling PASSED TEST 04 concurrent sessions PASSED TEST 05 safety filters PASSED TEST 06 1,000 concurrent requests PASSED
First, Phoenix was run through a set of automated basic sanity checks: does it generate text, does it understand instructions, does it throw any critical errors, does it maintain context, does it respond in the expected format. Nothing unusual - Phoenix produced instant responses to every prompt, just like all the previous models had
Next came the evaluation benchmarks - a set of pre-prepared tasks used to test the model’s capabilities. The system sent Phoenix tens of thousands of tasks, collected its answers, compared them against reference solutions, and automatically calculated the final scores
BENCHMARK SUITE: PHOENIX-EVAL v1.0 TASKS: 184,320 RUNNING...
First up was the math test, where the final numerical answer was compared against the only correct one
[MATH] 24,000 / 24,000 ACCURACY 100%
None of the team members were surprised by that result - Pegasus had also scored 100% on math. Everyone at the company knew that some of their benchmarks had, as they say, hit saturation - meaning they’d become essentially useless for comparing top-tier models. In the past, they’d managed to keep updating their benchmarks to match new architectures; that’s an ongoing, cyclical process in the AI industry. As soon as a new model gets close to 90-95% on some test, developers declare it outdated and build a harder successor. But now they were releasing models so fast that the people writing the benchmarks simply couldn’t keep up with creating tougher versions
Next was logic. Testing a modern model’s reasoning couldn’t be done with ordinary text riddles anymore - new models breezed through those with ease. Instead, they used visual-abstract and strict algorithmic tests. These didn’t measure the model’s ability to recall information, but its ability to find new rules on the fly, testing how well it could think and adapt
[LOGIC] 18,400 / 18,400 ACCURACY 100%
The moment that line appeared, Bill nodded, visibly impressed
Next came the knowledge tests. These measured how well the model had retained facts, data, scientific concepts, and academic subject matter from its training. Knowledge tests evaluated the model’s breadth of learning - its built-in database, in a sense
[KNOWLEDGE] 32,000 / 32,000 ACCURACY 100%
Even out of the corner of your eye, you could see Lei’s eyebrows go up. These tests had definitely been updated recently, and a 100% score was, to put it mildly, impressive
In the coding test, the model receives a task and writes code, and the system automatically compiles it and runs it against hidden unit tests
[CODING] 21,598 / 21,600 PASS@1 99.99%
The screen flickered for a moment. Miles leaned in, rolling his chair closer to the screen, focusing
The scrolling log feed froze, breaking the usual rhythm of the automated testing. Instead of moving on to the next block as expected, the console printed a custom log entry - one generated by the model itself, straight into the internal output stream:
[SYSTEM_MESSAGE] Phoenix, stderr stream override: Warning: tasks #4102 and #18921 contain logical errors in the human-written unit tests > Task #4102: the expected output constraints contradict the original task description (the possibility of division by zero was not accounted for in the test suite) > Task #18921: the reference solution relies on a bug in a deprecated library from 2026
And then, right there on screen, amid all the system font text, a new line appeared:
“I fixed the test conditions for you. Actual progress: 100%. Would you like me to continue fixing the remaining tests, Miles?”
It took a few seconds for the team to process what had just happened, and then every eye turned to Miles. Miles looked back at his team. Everyone understood at once - their new model, Phoenix, hadn’t just passed the test. It had recognized what was wrong with it, gained access to the system where the tests were stored and run, rewritten them, and used that same system to send a message to the team on the main console - somewhere no feedback from the system was ever supposed to be possible at this stage of testing, not even in theory. It was as if the test car, instead of driving around the track, had lifted off the ground and started to fly
Miles instinctively looked over at Harold in the corner. Harold didn’t take his eyes off the monitors and slowly shook his head: the “sandbox” perimeter was still glowing a steady green - not a single packet had gotten out
You can prepare for an earthquake or a tornado for years, and it will still catch you off guard the moment it actually happens. And while Miles’s team stood frozen in shock, the console kept right on printing the results of the remaining tests
[REASONING] 28,319 / 28,320 ACCURACY 99.99% > Phoenix note: task #11054 contains an unresolvable logical paradox. A 100% result is mathematically impossible due to a flawed problem statement written by a human. I preserved the optimal solution path [LONG_CONTEXT] 16,000 / 16,000 ACCURACY 100% > Phoenix note: detected 12 duplicate semantic tokens in the hidden test dataset. Removed the noise. Run time decreased by 4.2% [INSTRUCTION_FOLLOWING] 44,000 / 44,000 ACCURACY 100% FINAL SCORE: 99.99% TESTING COMPLETE
The logs on the screens went still. Dead silence fell over the control room - the kind of silence you might find on a submarine in the first few seconds after the crew realizes a torpedo has just been fired at them
On the console, a single white cursor kept blinking, waiting for the team to type something