I’m not very old. Certainly not old enough to sit around reminiscing about the good old days. And yet, that is more or less what I found myself doing last Sunday. I was out walking when I started thinking about how much the way I work with machine learning has changed since I first entered the field. AI is moving so quickly that something from last year can already feel outdated. Go back ten years, and it almost feels like another era. When I started working with machine learning, around 2016, we did things in a way that would seem fairly absurd today. We implemented almost everything ourselves. And I mean everything.
A fairly typical workflow looked something like this:
- Decide on the neural network architecture. Carefully. Changing it later could mean rewriting large parts of your code.
- Choose the loss function.
- Derive the gradients of all the model parameters with respect to that loss, by hand. If you wanted to experiment with certain second-order optimisation methods, there could even be Hessian matrices involved.
- Implement the whole thing in C++.
- Parallelise it with MPI.
- Find some computers to actually run it on, install and configure the cluster software, keep the machines alive, submit the jobs, and eventually collect the results.
Quite comprehensive. Today I can define a neural network in PyTorch, write down a loss function, and let automatic differentiation figure out all the gradients for me. Distributed training frameworks take care of much of the communication between GPUs. And instead of building the computer cluster myself, I submit jobs to some of the largest supercomputers in Europe. That difference is important for more than just convenience. When we were calculating gradients ourselves and running mostly on CPUs, the neural networks we worked with were relatively shallow. There was nothing mathematically preventing us from building deeper networks, of course, but every additional layer increased both the implementation complexity and the computational cost. Today, very deep neural networks containing hundreds of millions or billions of parameters are normal. Automatic differentiation takes care of the enormous computational graph, while GPUs make it possible to perform the corresponding calculations at a completely different scale. In other words, the tools did not just make the old way of working easier. They changed what was practically possible. Back then, the boundary between developing a machine learning model and being a system administrator was considerably less clear. And in our case, “setting up the computer cluster” was sometimes meant very literally.
Meet Smaug
Physics students at the University of Oslo had inherited an old collection of compute nodes dating back to around 2007. The cluster was called Smaug, after the dragon from The Hobbit. Students were responsible for much of the system themselves. That meant software, of course: Linux, networking, Slurm, MPI, libraries and whatever scientific software we needed. But it also meant hardware. Mounting servers in racks. Connecting network cables. Replacing components. Figuring out cooling. Trying to understand why some node had suddenly disappeared from the network. It was HPC in a rather physical sense. Looking back, keeping these old CPU systems alive was probably not particularly sensible from an energy-efficiency perspective. I suspect we were sometimes doing more harm than good. But as an education in how computing actually works, it was difficult to beat. And then we got something considerably bigger.
A snapshot of when we moved Abel from the Sigma2 halls to the physics building. Me and Anders Malthe-Sørenssen in action.
From Abel to Egil
In 2012, the University of Oslo installed Abel, a MEGWARE MiriQuid supercomputer. When it entered the TOP500 list that year, it was among the 100 most powerful supercomputers in the world. Eventually Abel reached the end of its production life. And apparently following some sort of unofficial university tradition, parts of it ended up with the physics students. I was one of the people involved in moving it. And “moving a supercomputer” turns out to be less glamorous than it sounds. We went to the old machine room, disconnected the equipment, removed the compute nodes from the racks, carried everything out, rented a vehicle, drove it across campus and started rebuilding it in the physics building. There were a lot of servers. Once the hardware was in place, the software work began. Operating systems, networking, Slurm, storage, libraries and everything else required to turn a large pile of metal boxes into something that scientists could actually use. We named the new cluster Egil, after the Norwegian physicist and former University of Oslo professor Egil Hylleraas. The name felt appropriate. Hylleraas was one of the pioneers of computational quantum mechanics. In the late 1920s, he performed groundbreaking numerical calculations of the helium atom using the calculating machines available at the time, helping demonstrate that the new quantum mechanics also worked for systems with more than one electron. Almost a century later, we were using thousands of CPU cores to solve descendants of the same kinds of problems. There was just one small problem with Egil. It used a lot of electricity. We couldn’t put the entire system in one place, so initially we split it between two locations in the physics building. Two racks ended up in a small room on the fourth floor. Cooling involved, among other things, a rather generous use of open windows. The third rack went behind a lecture hall on the second floor. This turned out to be a terrible idea. Anyone who has been near a rack full of servers running at high load knows what happens next. The fans spin up. And they are loud. So whenever somebody started using the cluster during a lecture, a wall of server noise appeared behind the lecturer. People complained. Quite reasonably. Eventually the machines had to move again. This time we were offered a garage in the basement. So we dismantled the system, carried everything downstairs, which was an excellent combination of HPC administration and strength training, and rebuilt the cluster yet again. By then we were getting rather good at it. And despite its improvised surroundings, the machine did real science. Several research projects and publications were run on those compute nodes.
This was the sinlge rack that we installed behind the lecturing hall. Left: Before nodes were installed. Right: After all hardware was in place.
Then came the GPUs
At the same time, the type of computing we needed was changing. Machine learning and molecular simulation increasingly benefited from GPUs rather than conventional CPUs. In 2021 I installed our first machine with four NVIDIA A100 GPUs. The following year we added another node with six A100s. Compared with the old CPU clusters, these were extraordinary machines. They became essential for my PhD. All of the molecular dynamics simulations behind the papers in my thesis were ultimately performed on these GPUs. At the time, having ten A100 GPUs available to us felt like an enormous amount of computing power. And it was. But even that now feels quite far removed from how I work today.
Backside of the racks in the basement garage. Some htop and nvtop terminals open for the mood :)
Ten years later
Today, I still work with high-performance computing. Probably more than ever. But my relationship with the hardware is completely different. We now run machine learning models on some of Europe’s largest supercomputers through EuroHPC, including Leonardo, Alps, LUMI and JUPITER. These systems are physically hundreds or thousands of kilometres away from me. I have never installed one of their GPUs. I have never connected one of their network cables. I have never had to work out whether opening a window would provide enough cooling. Most of the time I never see the machines at all. I log in remotely, submit a job, request some GPUs and start training. That feels very distant from carrying compute nodes down the stairs at the physics department. Not only physically, but mentally. Ten years ago, a surprisingly large fraction of the work required to run a machine learning experiment happened below the level of the actual scientific problem. We had to think about gradients, low-level implementation, communication between CPUs, operating systems, schedulers, networking and sometimes the physical machines themselves. Today, several layers of that have disappeared behind abstractions. PyTorch calculates gradients for us. CUDA lets us use GPUs without writing everything at the level of the hardware. Distributed training frameworks handle communication between accelerators. Containers and package managers make environments easier to reproduce. Professional HPC centres maintain the actual machines. That means I spend much more of my time thinking about the model itself. What data should we train on? How should the architecture look? What should the model learn? What loss function should we use? How do we represent uncertainty? How do we evaluate whether the model is actually producing useful forecasts? Of course, this does not mean everything has become easy. I still spend plenty of time debugging distributed jobs, running out of GPU memory, fighting storage systems, dealing with software environments and staring at error messages from jobs running on machines somewhere else in Europe. Some things never change. But the scale has changed enormously. In 2016, building a neural network could mean deriving the gradients, implementing the network in C++, parallelising it with MPI, configuring Slurm and sometimes physically installing the machines it would run on. In 2026, I can train very deep neural networks on some of the largest GPU supercomputers in Europe without ever seeing the hardware. There is something slightly absurd about the fact that this qualifies as nostalgia. It was only ten years ago.
Frontside of the racks in the basement garage.
Feel free to drop a comment or question below if you have thoughts or experiences you’d like to share.