Recently, Amazon Web Services experienced an outage – screens went dark, dashboards stopped blinking, and the internet once again discovered that the “cloud” periodically lives up to its name.
When so many systems stopped working, it was undeniably unpleasant, even painful. People will lose their jobs – including recognized professionals. A great deal of money was lost last week. And the question is not even whether it is possible to avoid this kind of catastrophe. The real question is whether such catastrophes are actually that bad. Considering how much has been saved since the early days of Amazon Web Services (AWS), the balance is still positive.
The sharp response to the company’s problems was immediate. But behind all the noise in the background, you could hear something familiar and somewhat hopeful: the system was performing a huge self‑check. Every great technology sooner or later reaches its limits. And then all that remains is to see how elegantly it learns from such incidents.
The “cloud” has always been tempting and irresistible. It promised endless growth, universal access, and a bill that ought to be more pleasant than a server room full of blinking lights. For most companies, it was a rational choice – cheaper than leased lines, easier to maintain than complex juggling with offices and remote workplaces, and certainly much more appealing than tinkering in a hot summer with a broken data‑center cooling system.
A pinch of irony is an essential ingredient of real progress. Every significant breakthrough creates its own things to become dependent on. The more we optimize to fit the “cloud,” the more responsibility for failures we hand over to others. When AWS goes down, the illusion of universal availability and presence collapses with it. And – note this – it has always been only an illusion.
The old promises of innovation
For a radically new technology to succeed, it must be able to plug into the economy quickly. It is so attractive, so elegant and clearly superior that it cannot be ignored. Engineers will therefore push its performance to the technical maximum, while business will squeeze its return on investment to the last drop. Until we reach a point where nothing more can be extracted from this technology.
But not all companies learn in the same way. Companies have organically split into three tribes.
The conscientious penny‑pinchers. They buy efficiency fully aware of the risks. They know they are saving money and at the same time accept that more or less regular system outages are inevitable. For them, the “cloud” is a conscious risk, not a religion. In the first five to ten years of AWS operations, the company’s customers were exactly this kind of people.
The skeptical engineers. First of all, they have never trusted anyone’s marketing. They know – if something can be built, it can definitely be broken. They also know that systems occasionally fail for no obvious reason or for some strange cause. The CrowdStrike incident was exactly this type of error. They possess a healthy distrust of any system – whether created by humans or by God – and this leads to an even healthier way of thinking about what else could break. One of AWS’s largest customers – Netflix, for example – did not suffer all that much from AWS issues. They created a tool called Chaos Monkey that induces chaos in the infrastructure and application components.
The perpetually busy. They were too busy to also be careful. Their business was growing faster than their contingency planning could adapt to the new situation. This is usually the group that suffers the greatest losses when the system stops.
Each system outage reshuffles the risks between these groups. And more importantly, it also redistributes risk between innovation and the status quo.
When economics meets engineering
Last year’s CrowdStrike fiasco – when a faulty software update paralyzed air traffic and hospitals – was not caused by some cosmic‑scale failure. It began with political choices. The European Union insisted that Microsoft give its competitors access to the core of its operating system. The idea was to promote competition. Initially, Microsoft objected, but then complied with EU requirements. The result was a much more open ecosystem and ultimately a system failure born out of that same openness.
Was that a mistake? Decisions made when lawyers advise engineers and vice versa rarely come without problems.
The same logic drives cloud architecture. Centralization provides speed and room for development, but at the same time makes systems more fragile. When a system goes down in Seattle, it is felt in Singapore. That is why, for at least the last five years, decentralization has been quietly underway.
The changing “cloud”
The once unquestioned dominance of near‑instant scale‑up providers is gradually fading, and two forces are driving this process.
First, when Amazon, Microsoft and Google stopped buying computers from traditional vendors such as Dell and HPE, the hardware market for computer systems began to push back. HPE, for example, was forced to create an entirely new business model. And they succeeded. HPE responded to market changes by offering better pricing and financing options, as well as equipment for managing and maintaining machine clusters.
Second, software vendors also rose to the challenge. Technologies like Kubernetes and new management systems have made running private clusters much easier. The cost gap between living in the “cloud” and running your own system has now shrunk significantly.
Both of these trends mean that local and cloud solutions are once again economically viable and growing in popularity. But the fact that the idea of keeping all your data in the cloud has lost its appeal is no longer much of a secret, even though people try very hard to keep it under wraps.
Even so, there are still many public hints. Elon Musk has publicly stated that by moving X/Twitter to a private data center, they will save 100 million a year. In certain cases, moving computing closer to the owners has once again become cheaper. Of course, one must take into account Elon’s team’s engineering capabilities.
Neo‑clouds are a new type of service provider that in many cases emerged from bitcoin “miners.” This is a high‑performance data center with private infrastructure that also uses public cloud services in its operations. The former Yandex team has been using this strategy very successfully. The pendulum, which once swung toward total centralization, is now seeking its balance.
All markets learn from mistakes
Keeping all your data in a single “cloud” has never been 100% safe. But – and this is a big BUT – it was safer than taking care of data security yourself. Over the last ten years, “multi‑cloud” has evolved from a theory into its own engineering discipline.
If the Amazon outage has a moral, it is this – systems become much more reliable after they have failed. Such a test reveals weak spots, and the market takes these lessons on board. Every such outage drives the next round of updates and improvements.
Be prepared for even greater diversity in infrastructure. Companies like Meta and Google or Amazon will strengthen their influence in AI and on the global network. Companies will develop hybrid solutions and private nodes. Engineering teams, sometimes making painful mistakes, will learn that backups are less about duplication and more about preparing for accidents.
Progress is once again performing its favorite trick – turning yesterday’s scandal into tomorrow’s best practice.
Originally published at https://inc-baltics.com/vai-amazon-problemas-patiesam-nesa-tikai-sliktu/
Like
Love
Happy
Haha
Sad
