Amazon’s shopfront has been wobbling, so now a pile of engineers are being hauled into a Tuesday meeting for a “deep dive” into the mess.
A briefing note for the session flags a “trend of incidents” in recent months, marked by a “high blast radius” and “Gen-AI assisted changes” among the usual gremlins.
Under “contributing factors”, it points to “novel GenAI usage for which best practices and safeguards are not yet fully established”, in other words, people don’t know what the hell they are doing.
Amazon senior vice-president Dave Treadwell wrote, “Folks, as you likely know, the availability of the site and related infrastructure has not been good recently,” and nobody needed a dashboard to confirm it.
The note does not specify which outages will receive post-mortem treatment, but the timing is awkward given that the main site and shopping app dropped for nearly six hours this month.
Amazon blamed an erroneous “software code deployment”, leaving customers unable to complete transactions or even check basic details like account details and product prices.
Treadwell told staff the weekly “This Week in Stores Tech” (TWiST) meeting would become a deep dive into some of the issues that got us here, as well as some short immediate-term initiatives aimed at limiting future outages.”
He wants people to turn up, even though the meeting is normally optional. Junior and mid-level engineers will now need more senior engineers to sign off on any AI-assisted changes.
Amazon insists the availability review was “part of normal business” and it aims for “continual improvement”, as opposed to improvement that does not continue.
The company added, “TWiST is our regular weekly operations meeting with a group of retail technology leaders and teams where we review operational performance across our store,” which is corporate for a weekly round of blame and spreadsheets.
Amazon Web Services has had its own fun, with at least two incidents linked to AI coding assistants the company has been rolling out to staff.
AWS suffered a 13-hour interruption to a cost calculator in mid-December after engineers allowed the group’s Kiro AI coding tool to make certain changes, after which the AI tool decided to “delete and recreate the environment”.
Amazon previously called the December incident an “extremely limited event” affecting only a single service in parts of mainland China, then said the second incident did not hit a “customer-facing AWS service”.
The FT has previously reported multiple Amazon engineers saying their teams were dealing with more “Sev2s”, incidents needing rapid response to avoid product outages, and they blamed job cuts for the churn.
Amazon has pushed through multiple rounds of layoffs in recent years, most recently cutting 16,000 corporate roles in January, while disputing that the recent spike in outages is due to headcount cuts.







