Front page
NewsService Operations25m · 12:01 BST · 13 min read
Salesforce Outage Chaos: The Week Uptime Lost Its Impact
During the September 16 Salesforce outage, some customers couldn’t even file a support case about Salesforce being unavailable. At the same time, other providers faced their own outage issues, proving that platform uptime tells CX leaders less than they think about the customer journey.
Before the news, this roundup is a bit personal for us. Back in May, Rhys Fisher sat down with Ty Givens, Rhona Bradshaw, and Mike Wehrs for a discussion that suggested most contact centers are one outage away from disasters, and this fortnight proved the panel right. Givens gave us the test we’ve been applying all week:
“There are two questions. Can the customers get help? And can the agents do their job? Those are the two things that actually make an organization resilient.”
Salesforce, Twilio, GitHub and Akamai answered those questions differently this fortnight, and the global Salesforce outage during Dreamforce definitely brought them into the spotlight.
The September 2026 outage started at 07:50 UTC on September 16 and lasted 7 hours and 36 minutes, spreading across multiple regions, with users reporting errors, delays, and services they simply couldn’t reach. Some customers couldn’t even submit new support cases through Salesforce Help. Salesforce later said requests were “stalling while waiting on a response from an internal login service,” using up available server resources.
Combine that outage with additional issues from Twilio and Google around the same time, and news roundup here is less about Salesforce embarrassment and more about an increasingly important question for CX resilience. If the system customers rely on and the place they go to complain about that system share the same blast radius, customer journey resilience has a very practical problem.
-
Salesforce’s September 16 outage hit access, support-case creation, and downstream workflows, with recovery continuing after logins returned.
-
Twilio’s carrier incidents show why message acceptance, delivery, and delivery confirmation need separate monitoring.
-
Twilio expanded its Salesforce Service Cloud Voice connection to SMS and WhatsApp on September 17, sharpening dependency questions.
-
Akamai and GitHub added two lessons: keep another support route, and test journeys instead of trusting one healthy metric.
Google Drive suffered a separate September 16 outage caused by reduced serving capacity, while a Gmail Android access issue began two days later and was resolved on September 24. At the same time, Salesforce’s disruption spilled into suppliers including Sage and Vonage, giving CX leaders another reason to map dependencies beyond the platform they directly buy.
What Caused the Salesforce Outage on September 16, 2026?
Salesforce traced the failure to requests getting stuck while they waited on an internal login service, eating through available server resources. Its own update shared:
“Requests are stalling while waiting on a response from an internal login service.”
During the incident, Salesforce also said it initially believed an external dependency was affecting a legacy login server. Salesforce then checked with its third-party infrastructure provider, which confirmed there were no issues on its side. Salesforce still hasn’t published the promised full investigation into the technical trigger and underlying cause.
The recovery was trickier than just getting a login screen back online. Salesforce tried rolling restarts, blocked an API endpoint as a temporary measure, and pushed fixes across regions. Some environments still needed manual restarts, while scheduled jobs continued misbehaving after interactive access returned. ThousandEyes independently saw HTTP 503 errors and timeouts from 07:50 UTC, with no evidence of network degradation, showing an application-layer failure.
The nastier CX detail came downstream. Some Salesforce customers couldn’t create support cases. In Japan, PayPay warned that its own customer inquiry forms could become unavailable because of the Salesforce failure. SynergyMarketing handled the same upstream problem differently: public Synergy!LEAD forms kept working, while submitted data waited to sync into Salesforce after recovery. Same Salesforce outage September 2026, very different customer impact.
Info-Tech’s Abbas Jaffery told CIO:
“Cloud does not eliminate architectural dependencies.”
He also asked the recovery question I think buyers should steal:
“What did the business expect to happen during the outage, and can we prove that it actually happened?”
Three days later, Salesforce Data Cloud activations in US-east-1 were disrupted for 13 hours and 41 minutes. This time, Salesforce confirmed the problem came from a third-party cloud storage service and said it was preparing an alternate storage option. For buyers, that’s a fairly pointed reminder to ask where critical dependencies sit and how far their failure can travel.
Why Can Twilio Messages Fail Even When the Core Platform Is Up?
Salesforce wasn’t the only one tackling embarrassing outages recently; Twilio also had a hand in showing how far problems can spread in the CX stack. Sending a message is a chain, and Twilio doesn’t own every link in it. The API can accept a request perfectly well while the carrier route, mobile network, or delivery-status path further downstream has a bad day.
On Twilio’s status page over the last couple of weeks, we’ve seen a number of messages warning about delayed SMS delivery receipts in various regions. Twilio said:
“Message delivery may succeed, but delivery receipts may be delayed.”
That sounds pretty minor overall, until an automated workflow assumes “no receipt” means “no delivery” and sends another OTP, reminder, or alert. Suddenly Twilio SMS delivery issues become a data-state problem too.
The Australia MMS incident was still being investigated on September 27. By September 28, Twilio was also reporting or monitoring carrier-specific SMS delivery problems in Brazil, Belgium, and Spain. Twilio also had its own brief API problem on September 21, touching SMS, MMS, Verify, Conversations, Studio, Sync, and Flex TaskRouter for about a minute. Those incidents needed completely different responses because they broke in completely different places.
The timing makes this more interesting. On September 17, Twilio expanded its Salesforce Service Cloud Voice connection to SMS and WhatsApp, routing conversations through TaskRouter into Salesforce Omni-Channel using the “same presence, skills, and capacity model as voice.”
That’s useful consolidation. It also gives buyers a sharper CPaaS carrier-dependency question: which parts of those channels actually fail independently? A shared agent desktop doesn’t make those failure domains disappear. Twilio itself is planning around that issue. Its new peak-season guidance says heavy traffic can strain carrier networks and cause queues to back up. For buyers, the thing to remember is that the more CPaaS coordinates the journey, the less useful a single “platform operational” badge becomes.
Why Does Google’s Bad Week Make the Failure Pattern More Familiar?
Alongside Salesforce and Twilio, Google’s been having its own problems, and they’re still worth mentioning, even if they don’t impact CX directly. Google Drive went down for 37 minutes on September 16, the same day as the Salesforce disruption.
Google later said a planned capacity test had reduced Drive’s serving capacity; peak traffic then overwhelmed what was left, triggering a protection system that started blocking legitimate requests. Users saw failed access, latency, timeouts, and unexpected CAPTCHA prompts.
Two days later, Google opened a separate Gmail Android incident. Some users on the latest app release were unexpectedly logged out, then found Gmail refusing to recognize the account after they signed back in. Google closed the incident on September 24 after rolling out a fixed Gmail app version. During the disruption, browser access and rolling back the app version were offered as workarounds.
Look at that in the context of the other recent news here, and it gives customer journey resilience another angle. Drive’s problem came from capacity management. Gmail’s came from the client experience. Salesforce had an internal login-service bottleneck. Twilio’s week involved carrier routes and delivery receipts. Different technical causes, same CX headache: customers hit the failure at the point where they were trying to get something done.
It’s also worth saying the Salesforce blast radius spread into other suppliers. Sage reported that its customer-service teams in the UK and Ireland couldn’t review tickets through the portal because of the Salesforce incident, while Vonage warned its own customers that Salesforce’s outage might affect Vonage services. Veeva CRM users also saw login failures, although its iOS offline mode still allowed call creation and media display while sync and online functionality were unavailable.
That’s telling for CX teams. Dependency mapping stops being an architecture exercise when one vendor’s incident starts changing what another company’s support team can see or do.
How Should Contact Centers Operate When CRM or Case History Is Unavailable?
These various incidents should push companies to rethink their contact center outage plan. A good one needs a stripped-back operating mode for the hours when agents can’t pull up the usual customer record. Give them enough verified context to handle urgent work safely, a separate place to capture cases, and clear limits on what they’re allowed to promise or change without current account data.
The Salesforce issue left agents worldwide working without a Salesforce case history while support portals wobbled for several hours. Similar incidents with any major provider can have the same lasting impact.
Identity is probably where things get a bit more complicated. We’ve already reported that around 68% of users abandon identity checks after a slow or failed process, and CRM outages can force customers through authentication all over again. Randy Layman, CTO at AVOXI, said:
“We need to stop thinking about authentication as a black-and-white yes or no. It’s a continuum of probability.”
I think that idea belongs in outage planning. An agent dealing with a recognized customer shouldn’t suddenly be working blind because one lookup fails.
Cross-channel failover has the same trap. Moving someone from chat to voice achieves little if the agent still can’t authenticate them or retrieve the case. Rhona Bradshaw put it well:
“It’s the smallest connector that can create the biggest problem.”
Keep emergency knowledge and temporary case capture outside the affected CRM, preserve approved read-only data where security rules allow it, and test whether backup channels share the failed identity, cloud, or routing layer. That’s graceful degradation in CX you can actually use.
Who Should Own Incident Communication During a CX Outage?
Engineering should own the technical facts and recovery updates. CX should decide how those facts are translated for customers: what’s broken, whether they should retry, which routes still work, and when they’ll hear more.
Akamai gave us a useful example on September 16. A third-party provider problem disrupted its case-ticketing service, leaving some customers unable to create new support requests. Akamai’s status update didn’t stop at “we’re investigating.” It said:
“The customers can use [the alternative support interface] or reach out to country-specific support numbers for urgent cases.”
That’s a good backup to have in place. If one door shuts, point people to another one immediately. Don’t make an already frustrated customer hunt around for it. That’s particularly important when pressure can arrive faster than expected.
Splunk’s 2026 research across 2,000 Global 2000 executives found 90% of technology leaders see customer-support demand rise after an incident. The same study found 81% cite customer loss as a consequence of downtime.
So incident communication best practices for customer experience need to sit inside the outage plan before anything fails. Engineering can explain the fault. CX has to tell a customer what to do with that information.
What Belongs on an Enterprise CX Resilience Checklist?
Follow the customer journey first, then work back through the systems holding it together. GitHub showed why that matters on September 13, when around 28 services degraded and web issue creation failed on roughly 96% of attempts at the worst point. The database cleanup safeguard was tracking replica lag. According to GitHub, it “stayed low the whole time.”
That’s exactly why customer journey resilience needs more than infrastructure dashboards. We’ve previously recommended using synthetic or controlled journey testing to prove whether real interactions still work for this reason. The useful service-management question isn’t only whether an alert fires, but how fast teams can connect it to the customer, queue or channel it’s actually hurting.
The checklist I’d put in front of every buyer after this week is:
-
Map the full journey, including vendor dependencies.
-
Find shared failure domains across channels.
-
Test real tasks, such as login, OTP delivery, case submission, and transfers.
-
Keep escalation outside the primary service stack.
-
Define what agents can safely do in degraded mode.
-
Separate platform failures from carrier failures in monitoring.
-
Test cross-channel failover and failback with identity and context intact.
-
Reconcile queued jobs, cases, and integrations before calling recovery complete.
-
Pre-write customer communications and assign ownership.
Also, keep in mind that Splunk found that 47% of technology leaders say customers are often or very often the first to spot service degradation or outages. If your customer is still your best monitoring tool, there’s more work to do.
In Other News
Two quick items worth reviewing from our In Brief feed alongside this roundup:
-
Cloudflare used its Birthday Week to announce its strategy to become a public certificate authority, targeting post-quantum Merkle Tree certificates for early 2027.
-
The same week brought cf, an agentic CLI covering Cloudflare’s full API, and 3,000+ operations with natural language discovery.
Look Ahead:AWS faced the region it lost at the Dubai Summit on 30 September, the first Middle East flagship since it lost Bahrain’s data. We’ll be watching what gets said there. We’re also keeping an eye out for Salesforce’s root-cause analysis of its recent issues.
Seen something on your own status boards we should know about? Tell us. This beat runs on reader signals, not just our research.
Stop Asking Whether the Platform Is Up
September handed CX buyers a free stress test. Salesforce exposed the reach of a login dependency. Twilio showed how carrier trouble can break communications while the CPaaS stays available. Google had separate Drive and Gmail failures, while Salesforce’s disruption spilled into services used by other suppliers. GitHub showed that even a healthy monitoring signal can sit beside a failing service.
Taken together, that’s a pretty good reason to stop treating uptime as something that belongs to one vendor at a time. Customer journey resilience lives across the handoffs between CRM, messaging, identity, cloud services, and whatever sits underneath them.
The next vendor uptime slide deserves a tougher follow-up: trace one real interaction from login to resolution, then name every dependency you don’t control.
Or, put another way: “Status pages measure components. Customers experience journeys.”
If the usual support route disappears with the service, customers shouldn’t be the ones figuring out where the second door is.
Was Google Responsible for the Salesforce Outage September 2026?
There’s no confirmed evidence that Google caused the Salesforce outage. Salesforce initially said an “external dependency failure” was affecting a legacy login server, but its third-party infrastructure provider confirmed there were no issues on its side. Google Drive did suffer a separate disruption on September 16, followed by a Gmail Android issue later that week.
What Is Graceful Degradation in CX?
Graceful degradation in CX means keeping a smaller, safe version of the customer journey working when part of the stack fails. An agent might lose full CRM history but retain enough verified information to record an urgent case and arrange follow-up. Done well, the customer gets a reduced service rather than hitting a dead end.
What Should a Contact Center Outage Plan Include?
It should make clear what agents can still do without the CRM, where urgent cases go, how customers hear about the disruption, and how temporary records get cleaned up afterward. Teams also need tested cross-channel failover and an escalation route that doesn’t depend on the failed platform. If support goes down with the service, the plan isn’t doing its job.
What Is Dependency Mapping in Customer Experience?
Dependency mapping means identifying every service a customer journey relies on, including systems supplied by vendors your business doesn’t directly control. A CRM failure can affect support tools elsewhere, while storage, email, carriers, and messaging routes can break independently. Mapping those handoffs helps teams see where one failure could interrupt several parts of the journey.
rate this storyhelps rank stories across CX Today
Read nextordered by techtelligence · every pick explained
same beat · Service Operations
HPE Says AI-Powered Networks Can Stop CX Failures Before They Start
3 Sept 2026same beat · Service OperationsWhat Elastic and OpenAI’s New Partnership Means for CX Service Management12 Aug 2026same beat · Service OperationsThe Inaugural Gartner Magic Quadrant for Customer Service Knowledge Management Systems 2026: The Rundown5 Aug 2026
More from the Service Operations desk
News
HPE highlights how AI-native networking, infrastructure performance and resilience could increasingly shape customer experience.
Nicole Willing · 3 Sept 2026News
What Elastic and OpenAI’s New Partnership Means for CX Service Management
Elastic and OpenAI expand their partnership to boost service resilience. This is what it means for CX infrastructure and observability.
Sean Nolan · 12 Aug 2026Guide
The Inaugural Gartner Magic Quadrant for Customer Service Knowledge Management Systems 2026: The Rundown
Gartner’s Magic Quadrant for Customer Service Knowledge Management Systems is here. Meet the Leaders, Challengers, Visionaries & Niche Players
