Every major cloud, AI and SaaS provider publishes its incidents on a public status page. We collected every incident that 41 of them declared between July 1 and September 25, 2026, 87 days, and counted what they told their own customers. This is what the third quarter looked like, straight from the providers.
Key findings
- 1,319 incidents in 87 days. The 41 providers declared 1,319 incidents rated minor or worse. Twilio alone accounts for 484 of them, because it posts every carrier-level delivery delay individually. Without Twilio, the other 40 declared 835.
- 280 were outage-level. The provider rated 280 incidents as a major or critical outage, not just degraded performance. 79 were critical.
- The three biggest AI tools declared 245 incidents between them. Anthropic (86), Cursor (84) and OpenAI (75) together declared an incident about every 8.5 hours, 66 of them outage-level. The AI category as a whole is 7 of the 40 providers (excluding Twilio) but 32% of their incidents.
- Anthropic declared the most outage-level incidents of any provider we tracked (31), followed by GitHub (25), Grafana Cloud (23) and Cursor (22).
- GitHub declared the most critical incidents (12), ahead of Snowflake (10) and Sentry (7).
- The typical incident lasted about an hour and a half. The median incident, excluding Twilio’s carrier notices, ran 88 minutes from first post to resolution, and 64% lasted more than an hour.
- Your provider has providers. At least 10 of GitHub’s incidents named a third-party AI model in the title, from “Elevated errors on Fable 5 due to upstream provider” to “Elevated rate of errors for OpenAI models provided by Copilot”.
Every provider, ranked by incidents declared
“Outage-level” means the provider itself rated the incident major or critical. Median duration runs from the first post to the resolution notice. Each name links to the status page the data came from.
| Provider | Category | Incidents | Outage-level | Critical | Median duration |
|---|---|---|---|---|---|
| Twilio | Communication | 484 | 2 | 0 | 5h 49m |
| Anthropic (Claude) | AI | 86 | 31 | 5 | 1h 04m |
| Cursor | AI | 84 | 22 | 5 | 1h 12m |
| OpenAI | AI | 75 | 13 | 3 | 1h 25m |
| GitHub | Developer tools | 64 | 25 | 12 | 1h 22m |
| Grafana Cloud | Developer tools | 62 | 23 | 6 | 2h 10m |
| Supabase | Cloud & hosting | 49 | 18 | 4 | 2h 02m |
| Fly.io | Cloud & hosting | 42 | 19 | 3 | 1h 19m |
| Coinbase | Payments & commerce | 38 | 1 | 0 | 1h 11m |
| Zoom | Communication | 33 | 0 | 0 | 2h 02m |
| Sentry | Developer tools | 25 | 15 | 7 | 3h 00m |
| Vercel | Cloud & hosting | 25 | 10 | 1 | 1h 07m |
| DigitalOcean | Cloud & hosting | 23 | 4 | 2 | 5h 29m |
| Snowflake | Cloud & hosting | 22 | 17 | 10 | 1h 56m |
| CircleCI | Developer tools | 22 | 2 | 1 | 1h 20m |
| Notion | Productivity | 21 | 5 | 0 | 1h 12m |
| Discord | Communication | 17 | 16 | 2 | 1h 22m |
| Render | Cloud & hosting | 16 | 10 | 4 | 58m |
| Jira | Developer tools | 15 | 6 | 3 | 2h 32m |
| MongoDB Atlas | Cloud & hosting | 15 | 6 | 0 | 1h 50m |
| Datadog | Developer tools | 12 | 3 | 1 | 56m |
| Netlify | Cloud & hosting | 10 | 4 | 1 | 26m |
| Mailgun | Communication | 9 | 5 | 3 | 1h 22m |
| Linode | Cloud & hosting | 9 | 0 | 0 | 2h 11m |
| Pinecone | AI | 8 | 6 | 1 | 2h 56m |
| Perplexity | AI | 8 | 5 | 3 | 1h 01m |
| Zapier | Productivity | 8 | 2 | 0 | 3h 30m |
| npm | Developer tools | 7 | 1 | 0 | 1h 39m |
| 1Password | Identity & security | 5 | 3 | 1 | 46m |
| Dropbox | Productivity | 5 | 1 | 0 | 1h 54m |
| Runway | AI | 5 | 1 | 1 | 1h 01m |
| Weights & Biases | AI | 5 | 1 | 0 | 56m |
| Duo Security | Identity & security | 4 | 0 | 0 | 9h 20m |
| New Relic | Developer tools | 2 | 1 | 0 | 1h 26m |
| Bitbucket | Developer tools | 2 | 0 | 0 | 1h 06m |
| Airtable | Productivity | 1 | 1 | 0 | 1h 42m |
| Figma | Productivity | 1 | 1 | 0 | 4h 11m |
| Atlassian Statuspage | Developer tools | 0 | 0 | 0 | – |
| HubSpot | Productivity | 0 | 0 | 0 | – |
| Shopify | Payments & commerce | 0 | 0 | 0 | – |
| Trello | Productivity | 0 | 0 | 0 | – |
Read the counts carefully: they measure disclosure as much as reliability
An incident count is what a provider chose to disclose, and providers disclose very differently. Twilio posts a separate incident whenever SMS delivery to one mobile network in one country slows down: 484 incidents this quarter, only 2 of them outage-level. Shopify and Trello posted no incidents at all in the window, and HubSpot posted only informational notices. A long list on a status page often means a provider is being transparent about small, scoped problems, not that it is less reliable than a quieter peer.
So compare providers within a category, weight the outage-level column more than the total, and treat a zero with some suspicion. The numbers are still useful: they are the providers’ own account of the quarter, and they are the only account you get.
AI providers carried the heaviest load
The AI category split in two. Anthropic, Cursor and OpenAI declared 245 incidents, while the other four AI providers we could include (Pinecone, Perplexity, Runway and Weights & Biases) declared 26 between them. The big three are also the tools developers now wire into their editors, CI pipelines and products, which means their incidents rarely stay theirs.
Many of those incidents were scoped: elevated errors on one model, one surface such as a desktop app or a coding agent, or one region. That is exactly why they are hard to plan around. A partial outage on the one model your feature calls is a full outage for your feature.
Your provider’s provider
GitHub’s status page shows how the dependency chain now runs. At least 10 of its 64 incidents this quarter named a third-party AI model or provider in the title, including degraded availability of GPT models, errors on Gemini, and a Copilot model provider incident with Grok. DigitalOcean posted incidents for hosted Llama, DeepSeek and Kimi models, and Twilio logged Flex AI features timing out on OpenAI. None of those companies’ own infrastructure had to fail for their customers to see errors.
What our own checks saw
Status pages tell you what a provider has declared. To see what a client actually gets, we also send an unauthenticated request every 10 minutes from AWS us-east-1 to 20 production API endpoints, including GitHub, Stripe, npm, PyPI, Cloudflare and eight AI providers. Between August 13 and September 25 that came to about 6,300 checks per endpoint.
- 15 of the 20 endpoints answered every check exactly as expected.
- GitHub’s REST API failed 13 checks, 12 with HTTP 504 and 1 with HTTP 500, spread across eight days. That is tiny next to the 64 incidents on GitHub’s status page, because most of those hit a single feature such as Actions or a Copilot model, not the API every integration calls.
- A failed check is not always an outage. Perplexity’s API answered 193 of our checks with HTTP 429, “too many requests”. The API was up and rate-limiting us. A monitor that treats every unexpected status code as downtime would have logged about 32 hours of outage that never happened.
- The rest were isolated network failures: two timeouts each for the Fastly and Meta Llama APIs, and one failed TCP connection to Spotify’s API.
An answer is not the same as a working service. Our checks are deliberately unauthenticated, so they never see the errors that only happen once you are signed in and calling a model. That is why our numbers run higher than some providers’ own: DeepSeek’s API answered every one of our checks, while its own status page reports 99.66% to 99.89% uptime for its APIs over roughly the same months. Both are true, and the second is the one that matters to an app calling the model.
You can watch these checks live on our public status board, and look up any single service on our Is it down? pages.
What to do with this
- Monitor the endpoint you call, not the vendor’s status page. A status page lags the failure and describes the vendor’s view of it; we covered why in Your Provider’s Status Page Is the Last Place You’ll Learn You’re Down. A check against the exact API and model you depend on tells you what your users are about to see.
- Decide in advance what “healthy” means for each dependency. A 401, a 404 or a 429 can all be a perfectly healthy answer from an API you are not authenticated to. Configure the expected codes, or you will be paged for rate limits.
- Plan for scoped failures. Many incidents in this data hit one model, feature or region. Fallbacks at that level (a second model, a second region, a queued retry) cover far more real incidents than a full multi-cloud failover.
- Find out before your customers do. In New Relic’s 2026 Observability Forecast, 42% of organizations still learn about disruptions through manual checks or customer complaints, and a high-impact outage takes an average of 41 minutes just to detect (New Relic).
Watch the APIs your product depends on
Site Qwality checks any endpoint you rely on, with the status codes you define as healthy, and alerts you on Slack, email, SMS or webhook the moment it stops answering. Free to start, no card required.
Start freeMethodology
- Source. Each provider’s public status page, collected on September 26, 2026. For pages hosted on Atlassian Statuspage we used the page’s incident history feed, which lists every incident of the last three months; for pages hosted on incident.io we used the page’s public incident feed.
- Window. Incidents that started between July 1 and September 25, 2026 (UTC).
- Severity. Statuspage impact levels as the provider set them (minor, major, critical). incident.io levels were mapped as degraded performance to minor, partial outage to major and full outage to critical. “Outage-level” means major or critical. Informational notices and scheduled maintenance are excluded.
- Duration. From the incident’s start to its resolution as shown on the status page, to the minute. Incidents still open on September 26 are counted but left out of the medians.
- Which providers are included. Only providers whose published history provably covers the whole window. Not included: Cloudflare, Fastly, Hugging Face, Together AI, DeepSeek, Mistral AI, xAI, Midjourney, Replicate, Auth0, Adyen and Intercom, whose incident history we could not retrieve in a machine-readable form; and Cohere, Groq, ElevenLabs, Fireworks AI, Stability AI, Miro and Square, whose current status page history starts after July 1, so their counts would be incomplete. SendGrid shares Twilio’s status page and is counted once, as Twilio.
- Our own checks. One unauthenticated HTTP request every 10 minutes from AWS us-east-1 per endpoint, August 13 to September 25, 2026, with each endpoint’s unauthenticated answer (for example 401 or 404) configured as healthy.
Using this data
You are welcome to use these figures in your own writing. Please credit “Site Qwality Q3 2026 Outage Report” with a link to this page. The full dataset is available as a CSV download.