On 16 September 2026, during the opening day of Dreamforce, Salesforce had one of its largest service disruptions in years. According to Salesforce's own incident record (incident 20004433), users on affected Hyperforce instances saw severe delays, intermittent errors and trouble logging in or creating support cases from 07:50 to 15:26 UTC, a total of 7 hours and 36 minutes.
Salesforce's updates during the incident said a core system component came under increased load while requests stalled waiting on an internal login service. A fix was rolled out region by region, and some instances needed manual restarts. Salesforce has said it will publish a full investigation.
For admins, the most important line in the incident log came late in the day: Salesforce had "received reports of scheduled jobs not running as expected" for customers who could log in again. The service came back, but that does not mean your org's data is back to normal.
This checklist is for the days after an outage like this one. It is written around the September incident, but the same steps apply to any long disruption.
First, work out your own impact window
The headline times are fleet-wide. Your org's window depends on its instance.
- Open the incident on Trust Status and check whether your instance is listed. Salesforce narrowed the list twice during the incident, so an instance that received early notifications may not have been affected.
- Note the start and end times for your instance in UTC. Every query below uses them.
- Salesforce later confirmed that sandboxes were not impacted, so you can focus on production.
Write the window down in one place. When someone asks in three weeks why a customer received two welcome emails, you will want it.
1. Find the scheduled and async jobs that failed or never ran
This is the step most teams miss. An integration that failed usually tells someone. A nightly batch job that silently did not start does not.
Start with Apex jobs that ran, or tried to run, during the window:
SELECT Id, ApexClass.Name, JobType, Status, NumberOfErrors,
ExtendedStatus, CreatedDate, CompletedDate
FROM AsyncApexJob
WHERE CreatedDate >= 2026-09-16T07:00:00Z
AND CreatedDate <= 2026-09-16T18:00:00Z
AND Status IN ('Failed', 'Aborted')
ORDER BY CreatedDate
Then look for scheduled jobs whose last run is older than it should be. A job scheduled to run hourly that last fired before the outage has not recovered on its own:
SELECT CronJobDetail.Name, CronJobDetail.JobType, State,
PreviousFireTime, NextFireTime
FROM CronTrigger
ORDER BY PreviousFireTime
Do not assume a missed run was caught up automatically. For each job, decide whether it is safe to run it again by hand. Jobs that send emails, create records or call external systems are the ones to be careful with.
Scheduled flows are harder to spot, because a run that never started leaves no error behind. The clearest signal is a gap: compare how many times each flow ran on 16 September with a normal weekday.
In orgadmin.ai, Flow Health charts runs against failures per flow, bucketed by day. Set the window to 30 days and the outage shows up as a failure spike or a dip in runs for the affected flows, along with which users hit the errors. It queries live, so there is nothing to set up first.
2. Re-run the integrations that failed, without creating duplicates
Most integrations retry. That is the good news. The bad news is that a retry during a partial outage can leave data behind twice. If the first attempt timed out after Salesforce committed the record but before the response reached the caller, the retry creates a second copy.
For each integration that writes to Salesforce:
- Ask the owner for its error log for the window. Middleware, ETL tools and custom services usually keep one.
- Check whether it inserts or upserts. An upsert on an external ID is safe to retry, because the second attempt updates the record the first attempt created. A plain insert is not.
- Replay the failed rows only, and use an upsert on an external ID wherever the object has one.
If you are replaying rows by hand from a spreadsheet, the orgadmin.ai Data Loader runs upserts on an external ID field you choose, in the browser, and reports errors row by row. You fix the rows that failed and keep the rest.
3. Look for duplicates created by retries
Next, look for records created twice during the window. A quick way to find Contacts that share an email and were created in the window:
SELECT Email, COUNT(Id) copies
FROM Contact
WHERE CreatedDate >= 2026-09-16T07:00:00Z
AND CreatedDate <= 2026-09-16T18:00:00Z
AND Email != null
GROUP BY Email
HAVING COUNT(Id) > 1
Repeat the same idea for the objects your integrations write to: Leads, Cases, Orders and any custom objects. Pick a field that should be unique, such as an external ID, an order number or an email address.
Exact-match queries only find exact copies, though. A retry that wrote slightly different data, like a trimmed phone number or a changed timestamp, will not show up in them. The Duplicate Manager scans any object with fuzzy, phonetic, email and phone matching, puts each pair side by side for a person to review, and keeps a 15-day undo on every merge. That undo matters when you are cleaning up quickly after an incident. We covered the longer-term cost of leaving duplicates in place in what duplicate data really costs.
4. Spot-check record counts and timestamps
Use record counts as a sanity check. Compare how many records each key object received in the window with the same hours on a normal day:
SELECT COUNT()
FROM Case
WHERE CreatedDate >= 2026-09-16T07:50:00Z
AND CreatedDate <= 2026-09-16T15:26:00Z
A count that is far lower than normal points to lost inbound traffic: web-to-lead forms, email-to-case messages or API calls that never arrived. Ask whether the source system queued them. A count that is far higher than normal points to retries (see step 3).
Email-to-Case deserves special attention because customers are directly involved. If a mailbox forwarded messages that Salesforce rejected, the sending server may have bounced them or may still be retrying. Check with whoever owns the mailbox.
If writing SOQL is not something you do every day, the SOQL Assistant turns a plain-English question ("how many Cases were created between 07:50 and 15:26 UTC on 16 September, by origin?") into a query grounded in your org's real fields, and runs it read-only.
5. Confirm the support cases you raised actually exist
Salesforce's own support case creation was affected during the incident. If you or your team tried to raise a case with Salesforce during the window, check that it exists before waiting on a reply.
6. Write it down
Before moving on, record:
- your instance and its real impact window
- which jobs, flows and integrations were affected
- what you re-ran, what you merged and what you decided not to touch, with the reason
This takes ten minutes and pays for itself the next time you have an incident. It is also the evidence you will need if you raise the outage with your account team at renewal.
What to change before the next outage
The September disruption was Salesforce's problem to fix. How much damage an outage does inside your org is still something you can control:
- Make integrations idempotent. Use upserts on external IDs instead of inserts wherever possible. This change alone removes most post-outage clean-up.
- Alert on missed runs, not just on failures. A job that never started does not throw an error. Something should notice when a nightly job has not run by morning.
- Know your critical jobs. Keep a short list of the scheduled jobs, flows and integrations that matter most, so the next post-incident check starts from a list instead of a search.
- Keep failure trends visible. Teams that already watch their flow failure rates can tell right away whether a spike came from the outage or was already there. Teams that do not are guessing.
Outages will happen on any cloud platform. The teams that recover fastest are the ones that already know what "normal" looks like in their org.