Skip to content

Service outages and slowdowns disrupt work

Outages and change21 posts from 21 people13d active+1010 posts in the last 7 days, 5 the 7 days before (rising)

Posts per day

Posts per day21 posts, Aug 26 to Sep 25

The posts behind it

21, newest first
PostDate
Please Help - AVD hosts intermittently freezing completely/unresponsive...now, barely any users on when it happens so it's not load issue.. When it hits: Users unable to login/disconnected, RDP dead, Bastion dead, Serial console dead, can't get in any way. Portal still shows it as "Running" the whole time and Azure Monitor telemetry just goes dark....r/sysadminu/VerydxSep 24yesterday
Jira OutageAnyone else experiencing issues with Atlassian/Jira? We are on the East Coast.. Edit: It seems to be back up now. It was down for ~20 minutes..r/sysadminu/Character-Act-7826Sep 24yesterday
[Help] SQL Server Auto-Failover Failing...in the AG and showing as (Synchronized).. The Problem: A critical database (R26UAT22) is failing to auto-failover. When Node-1 goes down, the cluster fails over to Node-2, but the DB doesn't come online.. What I Found :. * The client originally provided an EMC Shared Storage LUN...r/sysadminu/ICYGOD2Sep 24yesterday
HashiCorp: HCP Terraform Returning 404sIncidents (incidents.fru.dev)Sep 232 days ago
Move from Internal IT to MSP advice...projects and issues non stop but then internal IT I seem bored on quite days and struggle to stay productive during downtime even when studying for certifications and so on. The MSP are offering a 5k pay rise but require me onsite for the first 3 months until onboarding is...r/sysadminu/Shoddy-Anteater-1431Sep 232 days ago
Grafana Labs: IRM Access Issues for a Small Group of UsersIncidents (incidents.fru.dev)Sep 223 days ago
Simulating fault tolerance with stage skipping in pipeline-parallel training [R]...hypothesis.. These results point toward training on a broader pool of compute, including unreliable workers and spot instances. This is a simulation of the learning effects of stage failures, rather than a measurement of physical worker replacement or production cost savings.....r/MachineLearningu/covenant_aiSep 223 days ago
Transaction pooler (Supavisor, port 6543) intermittently hangs on trivial, already-indexed queries — project jtctezujqbzwgcsufqsj, eu-west-1# Bug report We've had recurring production outages on this project caused by the transaction pooler (port 6543) intermittently hanging on requests, even for trivial, fast, indexed queries. The direct/session pooler (port 5432) has never once failed under the same conditions....supabase/supabaseHIPIKASep 205 days ago
Downgrade from older versions...during voice mode and when using the read aloud feature in writing it has a lot of trouble with pronunciation. Please fix this, why did these features go downhill?Perplexity App Store reviewsSep 196 days ago
ElevenLabs: Some ElevenAgents calls failed to startIncidents (incidents.fru.dev)Sep 196 days ago
dbt Labs: Schema Hydration PanicIncidents (incidents.fru.dev)Sep 187 days ago
Vercel: Elevated Errors Triggering DeploymentsIncidents (incidents.fru.dev)Sep 187 days ago
LinkedIn node still sends LinkedIn-Version 202604### Bug Description The LinkedIn node hardcodes `LinkedIn-Version: 202604` in `packages/nodes-base/nodes/LinkedIn/GenericFunctions.ts`. LinkedIn Marketing / Community Management APIs use monthly YYYYMM versions. Each version stays active for about one year, then LinkedIn returns...n8n-io/n8ncursor[bot]Sep 187 days ago
OpenAI: Elevated errors affecting ChatGPT Work modeIncidents (incidents.fru.dev)Sep 178 days ago
PSA: Do NOT buy OVH Dedicated Servers if your business actually relies on them (40+ days of delays)...24 (Day 19): Still no server. Opened a ticket. Support replied admitting a "delivery bottleneck," promised delivery would happen on September 15 , and immediately marked the ticket as resolved.. September 15: The promised delivery date is here. Still no server delivered, no...r/sysadminu/Acrobatic_Ad1147Sep 1510 days ago
K8s agent: one transient gRPC UNAVAILABLE during metadata upload leaves a code location in a permanent ERROR state...replaces the pod, but definitions take 60 to 90 seconds to import. 4. The agent logs `Unable to update ... UNAVAILABLE` and `Updated with error`. The location shows `ERROR` in the UI. 5. Wait for the replacement pod to pass the health check. The location stays in `ERROR` with no...dagster-io/dagstercbiniSep 92 weeks ago
...away. You'd have to delegate the whole domain to your DNS provider and then you have no way to manage an outage of your fancy DNS provider. Especially if you go back to the days where NetworkSolutions did a single daily zone update for .com ... if you wanted to switch to a new...Hacker News commentstoast0Sep 33 weeks ago
I did migrate 80% of my tokens. But for some tasks claude models are still the best. Easy workaround is to work outside US peak hours (europe morning). I love this outages, I am hardly affected, and weekly reset usually promptly follows!Hacker News commentsthrow83930489Sep 33 weeks ago
OpenAI having an outage with 5.6 sol?For about 30 minutes now I keep getting the following from codex: Selected model is at capacity. Please try a different model. Anyone else seeing this? Their status page looks like everything is fine at the moment.Ask HNrclevengSep 23 weeks ago
critical_service_loop ignores exponential backoff when jitter_range is set### Bug summary `critical_service_loop` doubles its interval on each backoff, but when `jitter_range` is set the sleep is drawn around the *base* `interval` and `backoff_count` is never applied: The two branches are mutually exclusive, so passing `jitter_range` silently disables...PrefectHQ/prefectethanstonerAug 303 weeks ago
How do you deal with passive aggressive users?...us.. As an example, 3 days back, program x, which is hosted on a client on or VM server, stopped working. For some reason, the host machine was shut down. Not a big deal. Restart the machine, inform user what happened, Bob's your uncle. Yesterday, program x went down again. This...r/sysadminu/PublikEnemyNumber1Aug 264 weeks ago

Companies and products named

Company or productPosts naming it
Perplexity1
Terraform1
Grafana1
ElevenLabs1
dbt1
v01
About this problem

Evidence

21 posts from 21 people in 10 places, about 4 a week over 30 days. Mostly on Incidents (incidents.fru.dev), r/sysadmin, Hacker News comments. Tools named alongside: Perplexity, Terraform, Grafana, ElevenLabs.

Frustration Frustration 0 of 3· Seen on Reddit, fru.dev sites, GitHub issues, Hacker News, App Store reviews

How it was grouped

Posts that state a pain and match the "outage" rule. First post Aug 26, 2026, latest Sep 24, 2026. Corroborated: 3 or more posts from 2 or more people or places. Method

History

  • 2026-09-25 Added: Service outages and slowdowns disrupt work (11 posts)

Rising problems by email

Mondays: the problems in data, tech and AI that grew fastest that week.

Double opt-in. Unsubscribe any time.