Skip to content

Backfills are slow, costly and risky

Data engineering6 posts from 6 people5d active+11 post in the last 7 days, 3 the 7 days before (fading)

Posts per day

Posts per day6 posts, Sep 9 to Sep 25

The posts behind it

6, newest first
PostDate
[BUG] A budget crossing observed through backfill never fires the exceeded webhook> [!WARNING] > Before submitting a PR, please make sure that: > - A maintainer has triaged this issue and applied the `ready` label > - This issue has no assignee > - No duplicate PR exists > > PRs not meeting these requirements may be automatically closed. ### Issues Policy...mlflow/mlflowrhamzatovSep 214 days ago
Source Recharge: `events` stream's initial start_datetime crashes the entire sync (not clamped to API's 7-day limit)### Connector Name source-recharge ### Connector Version 3.0.11 ### What step the error happened? During the sync ### Relevant information ## What happened? When the connector's global `start_date` config is older than 7 days (e.g. a typical historical backfill start date like...airbytehq/airbyteianadanielSep 1510 days ago
Pending job offer - devil you know vs devil you don't know? (details inside)...a data engineering point of view, our data isn't terribly complex, but our warehouse and lack of design practice over the past 10+ years has overcomplicated a lot of seemingly simple things. We have basically zero documentation, our entire warehouse relies on tribal knowledge....r/dataengineeringu/lostmyway573Sep 1411 days ago
Backfills can complete before a failed run’s retry decision is publishedCodex (GPT-6): > ### What is the issue? > > An asset backfill can be considered complete after a run's `FAILURE` status is stored but before its `dagster/will_retry` decision is published. This is an ordinary publication race; no process crash is required. > > `handle_new_event`...dagster-io/dagstercarl-distillSep 1411 days ago
v3 stream: parallel tool calls collide on the positional fallback key and one is silently dropped### Submission checklist - [x] This is a bug, not a usage question. - [x] I added a clear and descriptive title that summarizes this issue. - [x] I used the GitHub search to find a similar question and didn't find it. - [x] I am sure that this is a bug in LangChain rather than...langchain-ai/langchainlong2bui-andpadSep 112 weeks ago
Spark: vectorized reads probe the position delete index once per row - 43–58% of scan CPU on narrow projections### Summary Spark's vectorized reader asks the position delete index "is row N deleted?" once per row, 5,000 times per batch, although the positions in a batch are a contiguous ascending range. The index can be traversed once for the whole range instead. - Delete-check CPU drops...apache/icebergJeonDaehongSep 92 weeks ago

Companies and products named

Company or productPosts naming it
GitHub2
Snowflake1
Azure Data Factory1
SQL Server1
Anthropic1
OpenAI1
About this problem

Evidence

6 posts from 6 people in 6 places, about 3 a week over 13 days. Mostly on r/dataengineering, langchain-ai/langchain, dagster-io/dagster. Tools named alongside: GitHub, Snowflake, Azure Data Factory, SQL Server.

Frustration Frustration 1 of 3· Seen on GitHub issues, Reddit

How it was grouped

Posts that state a pain and match the "backfills" rule. First post Sep 9, 2026, latest Sep 21, 2026. Corroborated: 3 or more posts from 2 or more people or places. Method

History

  • 2026-09-25 Corroboration: unverified to corroborated
  • 2026-09-25 Added: Backfills are slow, costly and risky (2 posts)

Rising problems by email

Mondays: the problems in data, tech and AI that grew fastest that week.

Double opt-in. Unsubscribe any time.