Skip to content

PyTorch upgrades break working setups

Outages and change6 posts from 6 people5d active+22 posts in the last 7 days, 1 the 7 days before (steady)

Posts per day

Posts per day6 posts, Sep 3 to Sep 25

The posts behind it

6, newest first
PostDate
`SequenceBiasLogitsProcessor`: two edge cases (token id 0 rejected in list format; prefix equal to full context skipped)### System Info - `transformers` version: 5.17.0 - Platform: macOS-26.6.2-arm64-arm-64bit-Mach-O - Python version: 3.14.7 - Huggingface_hub version: 1.33.0 - Safetensors version: 0.8.0 - Accelerate version: not installed - Accelerate config: not found - DeepSpeed version: not...huggingface/transformerslucaluo925Sep 24yesterday
`save_pretrained(state_dict=...)` can duplicate tied weights after state-dict materialization### System Info - transformers: main @ c587bc884db2c2e31fc2b8102314656b17aa07b1 - PyTorch: 2.14.0 - Python: 3.14 - Platform: macOS arm64 ### Who can help? @ArthurZucker @Cyrilvallez ### Information - [ ] The official example scripts - [x] My own modified scripts ### Tasks - [ ]...huggingface/transformersUnitedSnakesSep 214 days ago
[P] Wine synthesis using VAE [P].... Is the loss too large? I know that it never could reach perfect zero by how do I know if the loss is good enough? After reaching the plateau? I use MSELoss.. The repo itself: https://github.com/theaidenmax/tabular-vae-wine-generator . This is my second project in VAE (after...r/MachineLearningu/Dangerous-Pilot-6065Sep 1411 days ago
matrix_exp T8 approximation uses an incorrect x7 coefficient...by `compute_T8`. For a small 2D rotation, this causes a reproducible numerical error when a singleton matrix selects the T8 path. A minimal reproducer is: On macOS ARM64 CPU with PyTorch 2.14.0: The analytic reference is the closed-form exponential of which is the 2D rotation...pytorch/pytorchConcode0Sep 102 weeks ago
[BUG] DTensor RNG tracker is applied to CPU tensors under FSDP2 CPUOffloadPolicy, resetting the CPU RNG state## πŸ› Describe the bug When a module is wrapped with `fully_shard(..., offload_policy=CPUOffloadPolicy())`, its parameters keep a CUDA-typed `DeviceMesh` even though the local shards live on CPU. Random ops on those parameters are still dispatched through the process-wide...pytorch/pytorchfrancesco-bertolottiSep 102 weeks ago
[regression] torch.compile(fullgraph=True) fails on dataclasses.replace for a sourceless instance of a sourced class### πŸ› Describe the bug ### Issue `torch.compile(fullgraph=True)` fails to trace `dataclasses.replace(obj, **changes)` when `obj` is a locally-constructed (sourceless) instance of a module-level (sourced) dataclass, as long as `torch._dynamo.config.enable_trace_load_build_class`...pytorch/pytorchkshitij12345Sep 33 weeks ago

Companies and products named

Company or productPosts naming it
PyTorch6
GitHub3
Hugging Face2
PyPI1
NVIDIA1
About this problem

Evidence

6 posts from 6 people in 3 places, about 2 a week over 22 days. Mostly on pytorch/pytorch, huggingface/transformers, r/MachineLearning. Tools named alongside: GitHub, Hugging Face, PyPI, NVIDIA.

Frustration Frustration 0 of 3Β· Seen on GitHub issues, Reddit

How it was grouped

Posts that state a pain and match the "breaking-changes" rule about PyTorch. First post Sep 3, 2026, latest Sep 24, 2026. Corroborated: 3 or more posts from 2 or more people or places. Method

History

  • 2026-09-25 Added: PyTorch upgrades break working setups (6 posts)

Rising problems by email

Mondays: the problems in data, tech and AI that grew fastest that week.

Double opt-in. Unsubscribe any time.