Skip to content

Fine-tuning is costly and results are unpredictable

MLOps5 posts from 5 people5d active+44 posts in the last 7 days, 1 the 7 days before (rising)

Posts per day

Posts per day5 posts, Sep 14 to Sep 25

The posts behind it

5, newest first
PostDate
MusicGen: hub configs set decoder dropout=0.1 but the checkpoints were trained with 0 — train() doubles the loss and silently breaks fine-tuning### System Info | item | value | | --- | --- | | transformers | 5.17.0 (latest release); also reproduced on 4.57.6 | | torch | 2.14.0 | | Python | 3.11.16 | | platform | macOS 26 (Darwin 25.6.0), Apple M5 Max, arm64 | | devices tested | cpu and mps - same result on both | |...huggingface/transformersKhalidIbrahemSep 24yesterday
Modernizing the server environment...up-to-date. It was not a smooth journey. Mostly because the budget in academia sucks. Now, I get to start putting plans in place so we don't run EoL items in production again. I'm feeling pretty accomplished. 💪 Also, it pays a tiny bit better than the MSP, but not much. Thank...r/sysadminu/mavric15Sep 232 days ago
...their first release. BTL-2 (or -3 maybe? I forget) looked kind of interesting but it was annoying to run, needed a patch on llama.cpp, looks like they're improving their tooling. I'm certainly interested in how this actually performs.Hacker News commentskadobanSep 223 days ago
Issue doing PEFT on Embedding Gemma w/ new classifier head### System Info Copy-and-paste the text below in your GitHub issue and FILL OUT the two last points. - `transformers` version: 5.16.1 - Platform: Linux-6.6.122+-x86_64-with-glibc2.39 - Python version: 3.13.15 - Huggingface_hub version: 1.29.0 - Safetensors version: 0.8.0 -...huggingface/transformerssr5434Sep 205 days ago
Support 2D bool mask for logits_to_keep, expose last_hidden_state### Feature request Two changes to `Gemma4ForConditionalGeneration.forward`: 1. Accept a bool tensor for `logits_to_keep` to support non-contiguous masking. 2. Add a flag independent of `output_hidden_states` that returns `last_hidden_state`. ### Motivation SFT training with...huggingface/transformersjp1924Sep 1411 days ago

Companies and products named

Company or productPosts naming it
Hugging Face3
GitHub2
llama.cpp1
PyTorch1
NVIDIA1
Tesla1
About this problem

Evidence

5 posts from 5 people in 3 places, about 3 a week over 11 days. Mostly on huggingface/transformers, Hacker News comments, r/sysadmin. Tools named alongside: Hugging Face, GitHub, llama.cpp, PyTorch.

Frustration Frustration 2 of 3· Seen on GitHub issues, Hacker News, Reddit

How it was grouped

Posts that state a pain and match the "fine-tuning" rule. First post Sep 14, 2026, latest Sep 24, 2026. Corroborated: 3 or more posts from 2 or more people or places. Method

History

  • 2026-09-25 Added: Fine-tuning is costly and results are unpredictable (4 posts)

Rising problems by email

Mondays: the problems in data, tech and AI that grew fastest that week.

Double opt-in. Unsubscribe any time.