Skip to main content

Pretraining stack

Full pretraining is rare outside foundation-model teams, but Hugging Face documents compatible stacks.

Adapted for this playbook from the 🤗 Transformers documentation by Hugging Face. Official page: https://huggingface.co/docs/transformers/community_integrations/nanotron. Images © Hugging Face (documentation-images) unless noted. This is not a substitute for the upstream docs — verify against the current version.

Covers Hugging Face pages: community_integrations/nanotron, torchtitan, nemo_automodel_pretraining

StackDoc
NanotronNanotron
torchtitantorchtitan
NeMo AutomodelNeMo Automodel pretraining

Most enterprise programmes should fine-tune or RAG rather than pretrain — see Fine-tuning and the playbook RAG guides.

Discussion

Comments​

Share feedback or questions about this page. No account required.

Loading comments…