ELF-B-T5Gemma2distilled

ELF-B, a continuous diffusion language model, trained on the embeddings of T5Gemma-2-270M-OWTdistilled, distilled from T5Gemma-2. From the paper Scaling and Distilling Text Embeddings for Better Diffusibility.

Paper · Code · Project page · All models of the release

Model description

ELF (Embedded Language Flow) is a continuous diffusion language model that generates a sequence of text embeddings and decodes them to tokens. This model is the unchanged ELF-B trained on OpenWebText-1024 (sequences of 1024 tokens) for 5 epochs at a batch of 512, on the frozen embeddings of the distilled student T5Gemma-2-270M-OWTdistilled. Sampling uses the SDE sampler with self-conditioning guidance; --sc sets the guidance scale and --nfe the number of steps.

How to use

With the code of the paper; the checkpoint and the tokenizer are downloaded from this repository on first use, no login needed:

git clone https://github.com/la0ka1/diffusing-scaled-text-embeddings
cd diffusing-scaled-text-embeddings
pip install -r requirements.txt

python sample.py --ckpt ELF-B-T5Gemma2distilled --out samples.json --n 1024 --sc 1 --nfe 256
python evaluate.py --samples samples.json

Evaluation

On OpenWebText-1024, sampling at --sc 1 --nfe 256 gives Gen. PPL 31.2 at entropy 5.44 (real text: 15.4 at 5.43; GPT-2-S: 34.1 at 5.45). Gen. PPL is the perplexity of the samples under GPT-2-Large, and entropy is their unigram entropy. See the full sweep over --sc and --nfe in the paper.

Files

ELF-B-T5Gemma2distilled.pt holds the averaged (EMA) weights and a small config (model size, embedding dimension, sequence length, vocabulary, tokenizer, embedding mean and std). The T5Gemma-2 tokenizer is included unchanged.

Limitations

The model generates unconditional web-style English text, unfiltered; it can be false or biased. It is a research artifact for studying text embeddings as latent spaces, not for any downstream use.

License

The model is trained on the outputs of a Model Derivative of T5Gemma-2 and is released under the Gemma Terms of Use. Gemma is provided under and subject to the Gemma Terms of Use found at ai.google.dev/gemma/terms (see NOTICE). Use of this repository is subject to those terms, including the Gemma Prohibited Use Policy. The code is released under the MIT License. This is a research release by the authors of the paper; it is not a Google product and is not endorsed by Google.

Citation

@article{zhang2026scaling,
  title={Scaling and Distilling Text Embeddings for Better Diffusibility},
  author={Zhang, Zekai and Tian, Yunjie and He, Yanjin and Zhang, Xiaoyan and Zhao, Dongdi and Qu, Qing and Fu, Di},
  journal={arXiv preprint arXiv:2610.01016},
  year={2026}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train la0ka1/ELF-B-T5Gemma2distilled

Collection including la0ka1/ELF-B-T5Gemma2distilled

Paper for la0ka1/ELF-B-T5Gemma2distilled