ELF-B-T5Gemma2distilled
ELF-B, a continuous diffusion language model, trained on the embeddings of T5Gemma-2-270M-OWTdistilled, distilled from T5Gemma-2. From the paper Scaling and Distilling Text Embeddings for Better Diffusibility.
Paper · Code · Project page · All models of the release
Model description
ELF (Embedded Language Flow) is a continuous diffusion language model that generates a sequence of text embeddings and decodes them to tokens. This model is the unchanged ELF-B trained on OpenWebText-1024 (sequences of 1024 tokens) for 5 epochs at a batch of 512, on the frozen embeddings of the distilled student T5Gemma-2-270M-OWTdistilled. Sampling uses the SDE sampler with self-conditioning guidance; --sc sets the guidance scale and --nfe the number of steps.
How to use
With the code of the paper; the checkpoint and the tokenizer are downloaded from this repository on first use, no login needed:
git clone https://github.com/la0ka1/diffusing-scaled-text-embeddings
cd diffusing-scaled-text-embeddings
pip install -r requirements.txt
python sample.py --ckpt ELF-B-T5Gemma2distilled --out samples.json --n 1024 --sc 1 --nfe 256
python evaluate.py --samples samples.json
Evaluation
On OpenWebText-1024, sampling at --sc 1 --nfe 256 gives Gen. PPL 31.2 at entropy 5.44 (real text: 15.4 at 5.43; GPT-2-S: 34.1 at 5.45). Gen. PPL is the perplexity of the samples under GPT-2-Large, and entropy is their unigram entropy. See the full sweep over --sc and --nfe in the paper.
Files
ELF-B-T5Gemma2distilled.pt holds the averaged (EMA) weights and a small config (model size, embedding dimension, sequence
length, vocabulary, tokenizer, embedding mean and std). The T5Gemma-2 tokenizer is included unchanged.
Limitations
The model generates unconditional web-style English text, unfiltered; it can be false or biased. It is a research artifact for studying text embeddings as latent spaces, not for any downstream use.
License
The model is trained on the outputs of a Model Derivative of T5Gemma-2 and is released under the Gemma Terms of Use.
Gemma is provided under and subject to the Gemma Terms of Use found at
ai.google.dev/gemma/terms (see NOTICE). Use of this repository is
subject to those terms, including the Gemma Prohibited Use Policy.
The code is released under the MIT License. This is a research release by the authors of the paper; it is not
a Google product and is not endorsed by Google.
Citation
@article{zhang2026scaling,
title={Scaling and Distilling Text Embeddings for Better Diffusibility},
author={Zhang, Zekai and Tian, Yunjie and He, Yanjin and Zhang, Xiaoyan and Zhao, Dongdi and Qu, Qing and Fu, Di},
journal={arXiv preprint arXiv:2610.01016},
year={2026}
}