Bi-directional

community
Activity Feed

AI & ML interests

encoders

Ihorย 
posted an update 8 months ago
view post
Post
575

๐Ÿง  One Model to Classify, Verify, and Guard โ€” Meet GLiClass-Instruct

When we first released GLiClass, it was a fast, zero-shot text classifier that could rival cross-encoders at a fraction of the cost. But classification alone wasn't enough. Our real ambition was a single, lightweight model that could handle the diverse range of text-understanding tasks via classification.
We are excited to announce GLiClass-Instruct โ€“ a leap forward that transforms GLiClass from a classifier into an instruction-following, multi-task engine.
What's new:
โ–ช๏ธ Hierarchical labeling: organize labels into structured groups for complex taxonomies
โ–ช๏ธ In-context learning via examples: provide few-shot examples to adapt on the fly, no fine-tuning needed
โ–ช๏ธ Prompting support: guide classification behavior with natural-language task descriptions
โ–ช๏ธ EWC for preventing catastrophic forgetting: add new capabilities without losing old ones
โ–ช๏ธ 3x faster inference thanks to FlashDeBERTa
New multi-task capabilities:
Beyond topic classification and sentiment analysis, GLiClass now supports:
โ–ช๏ธ Hallucination detection: verify whether LLM outputs are grounded in context
โ–ช๏ธ Rule-following verification: check if content complies with custom guidelines
โ–ช๏ธ Safety classification: detect prompt injections, jailbreaks, and harmful requests
These tasks are crucial for building reliable and efficient agentic systems, where every LLM output needs to be verified, every user input needs to be screened, and every response needs to follow the rules, all at minimal latency.
We release 3 instruction-following models (edge, base, large), with the large model matching SoTA classification models while unlocking entirely new task categories.
๐Ÿ”— Explore more:
GitHub repo: https://github.com/Knowledgator/GLiClass
Models: https://huggingface.co/knowledgator/gliclass-multitask-large-v1.0
Our other solutions: https://www.knowledgator.com/
Ihorย 
posted an update 8 months ago
view post
Post
339
Meet **GLinker** โ€” an ultra-fast, modular, **zero-shot entity linking** framework ๐Ÿš€

When we introduced the **GLiNER bi-encoder** in 2024, it enabled efficient zero-shot NER across hundreds of entity types. But that was just the beginning. Our bigger goal was always clear: **linking text to millions of entities dynamically, without retraining**.

In other words: **true entity linking at scale** โšก

This unlocks powerful applications:
โ–ช๏ธ More precise search with real-world entity disambiguation
โ–ช๏ธ Knowledge graph construction across diverse document collections
โ–ช๏ธ Wikification โ€” turning raw text into richly linked, navigable knowledge

After nearly two years of research + engineering, this vision is now real.

Weโ€™re excited to release **GLinker** โ€” a **production-ready**, zero-shot entity linking system powered by our novel **GLiNER bi-encoder**. It efficiently detects entity spans of any length and matches them directly to entity descriptions โ€” **no retraining required**.

**Why GLinker?**
โ–ช๏ธ Production-ready: multi-layer caching (Redis โ†’ Elasticsearch โ†’ PostgreSQL)
โ–ช๏ธ Research-friendly: fully configurable YAML pipelines
โ–ช๏ธ High performance: precomputed embeddings for bi-encoder models
โ–ช๏ธ Scalable by design: DAG-based execution + efficient batch processing

GLinker transforms raw text into **structured, disambiguated entity mentions**, bridging unstructured language with large, evolving knowledge bases.

๐Ÿ”— Explore more:
GitHub: https://github.com/Knowledgator/GLinker
Report: https://github.com/Knowledgator/GLinker/blob/main/papers/GLiNER_bi_Encoder_paper.pdf
Linking models: https://huggingface.co/collections/knowledgator/gliner-linker
Bi-encoder models: https://huggingface.co/collections/knowledgator/gliner-bi-encoder
Ihorย 
posted an update 11 months ago
view post
Post
1348
Hey builders ๐Ÿ‘ทโ€โ™€๏ธ

Weโ€™re Knowledgator, the team behind open-source NLP models like GLiNER, GLiClass, and many other used for zero-shot text classification and information extraction.

If youโ€™ve explored them on Hugging Face or used our frameworks from GitHub, weโ€™d love your input:
๐Ÿงฉ Which of our models, like GLiNER or zero-shot classifiers, do you find helpful in your practical workflows?
๐Ÿงฉ Howโ€™s the setup, performance, and accuracy been for you?
๐Ÿงฉ Anything confusing, buggy, or missing that would make your workflow smoother?

Your feedback helps us improve speed, clarity, and stability for everyone in the open-source community.

๐Ÿ’ฌ Comment directly here or join the discussion. We read every one ๐Ÿ˜‰:
GitHub: https://github.com/Knowledgator
Discord: https://discord.gg/GXRcAVJQ
HuggingFace:
knowledgator


๐Ÿ“ Want to shape our next release?
Click here to complete this 2-min survey: https://docs.google.com/forms/d/e/1FAIpQLSdyz2UMHrMDX8S9stpBk0wyfngtKSYzwk-02mN1VNYDdTw8OQ/viewform
stefan-itย 
posted an update over 1 year ago
view post
Post
5526
Wohoo ๐Ÿฅณ I have finished my 2025 GPU workstation build and I am very excited to train new awesome open source models on it.

I built my last GPU workstation 5 years ago featuring an AMD Ryzen 5900X, 64GB of G.SKILL Trident Z RGB on an ASRock X570 Taichi cooled by an Alphacool Eisbรคr 420. GPU was a Zotac RTX 3090 AMP Extreme. Unfortunately, I was never satisfied with the case - some Fractal Define 7, as it is definitely too small, airflow is not optimal as I had to open the front door all the time and it also arrived with a partly damaged side panel.

For my new build, I've used the following components: an outstanding new AMD Ryzen 9950X3D with 64GB of Corsair Dominator Titanium (what a name). As a huge Noctua fan - warm greetings to my Austrian neighbors - I am using the brand new Noctua NH-D15 G2 on an ASRock X870E Taichi in an amazing Lian Li LANCOOL III chassis. One joke that only NVIDIA Blackwell users will understand: you definitely need a tempered glass panel to check if your GPU cables/connectors start melting ๐Ÿ˜‚ And the best is yet to come: I returned my previously bought Zotac RTX 5090 Solid to the eBay seller (because of... missing ROPs, only NVIDIA Blackwell users will again understand) and bought a Zotac 5090 AMP Extreme INFINITY (yes, the long name indicates that this is the flagship model from Zotac) from a more trustworthy source (NBB in Germany).

I am so happy to start training and fine-tuning new open source models - stay tuned!!!
  • 3 replies
ยท
stefan-itย 
posted an update over 1 year ago
view post
Post
1054
๐Ÿ‡น๐Ÿ‡ท ๐Ÿ˜ I'm very happy to finally announce my new Turkish LM called "BERT5urk":

stefan-it/bert5urk

It is a 1.42B T5-based model, trained with UL2 pretraining objective on the Turkish part of the awesome HuggingFaceFW/fineweb-2 dataset.

Feel free to check it out!
  • 1 reply
ยท
stefan-itย 
posted an update over 1 year ago
view post
Post
3350
After running some 3DMark and FurMark benchmarks on Windows to make sure that my new 5090 is not causing melting cables [1] and some nice shots with a thermal camera (I don't think that's too much), running some fine-tuning experiments with my favorite Flair & Transformers libraries are very easy to perform.

Important steps:

Good idea is to start with a fresh Ubuntu 24.04 installation with latest CUDA 12.8 and the open NVIDIA driver - follow more advices from [2]:

sudo apt -y install cuda-toolkit-12-8 nvidia-open

I tried update from an existing Ubuntu installation with an older CUDA and driver version and it resulted in a non-startable system.

If you are using PyTorch 2.6 with built CUDA 12.6 it will result in:

NVIDIA Graphics Device with CUDA capability sm_120 is not compatible with the current PyTorch installation.
The current PyTorch install supports CUDA capabilities sm_50 sm_60 sm_70 sm_75 sm_80 sm_86 sm_90.

But no worries! For PyTorch you need just to use a nightly 2.7 version that was built with CUDA 12.8. This can easily done via:

pip install --pre torch --index-url https://download.pytorch.org/whl/nightly/cu128

After that the latest Flair version can be installed and fine-tuning will work!

References:

[1]: https://www.reddit.com/r/nvidia/comments/1inpox7/rtx_50_series_12vhpwr_megathread/
[2]: https://developer.nvidia.com/cuda-downloads?target_os=Linux&target_arch=x86_64&Distribution=Ubuntu&target_version=24.04&target_type=deb_network
  • 1 reply
ยท
stefan-itย 
posted an update over 1 year ago
view post
Post
5205
She arrived ๐Ÿ˜

[Expect more models soon...]
  • 2 replies
ยท
Ihorย 
posted an update over 1 year ago
view post
Post
1945
๐Ÿš€ Reproducing DeepSeek R1 for Text-to-Graph Extraction

Iโ€™ve been working on replicating DeepSeek R1, focusing on zero-shot text-to-graph extractionโ€”a challenging task where LMs extract entities and relations from text based on predefined types.

๐Ÿง  Key Insight:
Language models struggle when constrained by entity/relation types. Supervised training alone isnโ€™t enough, but reinforcement learning (RL), specifically Guided Reward Policy Optimization (GRPO), shows promise.

๐Ÿ’ก Why GRPO?
It trains the model to generate structured graphs, optimizing multiple reward functions (format, JSON validity, and extraction accuracy).
It allows the model to learn from both positive and hard negative examples dynamically.
RL can be fine-tuned to emphasize relation extraction improvements.

๐Ÿ“Š Early Results:
Even with limited training, F1 scores consistently improved, and we saw clear benefits from RL-based optimization. More training = better performance!

๐Ÿ”ฌ Next Steps:
Weโ€™re scaling up experiments with larger models and high-quality data. Stay tuned for updates! Meanwhile, check out one of our experimental models here:
Ihor/Text2Graph-R1-Qwen2.5-0.5b

๐Ÿ“” Learn more details from the blog post: https://medium.com/p/d8b648d9f419

Feel free to share your thoughts and ask questions!
  • 2 replies
ยท
Ihorย 
posted an update almost 2 years ago
view post
Post
1256
๐Ÿš€ Welcome the New and Improved GLiNER-Multitask! ๐Ÿš€

Since the release of our beta version, GLiNER-Multitask has received many positive responses. It's been embraced in many consulting, research, and production environments. Thank you everyone for your feedback, it helped us rethink the strengths and weaknesses of the first model and we are excited to present the next iteration of this multi-task information extraction model.

๐Ÿ’ก Whatโ€™s New?
Here are the key improvements in this latest version:
๐Ÿ”น Expanded Task Support: Now includes text classification and other new capabilities.
๐Ÿ”น Enhanced Relation Extraction: Significantly improved accuracy and robustness.
๐Ÿ”น Improved Prompt Understanding: Optimized for open-information extraction tasks.
๐Ÿ”น Better Named Entity Recognition (NER): More accurate and reliable results.

๐Ÿ”ง How We Made It Better:
These advancements were made possible by:
๐Ÿ”น Leveraging a better and more diverse dataset.
๐Ÿ”น Using a larger backbone model for increased capacity.
๐Ÿ”น Implementing advanced model merging techniques.
๐Ÿ”น Employing self-learning strategies for continuous improvement.
๐Ÿ”น Better training strategies and hyperparameters tuning.

๐Ÿ“„ Read the Paper: https://arxiv.org/abs/2406.12925
โš™๏ธ Try the Model: knowledgator/gliner-multitask-v1.0
๐Ÿ’ป Test the Demo: knowledgator/GLiNER_HandyLab
๐Ÿ“Œ Explore the Repo: https://github.com/urchade/GLiNER