Fastino Labs Released GLiNER2.5-Decide AI Model
The 340 million parameter model enables local, non-generative classification for enterprise decision-making tasks.
Updated on Sept. 25, 2026 in Artificial Intelligence

Live Poll
Do you believe specialized small AI models are more practical for business tasks than large models?
Fastino Labs has released GLiNER2.5-Decide, a non-generative AI model designed for operational tasks like triage and tool selection. The model is currently available as a released software tool for local and air-gapped deployment.
Why it matters
By providing a specialized classification model, Fastino Labs offers an alternative to larger generative systems for workflow routing and guardrails. The architecture focuses on predictable, rule-enforced decision-making rather than open-ended text generation.
The 340 million parameter model achieved 167.3 ms p50 latency on an Intel Xeon CPU and 38.3 ms on an NVIDIA V100 GPU. It utilizes a DeBERTa-v3-large encoder architecture to score permitted answers before a constrained decoder enforces joint rules.
The players
Fastino Labs
A developer of specialized, small-scale machine learning models focused on classification and operational AI workflows.
The details
The model functions as a non-generative classifier, meaning it maps inputs to a predefined set of labels rather than synthesizing new text. It uses joint decoding, a process where the system evaluates multiple related outputs simultaneously to ensure the selected answer adheres to defined logical constraints. Developers can fine-tune the system locally or through the provided GLiNER API, with additional variants released at 1B and 287M parameters.
Timeline
September 24, 2026: Fastino Labs officially released the GLiNER2.5-Decide model.
The Tech Race
Fastino Labs is moving away from the industry trend of massive generative models to focus on highly optimized encoders based on the DeBERTa-v3 research. This trajectory prioritizes operational speed and deterministic classification over the conversational capabilities currently dominating AI development.
Developers can implement this model immediately under an Apache 2.0 license for tasks requiring strict output guardrails. It is designed to run locally on existing CPU or GPU hardware, making it suitable for environments where cloud-based API dependency is not feasible.
The takeaway
The shift toward smaller, intent-specific models demonstrates a growing demand for reliability in automated enterprise decisioning. Watch for further adoption benchmarks or comparative performance data against larger models as Fastino Labs integrates this into standard MLOps pipelines.
Further reading
Explore more developments in local and specialized machine learning models in our Artificial Intelligence section.
Source note: This article includes information reported by MarkTechPost.
Live Poll
Do you believe specialized small AI models are more practical for business tasks than large models?





