Perplexity pplx-decider-v1.1-27b: Price, Benchmarks and How It Works
Quick answer: Perplexity’s pplx-decider-v1.1-27b is an updated open-weight multimodal decision model based on Qwen3.8-27B. Perplexity’s model card reports a weighted Decision Index score of 61.56, up from 56.4 for v1 and above Jev’s 57.9 on that evaluation. The hosted Decisions API price for v1.1 is reported at $0.02 per million input tokens, with output free because the model returns decision probabilities rather than generated prose.
This is not a chat model. Its purpose is to make structured decisions from text, JSON-like state and images: choose between options, estimate a yes/no probability, or score an input against a defined scale. That makes it potentially useful for routing, moderation, ranking, triage, policy checks and automated workflows where a deterministic-looking probability distribution is more useful than a paragraph.
What is pplx-decider-v1.1-27b?
pplx-decider-v1.1-27b is the second public checkpoint in Perplexity’s 27B decision-model line. The Hugging Face model card identifies the backbone as Qwen3.8-27B and lists the model at roughly 26 billion parameters in BF16.
The model replaces ordinary language generation with a decision head. Instead of predicting arbitrary next tokens to write an answer, it produces probabilities over the valid choices defined for a decision task. That design is why decision models are attractive for software systems that need a score or route rather than a conversational response.
What changed from v1 to v1.1?
Perplexity attributes most of the improvement to two changes: lifting the causal mask in full-attention layers and training on more data, including data from the tasksource collection. The backbone remains the same.
| Metric | pplx-decider v1 | pplx-decider v1.1 |
|---|---|---|
| Base model | Qwen3.8-27B | Qwen3.8-27B |
| Decision Index, model card | 56.4 | 61.56 |
| Context | About 250K tokens | About 250K tokens |
| Weights | Open-weight | Open-weight |
| License | Apache 2.0 | Apache 2.0 |
| Hosted input price | $0.04 / 1M tokens at launch | $0.02 / 1M tokens |
pplx-decider v1.1 benchmark results
The official model card publishes category scores against v1 and Jev. The weighted overall score is the most quoted number, but the category breakdown is more useful because it shows where the gains come from.
| Decision Index category | Jev | v1 | v1.1 |
|---|---|---|---|
| Knowledge | 51.4 | 40.9 | 48.18 |
| Language | 62.0 | 63.5 | 69.45 |
| Retrieval | 55.4 | 54.9 | 61.26 |
| Tools | 75.1 | 79.3 | 78.88 |
| Arts | 37.7 | 39.4 | 44.66 |
| Weighted overall | 57.9 | 56.4 | 61.56 |
The important caveat is that benchmark scores are not a guarantee for a production dataset. If you route support tickets, prioritize leads or classify policy-sensitive content, evaluate the model on your own labeled examples and thresholds before replacing an existing system.
How the decision output differs from a chatbot
A traditional LLM might answer, “This looks like a billing issue because the customer mentions a charge.” A decision model can instead return probabilities such as billing 0.91, technical support 0.06 and sales 0.03. The application can then apply its own threshold and business logic.
This separation is useful because the model is not asked to both reason about the state and generate polished language. The downstream system remains responsible for deciding what to do with the probability.
Supported decision patterns
Perplexity’s decision-model family is designed around structured choices. In practice, the useful patterns include:
- Binary decisions: yes/no probability for a rule or condition.
- Multiple choice: probability distribution over a fixed set of candidate labels.
- Scoring: place an input on an ordered rubric or scale.
- Multimodal state: use images together with text or structured state when the decision depends on visual information.
Price: what $0.02 per million input tokens means
The v1.1 launch cuts the hosted input price in half from v1’s $0.04 per million input tokens to $0.02 per million input tokens. Output is free in the decision API model because the response is a compact decision distribution rather than a long stream of generated text.
For high-volume classification this can change the economics substantially, but token price is only one part of cost. A production system also needs to account for rate limits, retries, logging, evaluation and any human review path for low-confidence decisions.
Context window and self-hosting requirements
The model card describes a context size around 250K tokens. That is large enough for unusually long state inputs, but self-hosting is not lightweight. The checkpoint requires roughly 49 GiB of weights plus working memory on a CUDA GPU.
Perplexity also warns that the evaluated behavior depends on its supplied inference implementation. The checkpoint uses non-causal attention in full-attention layers and a separate decision readout. Loading it as if it were an ordinary Qwen causal language model will not reproduce the reported behavior.
How to run the open-weight checkpoint
The official repository requires Python 3.12 or newer, authenticated Hugging Face access and enough GPU memory. The recommended path is to download the checkpoint, install the pinned requirements and use the included DecisionModel implementation.
That implementation applies the saved calibration temperature and normalizes only over the valid candidates for the current question. If you build a custom serving layer, those details matter. A generic text-generation server that expects a normal vocabulary head is not a drop-in replacement.
Where pplx-decider v1.1 makes sense
- Support ticket routing before escalation to a larger agent.
- Content or transaction scoring where the application needs probabilities.
- Tool routing for an agent that must choose one action from a known set.
- Ranking or eligibility checks governed by an explicit rubric.
- Multimodal classification where images are part of the state.
Where it is a poor fit
Do not choose a decision model when the actual product requirement is to draft an answer, explain reasoning to a user, write code, summarize a document or carry on an open-ended conversation. Those tasks need a generative model. A decision model is best when the action space is known before the model runs.
How v1.1 compares with AVARIXO’s other recent AI coverage
The release sits in a different category from general-purpose models such as Mistral Large 4 or fast conversational models covered in our GPT-6 speed update. If you are building search-heavy workflows instead, see the Cloudflare Web Search API guide.
FAQ
Is pplx-decider-v1.1-27b open source?
The model weights are released under the Apache 2.0 license. “Open-weight” is the more precise term because the training data and full training process are not necessarily packaged as a complete reproducible open-source project.
Can it write normal answers?
No. Its native purpose is structured decision output, not free-form language generation.
Does v1.1 beat Jev?
On the weighted Decision Index table published in Perplexity’s model card, v1.1 scores 61.56 versus Jev’s 57.9. Results on your own workload can differ.
How much GPU memory is needed to self-host?
The official card says to plan for roughly 49 GiB for the weights plus additional working memory.
What is the hosted API price?
Perplexity announced $0.02 per million input tokens for v1.1, half the launch price of v1, with output free.
Sources
Primary source: Perplexity AI — pplx-decider-v1.1-27b model card. Additional context: Perplexity API documentation. Checked October 7, 2026.
