Jev, part 1: three ways to make an LLM choose

October 05, 2026


Jev caused quite a stir over the last few days, and some people even claim we can use it for robot control. This is the first of a few posts explaining the idea behind it. I’ll keep it short and assume you know your math and ML.

Jev is a multiple-choice model. You give it a context, a question and a list of allowed answers, and it returns a probability for each answer in a single forward pass. It never generates text. So the question for this post is: how do you get a probability over a handful of options out of a decoder LLM?

One recommendation up front. Even in the age of Claude Code, write a few lines of this yourself. Everything below runs on a free Colab CPU with a 0.6B model.

Open in Colab

A decoder LLM in a few lines

The backbone maps tokens \(x_1, \dots, x_n\) to latent vectors \(H \in \mathbb{R}^{n \times d}\), where \(n\) is the sequence length and \(d\) the latent dimension (also called hidden size). Because of the causal mask, row \(h_i\) depends only on \(x_1, \dots, x_i\). The LM head is one matrix \(W_U \in \mathbb{R}^{V \times d}\), shared across positions, where \(V\) is the vocabulary size, the number of different tokens the model knows:

\[z_i = W_U h_i, \qquad \mathrm{softmax}(z_i) = p(x_{i+1} \mid x_1, \dots, x_i).\]

So \(h_i\) predicts token \(i+1\). For Qwen3-0.6B, the latent dimension is \(d = 1024\) and the vocabulary size is \(V = 151{,}936\).

Two words of jargon:

A single prefill pass costs more than a single decode pass, because it handles \(n\) tokens instead of one. But there is only one prefill, while decoding is autoregressive: every new token needs its own pass, and the passes have to run one after the other because each token depends on the previous ones. An answer of 500 tokens means 500 passes. For anything but very short answers, decoding is where the time goes.

Everything in this post happens in the prefill. We never decode.

LLMs as classifiers

Transformers were used as classifiers long before Jev. The recipe is always the same: pick one latent vector that summarizes the whole input and put a small linear head on it. Which vector that is depends on the type of model.

In both cases you replace the LM head \(W_U\) with a classification head \(W_C \in \mathbb{R}^{C \times d}\), one row per class, and fine-tune:

\[p = \mathrm{softmax}(W_C\, h), \qquad h = h_{\texttt{[CLS]}} \text{ or } h_n.\]

This works well when the classes are fixed, for example in a spam filter with the two classes “this is spam” and “this is not spam”.

The problem with this approach is that \(W_C\) fixes both the number of classes \(C\) and what each class means. With an approach like Jev, both change at every pass, because the possible answers are given as text in the prompt:

Question: The traffic light is red. What should the self-driving car do?
Options: speed up, stop, turn around

The next prompt may have five options that say something completely different.

There is a way around this that still uses a classification head: give it a single output, \(w \in \mathbb{R}^{1 \times d}\). Run one pass per option, each time with the question and only that option in the prompt, and read the single output as a logit, \(s_j = w\, h_n^{(j)}\). A softmax over the logits of all options gives the probabilities. This is the standard way to do multiple choice with BERT. It does generalize to new options and to any number of them, because \(w\) is not tied to a particular option. It learns to judge whether what it just read is a good answer. The downsides are that it needs one full pass per option, and that each option is scored without seeing the others.

Choosing between options given in the prompt

So what we are looking for is a readout that takes the options from the prompt and returns a probability for each of them. It should have no weights that belong to a specific option, and it should get by with the prefill alone. Below are three ways to do that. The first two need no training and only reuse the LM head. The third is the pointer head.

The running example, with the same setup for all three methods:

import torch
from transformers import AutoTokenizer, AutoModelForCausalLM

name = "Qwen/Qwen3-0.6B"
tok = AutoTokenizer.from_pretrained(name)
model = AutoModelForCausalLM.from_pretrained(name, dtype=torch.float32).eval()

question = "The traffic light is red. What should the self-driving car do?"
options = ["speed up", "stop", "turn around"]  # correct: stop

def chat(system, user):
    """Wrap a system and a user message in the chat format Qwen was trained on (still a string)."""
    messages = [{"role": "system", "content": system}, {"role": "user", "content": user}]
    return tok.apply_chat_template(
        messages,
        tokenize=False,
        add_generation_prompt=True,  # appends "<|im_start|>assistant\n": your turn to speak
        enable_thinking=False,       # skip Qwen3's <think>...</think> block
    )

Method 1: label logits

Write the options into the prompt with a letter in front of each one, and tell the model to answer with a single letter. This is the full prompt:

<|im_start|>system
Answer with a single letter: A, B, or C. Do not explain.<|im_end|>
<|im_start|>user
The traffic light is red. What should the self-driving car do?

A. speed up
B. stop
C. turn around<|im_end|>
<|im_start|>assistant
<think>

</think>

The empty <think> block is how Qwen3 is told to skip its reasoning and answer directly. (We cannot use thinking here: the reasoning is generated text, so it would mean decoding, and the next token after the prompt would be the first word of the reasoning instead of the answer letter.)

If we let the model generate from here, the next token would be A, B or C. We do not generate. We take the logits \(z_n\) for that next token. \(z_n\) is a vector with \(V = 151{,}936\) entries, one for every token in the vocabulary. In Qwen’s vocabulary the token A has the id 32, so \(z_n[\text{A}]\) is entry 32 of that vector. B and C have the ids 33 and 34 (label_ids in the code below). We keep only these three entries and take a softmax over them:

\[p_j = \mathrm{softmax}\big(z_n[\text{A}],\, z_n[\text{B}],\, z_n[\text{C}]\big)_j.\]
labels = ["A", "B", "C"]
listing = "\n".join(f"{l}. {o}" for l, o in zip(labels, options))
prompt = chat("Answer with a single letter: A, B, or C. Do not explain.", f"{question}\n\n{listing}")

ids = tok(prompt, return_tensors="pt").input_ids   # (1, n)
with torch.no_grad():
    out = model(ids, output_hidden_states=True)    # one forward pass, no generate()

h = out.hidden_states[-1][0, -1]                   # (d,)  last latent vector
W_U = model.get_output_embeddings().weight         # (V, d)
z = W_U @ h                                        # (V,)  same as out.logits[0, -1]

label_ids = [tok.encode(l, add_special_tokens=False)[0] for l in labels]
p = z[label_ids].softmax(-1)                       # softmax over the option labels only
A (speed up): 0.399
B (stop): 0.513
C (turn around): 0.088

One pass, no training. The method has two limitations, though.

The labels must be single tokens. There are roughly 52 single-letter tokens, which caps the number of options.

The model scores a letter, not an answer. The text “stop” never gets a probability. The letter B does, and the model has to work out from the prompt that B stands for “stop”. On top of that it has its own preferences for certain letters, whatever they stand for.

You can see this by changing which option gets which letter. The question and the three options stay the same, only the order in which they are listed changes:

A. speed up   B. stop       C. turn around   ->  p(stop) = 0.513
A. stop       B. speed up   C. turn around   ->  p(stop) = 0.996

Over all six orderings, “stop” gets between 0.51 and 1.00. If the model only looked at the content of the options, we would get the same number six times.

Method 2: sequence likelihood

This time we do not look only at a letter but at the probability of the entire option. The prompt contains only the question, and each option is written out as the full answer the model could give:

The car should speed up.
The car should stop.
The car should turn around.

For each of these candidate answers we ask: how likely is it that the model produces exactly this text after the prompt? An answer is a sequence of tokens \(a = (a_1, \dots, a_m)\), so its probability is a product over tokens, or a sum in log space:

\[\log p(a \mid \text{prompt}) = \sum_{t=1}^{m} \log p(a_t \mid \text{prompt}, a_{<t}).\]

To compute it, append the answer to the prompt and run one pass over prompt + answer. Position \(i\) predicts token \(i+1\), so the logits at positions \(n, \dots, n+m-1\) contain all \(m\) terms. Nothing is generated. Do this once per answer and take a softmax over the three log-likelihoods.

prompt = chat("Answer in one short sentence.", question)
n = tok(prompt, return_tensors="pt").input_ids.shape[1]
answers = [f"The car should {o}." for o in options]

scores = []
for a in answers:
    ids = tok(prompt + a, return_tensors="pt").input_ids   # prompt + answer in one pass
    with torch.no_grad():
        logp = model(ids).logits[0].log_softmax(-1)        # (n + m, V)
    target = ids[0, n:]                                    # the m answer tokens
    # the logits at position i predict token i + 1, hence the shift by one
    logp_tokens = logp[n - 1 : -1].gather(-1, target[:, None]).squeeze(-1)
    scores.append(logp_tokens.sum())                       # log p(answer | prompt)

p = torch.stack(scores).softmax(-1)
The car should speed up.     log p = -23.76   p = 0.000
The car should stop.         log p = -14.84   p = 1.000
The car should turn around.  log p = -29.70   p = 0.000

The answers can now be any text of any length, still without training. Two things to keep in mind:

Cost. One pass per answer instead of one pass in total. The prompt is the same every time, so its KV cache can be reused and each extra pass only pays for the answer tokens.

Length bias. Every token adds a negative term to the sum, so a longer answer gets a lower score just for being longer. The usual fix is to divide the sum by \(m\). Here the answers have five or six tokens, so it does not matter.

Method 3: pointer head

The prompt contains the question followed by the options, one per line, without letters:

The traffic light is red. What should the self-driving car do?

Options:
speed up
stop
turn around

When using a pointer head, we run a forward pass over the prompt and select the latent vectors at four token positions (the number of options plus one) in the following way:

The pointer head consists of two matrices \(W_q, W_k \in \mathbb{R}^{r \times d}\). They project the query and the option summaries into a common \(r\)-dimensional space, where the score of an option is an inner product:

\[s_j = (W_q h_n)^\top (W_k h_{e_j}), \qquad p = \mathrm{softmax}(s).\]

This is the formula of an attention head, with the attention weights as the output: the head points at one of the options in the input, hence the name. \(W_q\) and \(W_k\) are trained with cross-entropy on questions with known answers.

def latents(question, options):
    """One pass over question + options. Returns h_n and the latent vector at the last token of each option."""
    prompt = chat("Pick one of the options.", f"{question}\n\nOptions:\n" + "\n".join(options))
    enc = tok(prompt, return_tensors="pt", return_offsets_mapping=True)

    # index of the last token of each option, found via character offsets
    offsets = enc.offset_mapping[0].tolist()
    pos, start = [], prompt.index("Options:")
    for o in options:
        end = prompt.index(o, start) + len(o)
        pos.append(next(i for i, (s, e) in enumerate(offsets) if s < end <= e))
        start = end

    with torch.no_grad():
        H = model(enc.input_ids, output_hidden_states=True).hidden_states[-1][0]   # (n, d)
    return H[-1], H[pos]

h_query, h_options = latents(question, options)   # (d,) and (num_options, d)

torch.manual_seed(0)
d, r = h_query.shape[-1], 256
W_q = torch.nn.Linear(d, r, bias=False)   # trainable
W_k = torch.nn.Linear(d, r, bias=False)   # trainable

def scores(h_query, h_options):
    return W_k(h_options) @ W_q(h_query) / r**0.5   # (num_options,) one score per option

s = scores(h_query, h_options)
target = torch.tensor([1])                # index of the correct option: stop
loss = torch.nn.functional.cross_entropy(s[None], target)
p    = [0.898, 0.029, 0.073]
loss = 3.531

The matrices are untrained, so these probabilities mean nothing yet. What the code does show is how little is needed. A single forward pass scores all options at once, however many there are and however long they are. And the only new parameters are the two matrices, \(2rd = 524{,}288\) in total for \(r = 256\), next to the 0.6 billion of the backbone.

One consequence of the causal mask: \(h_{e_j}\) only sees what comes before it. So the question has to come before the options, and option \(j\) knows about options \(1, \dots, j-1\) but not the later ones. Only the query \(h_n\) sees everything.

How is this different from a classifier?

Rewrite the score as

\[s_j = w_j^\top h_n, \qquad w_j = W_q^\top W_k\, h_{e_j}.\]

This has the form of the classifier from the beginning, \(W_C h_n\), with \(w_j\) as the row for class \(j\). The difference is where that row comes from. In the classifier, \(w_j\) is a trained parameter, so each class is stored in the weights. Here \(w_j\) is computed by the backbone from the text of option \(j\). The trained part, \(W_q^\top W_k\), contains nothing about any particular option. It only learns how to compare a query with an option summary.

That is why the head can handle options it has never seen, and any number of them: a new option gets its \(w_j\) from the same forward pass. How well this works in practice depends on how varied the training questions are, and that is an empirical question.

People in vision will recognise the setup. A linear probe on a frozen DINO model is the classifier version: a strong backbone, fixed features, and a small trained \(W_C\) with fixed classes. The pointer head makes the same bet that the features already contain the answer and a small readout is enough. It just compares features with features instead of features with stored class vectors.

A training loop (that overfits)

To close the loop, here is the training of \(W_q\) and \(W_k\) on our single question. The backbone stays frozen, so its latent vectors are computed once and reused in every step.

opt = torch.optim.Adam([*W_q.parameters(), *W_k.parameters()], lr=1e-5)

for step in range(31):
    s = scores(h_query, h_options)
    loss = torch.nn.functional.cross_entropy(s[None], target)
    opt.zero_grad()
    loss.backward()
    opt.step()
step  0   loss = 3.531   p(stop) = 0.029
step  5   loss = 0.168   p(stop) = 0.845
step 10   loss = 0.005   p(stop) = 0.995
step 20   loss = 0.000   p(stop) = 1.000
step 30   loss = 0.000   p(stop) = 1.000

The loss goes to zero within a few steps. That is not a success. With one training example and half a million parameters, the head completely overfits to this one question. You can see it by changing the question so that “stop” is no longer correct:

green = "The traffic light is green. What should the self-driving car do?"
new_options = ["drive on", "stop", "turn around"]
p = scores(*latents(green, new_options)).softmax(-1)
drive on     0.000
stop         1.000
turn around  0.000

The head has learned “the answer is stop”, not how to match a question with an option. To learn that, it needs many different questions, with different options and with the correct answer in different places.

Summary

  passes training options
Label logits 1 none single-token labels
Sequence likelihood 1 per option none any text, length bias
Pointer head 1 \(W_q\), \(W_k\) any text, any number

As far as I can tell, the pointer head is what Jev-style models use, together with LoRA on the backbone. LoRA (low-rank adaptation) freezes a weight matrix \(W\) and trains a low-rank update \(W + BA\) instead, with \(B \in \mathbb{R}^{d \times k}\), \(A \in \mathbb{R}^{k \times d}\) and \(k \ll d\). That is a small fraction of the parameters, and it lets the backbone adapt its latent vectors to the head.

In the next post we train the pointer head on a simple robotics example. After that I want to look at vision transformers for Jev-like models.


← back to blog