<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://phuembeli.com/feed.xml" rel="self" type="application/atom+xml" /><link href="https://phuembeli.com/" rel="alternate" type="text/html" /><updated>2026-10-05T08:26:34+00:00</updated><id>https://phuembeli.com/feed.xml</id><title type="html">Patrick Huembeli</title><subtitle>Patrick Huembeli — Founding Research Engineer at Noumenal Labs. Physical AI, probabilistic modelling, and physics.</subtitle><author><name>Patrick Huembeli</name></author><entry><title type="html">Jev, part 1: three ways to make an LLM choose</title><link href="https://phuembeli.com/blog/2026/10/05/jev-part-1/" rel="alternate" type="text/html" title="Jev, part 1: three ways to make an LLM choose" /><published>2026-10-05T00:00:00+00:00</published><updated>2026-10-05T00:00:00+00:00</updated><id>https://phuembeli.com/blog/2026/10/05/jev-part-1</id><content type="html" xml:base="https://phuembeli.com/blog/2026/10/05/jev-part-1/"><![CDATA[<p>Jev caused quite a stir over the last few days, and some people even claim we can use it for robot control. This is the first of a few posts explaining the idea behind it. I’ll keep it short and assume you know your math and ML.</p>

<p>Jev is a multiple-choice model. You give it a context, a question and a list of allowed answers, and it returns a probability for each answer in a single forward pass. It never generates text. So the question for this post is: how do you get a probability over a handful of options out of a decoder LLM?</p>

<p>One recommendation up front. Even in the age of Claude Code, write a few lines of this yourself. Everything below runs on a free Colab CPU with a 0.6B model.</p>

<p><a href="https://colab.research.google.com/drive/1CNEaJMmT9z5hgPG3fW5ftysNLui7PeDB" target="_blank"><img src="https://colab.research.google.com/assets/colab-badge.svg" alt="Open in Colab" style="margin: 0; border-radius: 0;" /></a></p>

<h1 id="a-decoder-llm-in-a-few-lines">A decoder LLM in a few lines</h1>

<p>The backbone maps tokens \(x_1, \dots, x_n\) to latent vectors \(H \in \mathbb{R}^{n \times d}\), where \(n\) is the sequence length and \(d\) the latent dimension (also called hidden size). Because of the causal mask, row \(h_i\) depends only on \(x_1, \dots, x_i\). The LM head is one matrix \(W_U \in \mathbb{R}^{V \times d}\), shared across positions, where \(V\) is the vocabulary size, the number of different tokens the model knows:</p>

\[z_i = W_U h_i, \qquad \mathrm{softmax}(z_i) = p(x_{i+1} \mid x_1, \dots, x_i).\]

<p>So \(h_i\) predicts token \(i+1\). For Qwen3-0.6B, the latent dimension is \(d = 1024\) and the vocabulary size is \(V = 151{,}936\).</p>

<p>Two words of jargon:</p>

<ul>
  <li><strong>Prefill</strong> is the first pass, over the whole prompt. All \(n\) prompt tokens are processed in parallel, and the pass gives the distribution of the first new token. It also leaves behind the keys and values of every layer, the KV cache.</li>
  <li><strong>Decode</strong> is every pass after that. Each one produces one more token and reuses the cache.</li>
</ul>

<p>A single prefill pass costs more than a single decode pass, because it handles \(n\) tokens instead of one. But there is only one prefill, while decoding is autoregressive: every new token needs its own pass, and the passes have to run one after the other because each token depends on the previous ones. An answer of 500 tokens means 500 passes. For anything but very short answers, decoding is where the time goes.</p>

<p>Everything in this post happens in the prefill. We never decode.</p>

<h1 id="llms-as-classifiers">LLMs as classifiers</h1>

<p>Transformers were used as classifiers long before Jev. The recipe is always the same: pick one latent vector that summarizes the whole input and put a small linear head on it. Which vector that is depends on the type of model.</p>

<ul>
  <li><strong>Encoder models</strong> like BERT have no causal mask, so every token attends to every other token. A special <code class="language-plaintext highlighter-rouge">[CLS]</code> token is added at the start of the input. It stands for no word, and its latent vector \(h_{\texttt{[CLS]}}\) is used as the summary of the sequence.</li>
  <li><strong>Decoder models</strong> like GPT or Qwen are trained to predict the next token, so they must not look into the future of a sentence. The causal mask enforces this: every token only attends to the tokens before it, so \(h_i\) only knows tokens \(1, \dots, i\). The only vector that has seen the whole input is the last one, \(h_n\), where \(n\) is the length of the input sequence.</li>
</ul>

<p>In both cases you replace the LM head \(W_U\) with a classification head \(W_C \in \mathbb{R}^{C \times d}\), one row per class, and fine-tune:</p>

\[p = \mathrm{softmax}(W_C\, h), \qquad h = h_{\texttt{[CLS]}} \text{ or } h_n.\]

<p>This works well when the classes are fixed, for example in a spam filter with the two classes “this is spam” and “this is not spam”.</p>

<p>The problem with this approach is that \(W_C\) fixes both the number of classes \(C\) and what each class means. With an approach like Jev, both change at every pass, because the possible answers are given as text in the prompt:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>Question: The traffic light is red. What should the self-driving car do?
Options: speed up, stop, turn around
</code></pre></div></div>

<p>The next prompt may have five options that say something completely different.</p>

<p>There is a way around this that still uses a classification head: give it a single output, \(w \in \mathbb{R}^{1 \times d}\). Run one pass per option, each time with the question and only that option in the prompt, and read the single output as a logit, \(s_j = w\, h_n^{(j)}\). A softmax over the logits of all options gives the probabilities. This is the standard way to do multiple choice with BERT. It does generalize to new options and to any number of them, because \(w\) is not tied to a particular option. It learns to judge whether what it just read is a good answer. The downsides are that it needs one full pass per option, and that each option is scored without seeing the others.</p>

<h1 id="choosing-between-options-given-in-the-prompt">Choosing between options given in the prompt</h1>

<p>So what we are looking for is a readout that takes the options from the prompt and returns a probability for each of them. It should have no weights that belong to a specific option, and it should get by with the prefill alone. Below are three ways to do that. The first two need no training and only reuse the LM head. The third is the pointer head.</p>

<p>The running example, with the same setup for all three methods:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kn">import</span> <span class="nn">torch</span>
<span class="kn">from</span> <span class="nn">transformers</span> <span class="kn">import</span> <span class="n">AutoTokenizer</span><span class="p">,</span> <span class="n">AutoModelForCausalLM</span>

<span class="n">name</span> <span class="o">=</span> <span class="s">"Qwen/Qwen3-0.6B"</span>
<span class="n">tok</span> <span class="o">=</span> <span class="n">AutoTokenizer</span><span class="p">.</span><span class="n">from_pretrained</span><span class="p">(</span><span class="n">name</span><span class="p">)</span>
<span class="n">model</span> <span class="o">=</span> <span class="n">AutoModelForCausalLM</span><span class="p">.</span><span class="n">from_pretrained</span><span class="p">(</span><span class="n">name</span><span class="p">,</span> <span class="n">dtype</span><span class="o">=</span><span class="n">torch</span><span class="p">.</span><span class="n">float32</span><span class="p">).</span><span class="nb">eval</span><span class="p">()</span>

<span class="n">question</span> <span class="o">=</span> <span class="s">"The traffic light is red. What should the self-driving car do?"</span>
<span class="n">options</span> <span class="o">=</span> <span class="p">[</span><span class="s">"speed up"</span><span class="p">,</span> <span class="s">"stop"</span><span class="p">,</span> <span class="s">"turn around"</span><span class="p">]</span>  <span class="c1"># correct: stop
</span>
<span class="k">def</span> <span class="nf">chat</span><span class="p">(</span><span class="n">system</span><span class="p">,</span> <span class="n">user</span><span class="p">):</span>
    <span class="s">"""Wrap a system and a user message in the chat format Qwen was trained on (still a string)."""</span>
    <span class="n">messages</span> <span class="o">=</span> <span class="p">[{</span><span class="s">"role"</span><span class="p">:</span> <span class="s">"system"</span><span class="p">,</span> <span class="s">"content"</span><span class="p">:</span> <span class="n">system</span><span class="p">},</span> <span class="p">{</span><span class="s">"role"</span><span class="p">:</span> <span class="s">"user"</span><span class="p">,</span> <span class="s">"content"</span><span class="p">:</span> <span class="n">user</span><span class="p">}]</span>
    <span class="k">return</span> <span class="n">tok</span><span class="p">.</span><span class="n">apply_chat_template</span><span class="p">(</span>
        <span class="n">messages</span><span class="p">,</span>
        <span class="n">tokenize</span><span class="o">=</span><span class="bp">False</span><span class="p">,</span>
        <span class="n">add_generation_prompt</span><span class="o">=</span><span class="bp">True</span><span class="p">,</span>  <span class="c1"># appends "&lt;|im_start|&gt;assistant\n": your turn to speak
</span>        <span class="n">enable_thinking</span><span class="o">=</span><span class="bp">False</span><span class="p">,</span>       <span class="c1"># skip Qwen3's &lt;think&gt;...&lt;/think&gt; block
</span>    <span class="p">)</span>
</code></pre></div></div>

<h1 id="method-1-label-logits">Method 1: label logits</h1>

<p>Write the options into the prompt with a letter in front of each one, and tell the model to answer with a single letter. This is the full prompt:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>&lt;|im_start|&gt;system
Answer with a single letter: A, B, or C. Do not explain.&lt;|im_end|&gt;
&lt;|im_start|&gt;user
The traffic light is red. What should the self-driving car do?

A. speed up
B. stop
C. turn around&lt;|im_end|&gt;
&lt;|im_start|&gt;assistant
&lt;think&gt;

&lt;/think&gt;
</code></pre></div></div>

<p>The empty <code class="language-plaintext highlighter-rouge">&lt;think&gt;</code> block is how Qwen3 is told to skip its reasoning and answer directly. (We cannot use thinking here: the reasoning is generated text, so it would mean decoding, and the next token after the prompt would be the first word of the reasoning instead of the answer letter.)</p>

<p>If we let the model generate from here, the next token would be <code class="language-plaintext highlighter-rouge">A</code>, <code class="language-plaintext highlighter-rouge">B</code> or <code class="language-plaintext highlighter-rouge">C</code>. We do not generate. We take the logits \(z_n\) for that next token. \(z_n\) is a vector with \(V = 151{,}936\) entries, one for every token in the vocabulary. In Qwen’s vocabulary the token <code class="language-plaintext highlighter-rouge">A</code> has the id 32, so \(z_n[\text{A}]\) is entry 32 of that vector. <code class="language-plaintext highlighter-rouge">B</code> and <code class="language-plaintext highlighter-rouge">C</code> have the ids 33 and 34 (<code class="language-plaintext highlighter-rouge">label_ids</code> in the code below). We keep only these three entries and take a softmax over them:</p>

\[p_j = \mathrm{softmax}\big(z_n[\text{A}],\, z_n[\text{B}],\, z_n[\text{C}]\big)_j.\]

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">labels</span> <span class="o">=</span> <span class="p">[</span><span class="s">"A"</span><span class="p">,</span> <span class="s">"B"</span><span class="p">,</span> <span class="s">"C"</span><span class="p">]</span>
<span class="n">listing</span> <span class="o">=</span> <span class="s">"</span><span class="se">\n</span><span class="s">"</span><span class="p">.</span><span class="n">join</span><span class="p">(</span><span class="sa">f</span><span class="s">"</span><span class="si">{</span><span class="n">l</span><span class="si">}</span><span class="s">. </span><span class="si">{</span><span class="n">o</span><span class="si">}</span><span class="s">"</span> <span class="k">for</span> <span class="n">l</span><span class="p">,</span> <span class="n">o</span> <span class="ow">in</span> <span class="nb">zip</span><span class="p">(</span><span class="n">labels</span><span class="p">,</span> <span class="n">options</span><span class="p">))</span>
<span class="n">prompt</span> <span class="o">=</span> <span class="n">chat</span><span class="p">(</span><span class="s">"Answer with a single letter: A, B, or C. Do not explain."</span><span class="p">,</span> <span class="sa">f</span><span class="s">"</span><span class="si">{</span><span class="n">question</span><span class="si">}</span><span class="se">\n\n</span><span class="si">{</span><span class="n">listing</span><span class="si">}</span><span class="s">"</span><span class="p">)</span>

<span class="n">ids</span> <span class="o">=</span> <span class="n">tok</span><span class="p">(</span><span class="n">prompt</span><span class="p">,</span> <span class="n">return_tensors</span><span class="o">=</span><span class="s">"pt"</span><span class="p">).</span><span class="n">input_ids</span>   <span class="c1"># (1, n)
</span><span class="k">with</span> <span class="n">torch</span><span class="p">.</span><span class="n">no_grad</span><span class="p">():</span>
    <span class="n">out</span> <span class="o">=</span> <span class="n">model</span><span class="p">(</span><span class="n">ids</span><span class="p">,</span> <span class="n">output_hidden_states</span><span class="o">=</span><span class="bp">True</span><span class="p">)</span>    <span class="c1"># one forward pass, no generate()
</span>
<span class="n">h</span> <span class="o">=</span> <span class="n">out</span><span class="p">.</span><span class="n">hidden_states</span><span class="p">[</span><span class="o">-</span><span class="mi">1</span><span class="p">][</span><span class="mi">0</span><span class="p">,</span> <span class="o">-</span><span class="mi">1</span><span class="p">]</span>                   <span class="c1"># (d,)  last latent vector
</span><span class="n">W_U</span> <span class="o">=</span> <span class="n">model</span><span class="p">.</span><span class="n">get_output_embeddings</span><span class="p">().</span><span class="n">weight</span>         <span class="c1"># (V, d)
</span><span class="n">z</span> <span class="o">=</span> <span class="n">W_U</span> <span class="o">@</span> <span class="n">h</span>                                        <span class="c1"># (V,)  same as out.logits[0, -1]
</span>
<span class="n">label_ids</span> <span class="o">=</span> <span class="p">[</span><span class="n">tok</span><span class="p">.</span><span class="n">encode</span><span class="p">(</span><span class="n">l</span><span class="p">,</span> <span class="n">add_special_tokens</span><span class="o">=</span><span class="bp">False</span><span class="p">)[</span><span class="mi">0</span><span class="p">]</span> <span class="k">for</span> <span class="n">l</span> <span class="ow">in</span> <span class="n">labels</span><span class="p">]</span>
<span class="n">p</span> <span class="o">=</span> <span class="n">z</span><span class="p">[</span><span class="n">label_ids</span><span class="p">].</span><span class="n">softmax</span><span class="p">(</span><span class="o">-</span><span class="mi">1</span><span class="p">)</span>                       <span class="c1"># softmax over the option labels only
</span></code></pre></div></div>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>A (speed up): 0.399
B (stop): 0.513
C (turn around): 0.088
</code></pre></div></div>

<p>One pass, no training. The method has two limitations, though.</p>

<p><strong>The labels must be single tokens.</strong> There are roughly 52 single-letter tokens, which caps the number of options.</p>

<p><strong>The model scores a letter, not an answer.</strong> The text “stop” never gets a probability. The letter <code class="language-plaintext highlighter-rouge">B</code> does, and the model has to work out from the prompt that <code class="language-plaintext highlighter-rouge">B</code> stands for “stop”. On top of that it has its own preferences for certain letters, whatever they stand for.</p>

<p>You can see this by changing which option gets which letter. The question and the three options stay the same, only the order in which they are listed changes:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>A. speed up   B. stop       C. turn around   -&gt;  p(stop) = 0.513
A. stop       B. speed up   C. turn around   -&gt;  p(stop) = 0.996
</code></pre></div></div>

<p>Over all six orderings, “stop” gets between 0.51 and 1.00. If the model only looked at the content of the options, we would get the same number six times.</p>

<h1 id="method-2-sequence-likelihood">Method 2: sequence likelihood</h1>

<p>This time we do not look only at a letter but at the probability of the entire option. The prompt contains only the question, and each option is written out as the full answer the model could give:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>The car should speed up.
The car should stop.
The car should turn around.
</code></pre></div></div>

<p>For each of these candidate answers we ask: how likely is it that the model produces exactly this text after the prompt? An answer is a sequence of tokens \(a = (a_1, \dots, a_m)\), so its probability is a product over tokens, or a sum in log space:</p>

\[\log p(a \mid \text{prompt}) = \sum_{t=1}^{m} \log p(a_t \mid \text{prompt}, a_{&lt;t}).\]

<p>To compute it, append the answer to the prompt and run one pass over prompt + answer. Position \(i\) predicts token \(i+1\), so the logits at positions \(n, \dots, n+m-1\) contain all \(m\) terms. Nothing is generated. Do this once per answer and take a softmax over the three log-likelihoods.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">prompt</span> <span class="o">=</span> <span class="n">chat</span><span class="p">(</span><span class="s">"Answer in one short sentence."</span><span class="p">,</span> <span class="n">question</span><span class="p">)</span>
<span class="n">n</span> <span class="o">=</span> <span class="n">tok</span><span class="p">(</span><span class="n">prompt</span><span class="p">,</span> <span class="n">return_tensors</span><span class="o">=</span><span class="s">"pt"</span><span class="p">).</span><span class="n">input_ids</span><span class="p">.</span><span class="n">shape</span><span class="p">[</span><span class="mi">1</span><span class="p">]</span>
<span class="n">answers</span> <span class="o">=</span> <span class="p">[</span><span class="sa">f</span><span class="s">"The car should </span><span class="si">{</span><span class="n">o</span><span class="si">}</span><span class="s">."</span> <span class="k">for</span> <span class="n">o</span> <span class="ow">in</span> <span class="n">options</span><span class="p">]</span>

<span class="n">scores</span> <span class="o">=</span> <span class="p">[]</span>
<span class="k">for</span> <span class="n">a</span> <span class="ow">in</span> <span class="n">answers</span><span class="p">:</span>
    <span class="n">ids</span> <span class="o">=</span> <span class="n">tok</span><span class="p">(</span><span class="n">prompt</span> <span class="o">+</span> <span class="n">a</span><span class="p">,</span> <span class="n">return_tensors</span><span class="o">=</span><span class="s">"pt"</span><span class="p">).</span><span class="n">input_ids</span>   <span class="c1"># prompt + answer in one pass
</span>    <span class="k">with</span> <span class="n">torch</span><span class="p">.</span><span class="n">no_grad</span><span class="p">():</span>
        <span class="n">logp</span> <span class="o">=</span> <span class="n">model</span><span class="p">(</span><span class="n">ids</span><span class="p">).</span><span class="n">logits</span><span class="p">[</span><span class="mi">0</span><span class="p">].</span><span class="n">log_softmax</span><span class="p">(</span><span class="o">-</span><span class="mi">1</span><span class="p">)</span>        <span class="c1"># (n + m, V)
</span>    <span class="n">target</span> <span class="o">=</span> <span class="n">ids</span><span class="p">[</span><span class="mi">0</span><span class="p">,</span> <span class="n">n</span><span class="p">:]</span>                                    <span class="c1"># the m answer tokens
</span>    <span class="c1"># the logits at position i predict token i + 1, hence the shift by one
</span>    <span class="n">logp_tokens</span> <span class="o">=</span> <span class="n">logp</span><span class="p">[</span><span class="n">n</span> <span class="o">-</span> <span class="mi">1</span> <span class="p">:</span> <span class="o">-</span><span class="mi">1</span><span class="p">].</span><span class="n">gather</span><span class="p">(</span><span class="o">-</span><span class="mi">1</span><span class="p">,</span> <span class="n">target</span><span class="p">[:,</span> <span class="bp">None</span><span class="p">]).</span><span class="n">squeeze</span><span class="p">(</span><span class="o">-</span><span class="mi">1</span><span class="p">)</span>
    <span class="n">scores</span><span class="p">.</span><span class="n">append</span><span class="p">(</span><span class="n">logp_tokens</span><span class="p">.</span><span class="nb">sum</span><span class="p">())</span>                       <span class="c1"># log p(answer | prompt)
</span>
<span class="n">p</span> <span class="o">=</span> <span class="n">torch</span><span class="p">.</span><span class="n">stack</span><span class="p">(</span><span class="n">scores</span><span class="p">).</span><span class="n">softmax</span><span class="p">(</span><span class="o">-</span><span class="mi">1</span><span class="p">)</span>
</code></pre></div></div>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>The car should speed up.     log p = -23.76   p = 0.000
The car should stop.         log p = -14.84   p = 1.000
The car should turn around.  log p = -29.70   p = 0.000
</code></pre></div></div>

<p>The answers can now be any text of any length, still without training. Two things to keep in mind:</p>

<p><strong>Cost.</strong> One pass per answer instead of one pass in total. The prompt is the same every time, so its KV cache can be reused and each extra pass only pays for the answer tokens.</p>

<p><strong>Length bias.</strong> Every token adds a negative term to the sum, so a longer answer gets a lower score just for being longer. The usual fix is to divide the sum by \(m\). Here the answers have five or six tokens, so it does not matter.</p>

<h1 id="method-3-pointer-head">Method 3: pointer head</h1>

<p>The prompt contains the question followed by the options, one per line, without letters:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>The traffic light is red. What should the self-driving car do?

Options:
speed up
stop
turn around
</code></pre></div></div>

<p>When using a pointer head, we run a forward pass over the prompt and select the latent vectors at four token positions (the number of options plus one) in the following way:</p>

<ul>
  <li>\(h_n\), the vector at the last token of the prompt. It has seen the question and all the options. This is the query: it encodes which answer the question calls for.</li>
  <li>\(h_{e_j}\) for \(j = 1, 2, 3\), the vector at the last token of option \(j\), where \(e_j\) is the index of that token. For “turn around” that is the vector at “around”. Because of the causal mask it is the first vector that has seen the whole option, so it is our summary of option \(j\).</li>
</ul>

<p>The pointer head consists of two matrices \(W_q, W_k \in \mathbb{R}^{r \times d}\). They project the query and the option summaries into a common \(r\)-dimensional space, where the score of an option is an inner product:</p>

\[s_j = (W_q h_n)^\top (W_k h_{e_j}), \qquad p = \mathrm{softmax}(s).\]

<p>This is the formula of an attention head, with the attention weights as the output: the head points at one of the options in the input, hence the name. \(W_q\) and \(W_k\) are trained with cross-entropy on questions with known answers.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">def</span> <span class="nf">latents</span><span class="p">(</span><span class="n">question</span><span class="p">,</span> <span class="n">options</span><span class="p">):</span>
    <span class="s">"""One pass over question + options. Returns h_n and the latent vector at the last token of each option."""</span>
    <span class="n">prompt</span> <span class="o">=</span> <span class="n">chat</span><span class="p">(</span><span class="s">"Pick one of the options."</span><span class="p">,</span> <span class="sa">f</span><span class="s">"</span><span class="si">{</span><span class="n">question</span><span class="si">}</span><span class="se">\n\n</span><span class="s">Options:</span><span class="se">\n</span><span class="s">"</span> <span class="o">+</span> <span class="s">"</span><span class="se">\n</span><span class="s">"</span><span class="p">.</span><span class="n">join</span><span class="p">(</span><span class="n">options</span><span class="p">))</span>
    <span class="n">enc</span> <span class="o">=</span> <span class="n">tok</span><span class="p">(</span><span class="n">prompt</span><span class="p">,</span> <span class="n">return_tensors</span><span class="o">=</span><span class="s">"pt"</span><span class="p">,</span> <span class="n">return_offsets_mapping</span><span class="o">=</span><span class="bp">True</span><span class="p">)</span>

    <span class="c1"># index of the last token of each option, found via character offsets
</span>    <span class="n">offsets</span> <span class="o">=</span> <span class="n">enc</span><span class="p">.</span><span class="n">offset_mapping</span><span class="p">[</span><span class="mi">0</span><span class="p">].</span><span class="n">tolist</span><span class="p">()</span>
    <span class="n">pos</span><span class="p">,</span> <span class="n">start</span> <span class="o">=</span> <span class="p">[],</span> <span class="n">prompt</span><span class="p">.</span><span class="n">index</span><span class="p">(</span><span class="s">"Options:"</span><span class="p">)</span>
    <span class="k">for</span> <span class="n">o</span> <span class="ow">in</span> <span class="n">options</span><span class="p">:</span>
        <span class="n">end</span> <span class="o">=</span> <span class="n">prompt</span><span class="p">.</span><span class="n">index</span><span class="p">(</span><span class="n">o</span><span class="p">,</span> <span class="n">start</span><span class="p">)</span> <span class="o">+</span> <span class="nb">len</span><span class="p">(</span><span class="n">o</span><span class="p">)</span>
        <span class="n">pos</span><span class="p">.</span><span class="n">append</span><span class="p">(</span><span class="nb">next</span><span class="p">(</span><span class="n">i</span> <span class="k">for</span> <span class="n">i</span><span class="p">,</span> <span class="p">(</span><span class="n">s</span><span class="p">,</span> <span class="n">e</span><span class="p">)</span> <span class="ow">in</span> <span class="nb">enumerate</span><span class="p">(</span><span class="n">offsets</span><span class="p">)</span> <span class="k">if</span> <span class="n">s</span> <span class="o">&lt;</span> <span class="n">end</span> <span class="o">&lt;=</span> <span class="n">e</span><span class="p">))</span>
        <span class="n">start</span> <span class="o">=</span> <span class="n">end</span>

    <span class="k">with</span> <span class="n">torch</span><span class="p">.</span><span class="n">no_grad</span><span class="p">():</span>
        <span class="n">H</span> <span class="o">=</span> <span class="n">model</span><span class="p">(</span><span class="n">enc</span><span class="p">.</span><span class="n">input_ids</span><span class="p">,</span> <span class="n">output_hidden_states</span><span class="o">=</span><span class="bp">True</span><span class="p">).</span><span class="n">hidden_states</span><span class="p">[</span><span class="o">-</span><span class="mi">1</span><span class="p">][</span><span class="mi">0</span><span class="p">]</span>   <span class="c1"># (n, d)
</span>    <span class="k">return</span> <span class="n">H</span><span class="p">[</span><span class="o">-</span><span class="mi">1</span><span class="p">],</span> <span class="n">H</span><span class="p">[</span><span class="n">pos</span><span class="p">]</span>

<span class="n">h_query</span><span class="p">,</span> <span class="n">h_options</span> <span class="o">=</span> <span class="n">latents</span><span class="p">(</span><span class="n">question</span><span class="p">,</span> <span class="n">options</span><span class="p">)</span>   <span class="c1"># (d,) and (num_options, d)
</span>
<span class="n">torch</span><span class="p">.</span><span class="n">manual_seed</span><span class="p">(</span><span class="mi">0</span><span class="p">)</span>
<span class="n">d</span><span class="p">,</span> <span class="n">r</span> <span class="o">=</span> <span class="n">h_query</span><span class="p">.</span><span class="n">shape</span><span class="p">[</span><span class="o">-</span><span class="mi">1</span><span class="p">],</span> <span class="mi">256</span>
<span class="n">W_q</span> <span class="o">=</span> <span class="n">torch</span><span class="p">.</span><span class="n">nn</span><span class="p">.</span><span class="n">Linear</span><span class="p">(</span><span class="n">d</span><span class="p">,</span> <span class="n">r</span><span class="p">,</span> <span class="n">bias</span><span class="o">=</span><span class="bp">False</span><span class="p">)</span>   <span class="c1"># trainable
</span><span class="n">W_k</span> <span class="o">=</span> <span class="n">torch</span><span class="p">.</span><span class="n">nn</span><span class="p">.</span><span class="n">Linear</span><span class="p">(</span><span class="n">d</span><span class="p">,</span> <span class="n">r</span><span class="p">,</span> <span class="n">bias</span><span class="o">=</span><span class="bp">False</span><span class="p">)</span>   <span class="c1"># trainable
</span>
<span class="k">def</span> <span class="nf">scores</span><span class="p">(</span><span class="n">h_query</span><span class="p">,</span> <span class="n">h_options</span><span class="p">):</span>
    <span class="k">return</span> <span class="n">W_k</span><span class="p">(</span><span class="n">h_options</span><span class="p">)</span> <span class="o">@</span> <span class="n">W_q</span><span class="p">(</span><span class="n">h_query</span><span class="p">)</span> <span class="o">/</span> <span class="n">r</span><span class="o">**</span><span class="mf">0.5</span>   <span class="c1"># (num_options,) one score per option
</span>
<span class="n">s</span> <span class="o">=</span> <span class="n">scores</span><span class="p">(</span><span class="n">h_query</span><span class="p">,</span> <span class="n">h_options</span><span class="p">)</span>
<span class="n">target</span> <span class="o">=</span> <span class="n">torch</span><span class="p">.</span><span class="n">tensor</span><span class="p">([</span><span class="mi">1</span><span class="p">])</span>                <span class="c1"># index of the correct option: stop
</span><span class="n">loss</span> <span class="o">=</span> <span class="n">torch</span><span class="p">.</span><span class="n">nn</span><span class="p">.</span><span class="n">functional</span><span class="p">.</span><span class="n">cross_entropy</span><span class="p">(</span><span class="n">s</span><span class="p">[</span><span class="bp">None</span><span class="p">],</span> <span class="n">target</span><span class="p">)</span>
</code></pre></div></div>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>p    = [0.898, 0.029, 0.073]
loss = 3.531
</code></pre></div></div>

<p>The matrices are untrained, so these probabilities mean nothing yet. What the code does show is how little is needed. A single forward pass scores all options at once, however many there are and however long they are. And the only new parameters are the two matrices, \(2rd = 524{,}288\) in total for \(r = 256\), next to the 0.6 billion of the backbone.</p>

<p>One consequence of the causal mask: \(h_{e_j}\) only sees what comes before it. So the question has to come before the options, and option \(j\) knows about options \(1, \dots, j-1\) but not the later ones. Only the query \(h_n\) sees everything.</p>

<h2 id="how-is-this-different-from-a-classifier">How is this different from a classifier?</h2>

<p>Rewrite the score as</p>

\[s_j = w_j^\top h_n, \qquad w_j = W_q^\top W_k\, h_{e_j}.\]

<p>This has the form of the classifier from the beginning, \(W_C h_n\), with \(w_j\) as the row for class \(j\). The difference is where that row comes from. In the classifier, \(w_j\) is a trained parameter, so each class is stored in the weights. Here \(w_j\) is computed by the backbone from the text of option \(j\). The trained part, \(W_q^\top W_k\), contains nothing about any particular option. It only learns how to compare a query with an option summary.</p>

<p>That is why the head can handle options it has never seen, and any number of them: a new option gets its \(w_j\) from the same forward pass. How well this works in practice depends on how varied the training questions are, and that is an empirical question.</p>

<p>People in vision will recognise the setup. A linear probe on a frozen DINO model is the classifier version: a strong backbone, fixed features, and a small trained \(W_C\) with fixed classes. The pointer head makes the same bet that the features already contain the answer and a small readout is enough. It just compares features with features instead of features with stored class vectors.</p>

<h1 id="a-training-loop-that-overfits">A training loop (that overfits)</h1>

<p>To close the loop, here is the training of \(W_q\) and \(W_k\) on our single question. The backbone stays frozen, so its latent vectors are computed once and reused in every step.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">opt</span> <span class="o">=</span> <span class="n">torch</span><span class="p">.</span><span class="n">optim</span><span class="p">.</span><span class="n">Adam</span><span class="p">([</span><span class="o">*</span><span class="n">W_q</span><span class="p">.</span><span class="n">parameters</span><span class="p">(),</span> <span class="o">*</span><span class="n">W_k</span><span class="p">.</span><span class="n">parameters</span><span class="p">()],</span> <span class="n">lr</span><span class="o">=</span><span class="mf">1e-5</span><span class="p">)</span>

<span class="k">for</span> <span class="n">step</span> <span class="ow">in</span> <span class="nb">range</span><span class="p">(</span><span class="mi">31</span><span class="p">):</span>
    <span class="n">s</span> <span class="o">=</span> <span class="n">scores</span><span class="p">(</span><span class="n">h_query</span><span class="p">,</span> <span class="n">h_options</span><span class="p">)</span>
    <span class="n">loss</span> <span class="o">=</span> <span class="n">torch</span><span class="p">.</span><span class="n">nn</span><span class="p">.</span><span class="n">functional</span><span class="p">.</span><span class="n">cross_entropy</span><span class="p">(</span><span class="n">s</span><span class="p">[</span><span class="bp">None</span><span class="p">],</span> <span class="n">target</span><span class="p">)</span>
    <span class="n">opt</span><span class="p">.</span><span class="n">zero_grad</span><span class="p">()</span>
    <span class="n">loss</span><span class="p">.</span><span class="n">backward</span><span class="p">()</span>
    <span class="n">opt</span><span class="p">.</span><span class="n">step</span><span class="p">()</span>
</code></pre></div></div>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>step  0   loss = 3.531   p(stop) = 0.029
step  5   loss = 0.168   p(stop) = 0.845
step 10   loss = 0.005   p(stop) = 0.995
step 20   loss = 0.000   p(stop) = 1.000
step 30   loss = 0.000   p(stop) = 1.000
</code></pre></div></div>

<p>The loss goes to zero within a few steps. That is not a success. With one training example and half a million parameters, the head completely overfits to this one question. You can see it by changing the question so that “stop” is no longer correct:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">green</span> <span class="o">=</span> <span class="s">"The traffic light is green. What should the self-driving car do?"</span>
<span class="n">new_options</span> <span class="o">=</span> <span class="p">[</span><span class="s">"drive on"</span><span class="p">,</span> <span class="s">"stop"</span><span class="p">,</span> <span class="s">"turn around"</span><span class="p">]</span>
<span class="n">p</span> <span class="o">=</span> <span class="n">scores</span><span class="p">(</span><span class="o">*</span><span class="n">latents</span><span class="p">(</span><span class="n">green</span><span class="p">,</span> <span class="n">new_options</span><span class="p">)).</span><span class="n">softmax</span><span class="p">(</span><span class="o">-</span><span class="mi">1</span><span class="p">)</span>
</code></pre></div></div>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>drive on     0.000
stop         1.000
turn around  0.000
</code></pre></div></div>

<p>The head has learned “the answer is stop”, not how to match a question with an option. To learn that, it needs many different questions, with different options and with the correct answer in different places.</p>

<h1 id="summary">Summary</h1>

<table>
  <thead>
    <tr>
      <th> </th>
      <th>passes</th>
      <th>training</th>
      <th>options</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Label logits</td>
      <td>1</td>
      <td>none</td>
      <td>single-token labels</td>
    </tr>
    <tr>
      <td>Sequence likelihood</td>
      <td>1 per option</td>
      <td>none</td>
      <td>any text, length bias</td>
    </tr>
    <tr>
      <td>Pointer head</td>
      <td>1</td>
      <td>\(W_q\), \(W_k\)</td>
      <td>any text, any number</td>
    </tr>
  </tbody>
</table>

<p>As far as I can tell, the pointer head is what Jev-style models use, together with LoRA on the backbone. LoRA (low-rank adaptation) freezes a weight matrix \(W\) and trains a low-rank update \(W + BA\) instead, with \(B \in \mathbb{R}^{d \times k}\), \(A \in \mathbb{R}^{k \times d}\) and \(k \ll d\). That is a small fraction of the parameters, and it lets the backbone adapt its latent vectors to the head.</p>

<p>In the next post we train the pointer head on a simple robotics example. After that I want to look at vision transformers for Jev-like models.</p>]]></content><author><name>Patrick Huembeli</name></author><summary type="html"><![CDATA[Jev caused quite a stir over the last few days, and some people even claim we can use it for robot control. This is the first of a few posts explaining the idea behind it. I’ll keep it short and assume you know your math and ML.]]></summary></entry><entry><title type="html">How I work in ML in 2026</title><link href="https://phuembeli.com/blog/2026/03/19/working-in-ml-in-2026/" rel="alternate" type="text/html" title="How I work in ML in 2026" /><published>2026-03-19T00:00:00+00:00</published><updated>2026-03-19T00:00:00+00:00</updated><id>https://phuembeli.com/blog/2026/03/19/working-in-ml-in-2026</id><content type="html" xml:base="https://phuembeli.com/blog/2026/03/19/working-in-ml-in-2026/"><![CDATA[<h1 id="what-skills-do-i-need">What skills do I need?</h1>

<p>I know I should take the discussions on social media with a huge grain of salt. But they’re hard to miss and there’s probably some truth in there. The way we work is going to change a lot over the next few years. Parts of jobs, probably not entire jobs, will get automated. We already see it in software engineering. At my level of SWE, I’m barely writing any code myself anymore. This might be different for more security-sensitive parts of the architecture, or for people working on bigger, more stable codebases. But I’m pretty convinced most SWEs work the way I do now: find a bug or think of a feature, go to Claude Code planning mode, prompt back and forth a few times, then let it rip. Some testing. Done.</p>

<p>The loud version of this on social media gets exaggerated to “software engineering is dead, no need to hire SWEs, fire 50% of the workforce.” That’s wishful thinking from people selling a narrative. My job has changed, but I don’t have less work. My focus has shifted. I’m writing this mostly for myself, to figure out how I want to use my skills going forward and what new ones I should pick up.</p>

<h1 id="what-am-i-good-at">What am I good at?</h1>

<p>I consider myself an all-rounder. I did an apprenticeship as an electronics technician when I was 16: soldering, programming microprocessors, designing simple circuits, plus 10–14 weeks of mechanical shop work where we learned to drill, turn parts on a lathe, mill, and do a bit of CNC programming.</p>

<p>I’m not good at any of these anymore, but I’m sure I could pick them up again pretty quickly. More importantly, it was my first contact with the real business world, and it shaped how I approached university. To be frank, by the end of the apprenticeship I knew I didn’t want to stay in that profession at all, so I went back to do my high-school degree and then studied physics. Physics felt like a good bet because it teaches relevant math and a bit of programming, and I was vaguely considering finance later, where physicists are in demand.</p>

<p>While studying physics I started to enjoy the process of learning for its own sake, which is why I wanted to do a PhD. There’s a lot of discourse right now claiming PhDs are useless and you should just go to industry. I disagree, especially if you’re someone like me who just wants to keep learning. Being surrounded by brilliant people for more than five years was transformative. You can do that in industry too, but you’ll have much less freedom to chase rabbit holes; most of the time you’re on a tighter schedule. That has its own advantages, of course. If you have a clear career goal and a PhD is just a means to get there, you probably won’t enjoy the PhD itself. For context: I never planned to stay in academia. In Switzerland a PhD is a fairly common advanced degree.</p>

<p>After a one-year postdoc, I went to industry. First to a startup called Menten, where we tried to show how quantum computing could improve or accelerate protein design. The premise was cursed from the start: classical optimization on a quantum computer just doesn’t work. Menten realized this and pivoted away from quantum at the same time I did, and I had a gig lined up at Extropic where I started as a Staff Scientist. Extropic was the perfect blend of research and industry: figuring out how to build a new computing paradigm and designing algorithms for it. I left when they decided to relocate everyone to the US, which wasn’t possible for me at the time.</p>

<p>That brought me to Axiomatic as AI Lead, where I had a deep dive on agentic AI. The pace was insane. Late 2024 to December 2025 felt like a different decade every quarter. December 2025 in particular was when Claude Code became really good. So good that most people just started running it with <code class="language-plaintext highlighter-rouge">--dangerously-skip-permissions</code>. It changed the way I work.</p>

<p>I’m now at Noumenal and the work is extremely diverse. From general discussions about the direction of the company, to grant applications, to writing code for our robots, to testing them in the lab. We all do everything, and AI agents make that genuinely possible. I let an agent plan my meetings, give me briefs and reminders ahead of time, and help me track my own and others’ tasks. The main operational benefit is that it lowers the cost of switching tasks. I don’t yet let it do many things autonomously (like sending emails), but I use it as a second memory and as a virtual assistant that brings me up to speed before I context-switch into something.</p>

<p>On the development side, AI agents do a lot of the heavy lifting. Anyone who has worked with ROS knows how painful that software is, and an agent can just take care of wrapping our Python scripts into ROS nodes for me. I had never used ROS before this job, and I can take care of all of it myself. Same for quick prototypes of frontend interfaces, scripts to parse a new dataset, or one-off tools I need for an afternoon. All of that previously would have meant either a context-switch into an unfamiliar stack or asking someone else to do it.</p>

<h1 id="spare-a-thought-for-the-juniors">Spare a thought for the juniors</h1>

<p>I don’t envy the next generation. The uncertainty is big, and it’s genuinely hard to learn deep skills if you never get the chance to explore. Thinking about code, really thinking, including on an architectural level, is what teaches you to code. Watching an agent produce something that works isn’t the same thing.</p>

<p>It’s also why a lot of vibe-coded projects fail. People who don’t really know how software works make big architectural mistakes from the very beginning, and agents aren’t yet good enough to fix those ad hoc. Worse, a messy codebase actively makes the agent perform worse, which makes the codebase messier, which makes the agent worse. You end up in a hole you can’t get out of. Jank in, jank out.</p>

<p>I don’t have a clean answer to either of these. I think there will still be deeply skilled engineers in five years, and they’ll be more valuable, not less. But the path to becoming one looks much harder from where I’m standing.</p>]]></content><author><name>Patrick Huembeli</name></author><summary type="html"><![CDATA[What skills do I need?]]></summary></entry></feed>