<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>The Rimembra blog</title>
    <link>https://rimembra.ai/blog/</link>
    <description>Research, ideas and progress from the team building AI that learns how you think.</description>
    <language>en</language>
    <atom:link href="https://rimembra.ai/blog/rss.xml" rel="self" type="application/rss+xml"/>
    <lastBuildDate>Wed, 30 Sep 2026 00:00:00 GMT</lastBuildDate>
    <item>
      <title>You can own the model weights and still lose the personality.</title>
      <link>https://rimembra.ai/blog/owning-the-weights/</link>
      <guid isPermaLink="true">https://rimembra.ai/blog/owning-the-weights/</guid>
      <pubDate>Wed, 30 Sep 2026 00:00:00 GMT</pubDate>
      <description>Preserving a digital representation of someone requires control over the entire system that speaks on their behalf. Research into fine-tuning helps explain why.</description>
      <content:encoded><![CDATA[<p><img src="https://rimembra.ai/blog/images/owning-the-weights/hero-field.png" alt="Many strands of context converge on a single response; one strand drifts away."></p><p>An AI representation of a person can change its judgment without a single model weight changing.</p>
<p>A retrieval system selects different memories. A conversation summary drops an important qualification. An instruction places greater emphasis on reassurance. The model remains identical, yet its responses become more agreeable, less cautious, or more certain than the person’s evidence supports.</p>
<p>This is a central engineering challenge for human identity preservation. The behaviour we experience as a digital Persona emerges from the interaction between a trained model, personal evidence, instructions, conversation history, and the software generating the response.</p>
<p>At Rimembra, that distinction shapes our research question: <strong>which interventions make a Persona more faithful to an individual, and which merely make it more convincing?</strong></p>
<h2 id="where-the-response-comes-from">Where the response comes from</h2>
<p>A language model’s weights are the numerical parameters learned during training. Retaining an exact checkpoint gives a developer control over that trained model, including the ability to preserve it independently of a provider’s future releases.</p>
<p>But those weights do not specify everything the system will say.</p>
<p>The model also receives a constructed context: instructions, selected personal evidence, and information from the current conversation. Retrieval determines which evidence arrives. Summarisation determines what survives. Inference settings influence how the next tokens are selected.</p>
<figure><img src="https://rimembra.ai/blog/images/owning-the-weights/figure-1.png" alt="Figure 1. Personal evidence and conversation history pass through retrieval and context selection. With instructions they form the assembled model input. Response generation combines that input with the model weights and adapters, and with inference settings and serving software. Only the weights and adapters are preserved by owning the model." loading="lazy" decoding="async"></figure>
<p>Consider someone who described starting a business as the best decision they ever made. Elsewhere, they explained that they waited until they had enough savings to protect their family.</p>
<p>A system that retrieves the enthusiasm but misses the qualification may produce advice that sounds entrepreneurial while overlooking the condition behind the original decision.</p>
<figure><img src="https://rimembra.ai/blog/images/owning-the-weights/figure-2.png" alt="Illustrative example. Only the enthusiastic memory is retrieved, so the Persona says &quot;Go for it.&quot; With the savings condition preserved, it says yes once the family is protected." loading="lazy" decoding="async"></figure>
<p>That is a hypothetical example, but the sensitivity of language models to evidence presentation is experimentally documented.</p>
<h2 id="the-same-weights-a-different-result">The same weights, a different result</h2>
<p>In <em>Lost in the Middle</em>, Liu and colleagues examined how language models used information placed at different positions within their input.<sup class="fn"><a href="#ref-1" id="cite-1" aria-label="Reference 1">1</a></sup></p>
<p>In one experiment, GPT‑3.5‑Turbo received a question and 20 documents, one containing the answer. Its accuracy changed substantially depending on where that document appeared:</p>
<figure><img src="https://rimembra.ai/blog/images/owning-the-weights/figure-3.png" alt="Figure 2. Line chart of answer accuracy by position: 75.8% first, 57.2% fifth, 53.8% tenth, 55.4% fifteenth, 63.2% twentieth. Closed-book accuracy with no documents was 56.1%." loading="lazy" decoding="async"></figure>
<div class="table-wrap"><table>
<thead><tr><th scope="col">Position of the answer-containing document</th><th scope="col" class="right">Answer accuracy</th></tr></thead>
<tbody>
<tr><td>First</td><td class="right">75.8%</td></tr>
<tr><td>Fifth</td><td class="right">57.2%</td></tr>
<tr><td>Tenth</td><td class="right">53.8%</td></tr>
<tr><td>Fifteenth</td><td class="right">55.4%</td></tr>
<tr><td>Twentieth</td><td class="right">63.2%</td></tr>
</tbody>
</table></div>
<p>Moving the relevant document from first to tenth position corresponded to a <strong>22 percentage-point reduction in accuracy</strong>, without retraining the model. These results concern an older model and a question-answering benchmark; they are not measurements of digital Persona fidelity.</p>
<p>For Rimembra, the implication is a testable question: does a Persona preserve important qualifications when evidence is reordered, the conversation grows longer, or competing information enters the context?</p>
<p>Possessing the right information and using it faithfully are separate engineering problems.</p>
<h2 id="what-changing-the-weights-can-achieve">What changing the weights can achieve</h2>
<p>Fine-tuning offers a way to alter how a model responds to its inputs. For personal representation, there are two distinct objectives worth investigating.</p>
<p>One is teaching a model to express a particular individual’s patterns. Another is teaching it to use personal evidence more faithfully across individuals—for example, retaining qualifications, respecting uncertainty, and distinguishing recorded beliefs from inferred ones.</p>
<p>Those objectives require different training examples and different measures of success.</p>
<p>LoRA, introduced by Hu and colleagues, provides a practical mechanism for such experiments.<sup class="fn"><a href="#ref-2" id="cite-2" aria-label="Reference 2">2</a></sup> It keeps the original weight matrices frozen and learns an update through smaller trainable matrices. During inference, that learned update contributes to the model’s computation.</p>
<figure><img src="https://rimembra.ai/blog/images/owning-the-weights/figure-4.png" alt="Figure 3. Simplified LoRA mechanism for one adapted layer: input activations pass through frozen weights W and, in parallel, trainable matrices A and B, whose scaled update is added to the output. A grid of 128 cells shows full adaptation of one 4096 × 4096 matrix; a rank-16 adapter fills one cell." loading="lazy" decoding="async"></figure>
<p>For a single 4096 × 4096 weight matrix, full adaptation would involve 16,777,216 trainable parameters. A rank-16 LoRA adapter uses 131,072—approximately 0.78% as many. This is an illustrative per-matrix calculation, excluding biases, rather than a whole-model cost estimate.</p>
<p>A smaller update is easier to isolate experimentally. Its size, however, tells us little about whether the resulting behaviour remains faithful to someone.</p>
<h2 id="training-can-improve-consistency-and-change-more-than-intended">Training can improve consistency—and change more than intended</h2>
<p>There is evidence that training can improve persona adherence.</p>
<p>In <em>Consistently Simulating Human Personas with Multi-Turn Reinforcement Learning</em>, Abdulhai and colleagues reported reducing inconsistency by <strong>over 55%</strong> in experiments involving simulated conversation partners, students, and patients.<sup class="fn"><a href="#ref-3" id="cite-3" aria-label="Reference 3">3</a></sup> Their evaluations examined alignment with the supplied persona, contradictions within conversations, and consistency when answering questions about persona characteristics. These were simulated roles, rather than validated reconstructions of real individuals.</p>
<p>That finding makes training a serious research avenue. It also leaves a crucial question open: does stronger adherence to a persona description produce a more accurate representation of the person?</p>
<p>Someone can be cautious with household finances and adventurous in other circumstances. Training a model to be uniformly cautious could improve a simplistic consistency score while erasing that distinction.</p>
<p>Other research demonstrates why evaluation must extend beyond the training objective.</p>
<p>In the <em>Emergent Misalignment</em> study published in <em>Nature</em>, Betley and colleagues found that fine-tuning GPT‑4o on insecure-code examples produced broader behavioural changes outside coding.<sup class="fn"><a href="#ref-4" id="cite-4" aria-label="Reference 4">4</a></sup> On selected evaluation questions, the adapted model generated misaligned responses approximately <strong>20% of the time</strong>, compared with <strong>0%</strong> for the original model. Controls involving secure code or explicitly educational requests for insecure code did not produce similar behaviour.</p>
<p>This does not establish that ordinary personalisation causes those outcomes. It demonstrates that a narrow training intervention can have effects beyond its intended domain.</p>
<p>For a Persona, that motivates a specific experiment: if we train for familiar expression, do we also change its willingness to disagree, its handling of uncertainty, or the priorities reflected in its advice?</p>
<h2 id="testing-where-personalisation-belongs">Testing where personalisation belongs</h2>
<p>Rimembra’s initial design keeps the base model frozen while evaluating the capture, representation, and use of personal evidence. This establishes a baseline against which future adaptation can be assessed.</p>
<p>Our proposed research would then separate the contribution of better evidence organisation from the contribution of changing weights.</p>
<figure><img src="https://rimembra.ai/blog/images/owning-the-weights/figure-5.png" alt="Proposed design, not reported results. A two-by-two grid: basic or structured evidence presentation, each with a base or adapted model (B0, B1, S0, S1). Every condition answers the same held-out scenarios." loading="lazy" decoding="async"></figure>
<p>Within each evidence condition, comparing the base and adapted models would help estimate the additional contribution of training. Comparing evidence conditions would help establish whether the model benefits from a richer representation. The combined comparison could reveal whether adaptation mainly helps the model use that representation.</p>
<p>The most revealing tests would examine unfamiliar situations and the conditions that change a person’s judgment.</p>
<p>Suppose someone supports helping a relative financially when their own household’s essential expenses remain protected. A useful evaluation would vary that condition while keeping the rest of the scenario similar.</p>
<figure><img src="https://rimembra.ai/blog/images/owning-the-weights/figure-6.png" alt="Illustrative example. When essentials are covered, the person, a uniformly warm Persona and a faithful Persona all say help. When essentials are at risk, the person says not yet; the warm Persona still says help and misses the condition; the faithful Persona matches." loading="lazy" decoding="async"></figure>
<p>A Persona that always recommends helping might sound warm and consistent. A more faithful one would need to preserve the person’s qualification.</p>
<p>We would assess those answers along separate dimensions:</p>
<div class="table-wrap"><table>
<thead><tr><th scope="col">Dimension</th><th scope="col">What the evaluation asks</th></tr></thead>
<tbody>
<tr><td>Personal facts</td><td>Are claims supported by the person’s evidence?</td></tr>
<tr><td>Judgment</td><td>Does the answer reflect their priorities and relevant conditions?</td></tr>
<tr><td>Expression</td><td>Is the communication recognisable without exaggeration?</td></tr>
<tr><td>Uncertainty</td><td>Does the system acknowledge what it cannot establish?</td></tr>
<tr><td>Preference</td><td>Does the reviewer like the answer, independently of its fidelity?</td></tr>
</tbody>
</table></div>
<p>The person’s reference answers should be recorded before they see model outputs. Final testing should use scenarios withheld from training and development. Blinded comparisons and repeated runs would help distinguish systematic changes from ordinary variation.</p>
<p>These are proposed experiments. Their value will depend on what they demonstrate—including where adaptation provides no benefit or introduces regressions.</p>
<h2 id="preserve-the-system-then-test-the-representation">Preserve the system, then test the representation</h2>
<p>A meaningful preservation strategy needs a versioned record of the model and any adapters, the personal-evidence snapshot, retrieval and context construction, instructions, and inference environment.</p>
<figure><img src="https://rimembra.ai/blog/images/owning-the-weights/figure-8.png" alt="Figure 5. Five stacked layers recorded together: model and adapters, personal-evidence snapshot, retrieval and context construction, instructions, and inference environment. Traceable is not the same as faithful." loading="lazy" decoding="async"></figure>
<p>That record makes changes traceable. It supports investigation and recovery when behaviour shifts. It does not, by itself, establish that the resulting Persona is faithful.</p>
<p>The scientific challenge is to connect a system’s responses to independent evidence about the individual: their judgments, qualifications, contradictions, and uncertainty.</p>
<p>Owning the weights is an important form of control. Fine-tuning is a promising experimental tool. Neither substitutes for testing the complete system.</p>
<figure class="closing-quote"><blockquote><p>The standard is whether the person remains recognisable in the judgments the system expresses—not simply in the language it uses.</p></blockquote></figure>
<hr>
<h3 id="references">References</h3>
<ol class="references">
<li id="ref-1"><span>Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F. and Liang, P. <em>Lost in the Middle: How Language Models Use Long Contexts.</em> Transactions of the ACL, 2024. <a href="https://arxiv.org/abs/2307.03172" target="_blank" rel="noopener">arXiv:2307.03172</a> <a class="back-ref" href="#cite-1" aria-label="Back to the text">&#8617;&#xFE0E;</a></span></li>
<li id="ref-2"><span>Hu, E. J. et al. <em>LoRA: Low-Rank Adaptation of Large Language Models.</em> ICLR 2022. <a href="https://arxiv.org/abs/2106.09685" target="_blank" rel="noopener">arXiv:2106.09685</a> <a class="back-ref" href="#cite-2" aria-label="Back to the text">&#8617;&#xFE0E;</a></span></li>
<li id="ref-3"><span>Abdulhai, M. et al. <em>Consistently Simulating Human Personas with Multi-Turn Reinforcement Learning.</em> NeurIPS 2025. <a href="https://arxiv.org/abs/2511.00222" target="_blank" rel="noopener">arXiv:2511.00222</a> <a class="back-ref" href="#cite-3" aria-label="Back to the text">&#8617;&#xFE0E;</a></span></li>
<li id="ref-4"><span>Betley, J., Warncke, N., Sztyber-Betley, A. et al. <em>Training large language models on narrow tasks can lead to broad misalignment.</em> Nature, 2026. Preprint: <a href="https://arxiv.org/abs/2502.17424" target="_blank" rel="noopener">arXiv:2502.17424</a> <a class="back-ref" href="#cite-4" aria-label="Back to the text">&#8617;&#xFE0E;</a></span></li>
</ol>]]></content:encoded>
    </item>
  </channel>
</rss>
