You can own the model weights and still lose the personality.
Preserving a digital representation of someone requires control over the entire system that speaks on their behalf. Research into fine-tuning helps explain why.

An AI representation of a person can change its judgment without a single model weight changing.
A retrieval system selects different memories. A conversation summary drops an important qualification. An instruction places greater emphasis on reassurance. The model remains identical, yet its responses become more agreeable, less cautious, or more certain than the person’s evidence supports.
This is a central engineering challenge for human identity preservation. The behaviour we experience as a digital Persona emerges from the interaction between a trained model, personal evidence, instructions, conversation history, and the software generating the response.
At Rimembra, that distinction shapes our research question: which interventions make a Persona more faithful to an individual, and which merely make it more convincing?
Where the response comes from
A language model’s weights are the numerical parameters learned during training. Retaining an exact checkpoint gives a developer control over that trained model, including the ability to preserve it independently of a provider’s future releases.
But those weights do not specify everything the system will say.
The model also receives a constructed context: instructions, selected personal evidence, and information from the current conversation. Retrieval determines which evidence arrives. Summarisation determines what survives. Inference settings influence how the next tokens are selected.

Consider someone who described starting a business as the best decision they ever made. Elsewhere, they explained that they waited until they had enough savings to protect their family.
A system that retrieves the enthusiasm but misses the qualification may produce advice that sounds entrepreneurial while overlooking the condition behind the original decision.

That is a hypothetical example, but the sensitivity of language models to evidence presentation is experimentally documented.
The same weights, a different result
In Lost in the Middle, Liu and colleagues examined how language models used information placed at different positions within their input.1
In one experiment, GPT‑3.5‑Turbo received a question and 20 documents, one containing the answer. Its accuracy changed substantially depending on where that document appeared:

| Position of the answer-containing document | Answer accuracy |
|---|---|
| First | 75.8% |
| Fifth | 57.2% |
| Tenth | 53.8% |
| Fifteenth | 55.4% |
| Twentieth | 63.2% |
Moving the relevant document from first to tenth position corresponded to a 22 percentage-point reduction in accuracy, without retraining the model. These results concern an older model and a question-answering benchmark; they are not measurements of digital Persona fidelity.
For Rimembra, the implication is a testable question: does a Persona preserve important qualifications when evidence is reordered, the conversation grows longer, or competing information enters the context?
Possessing the right information and using it faithfully are separate engineering problems.
What changing the weights can achieve
Fine-tuning offers a way to alter how a model responds to its inputs. For personal representation, there are two distinct objectives worth investigating.
One is teaching a model to express a particular individual’s patterns. Another is teaching it to use personal evidence more faithfully across individuals—for example, retaining qualifications, respecting uncertainty, and distinguishing recorded beliefs from inferred ones.
Those objectives require different training examples and different measures of success.
LoRA, introduced by Hu and colleagues, provides a practical mechanism for such experiments.2 It keeps the original weight matrices frozen and learns an update through smaller trainable matrices. During inference, that learned update contributes to the model’s computation.

For a single 4096 × 4096 weight matrix, full adaptation would involve 16,777,216 trainable parameters. A rank-16 LoRA adapter uses 131,072—approximately 0.78% as many. This is an illustrative per-matrix calculation, excluding biases, rather than a whole-model cost estimate.
A smaller update is easier to isolate experimentally. Its size, however, tells us little about whether the resulting behaviour remains faithful to someone.
Training can improve consistency—and change more than intended
There is evidence that training can improve persona adherence.
In Consistently Simulating Human Personas with Multi-Turn Reinforcement Learning, Abdulhai and colleagues reported reducing inconsistency by over 55% in experiments involving simulated conversation partners, students, and patients.3 Their evaluations examined alignment with the supplied persona, contradictions within conversations, and consistency when answering questions about persona characteristics. These were simulated roles, rather than validated reconstructions of real individuals.
That finding makes training a serious research avenue. It also leaves a crucial question open: does stronger adherence to a persona description produce a more accurate representation of the person?
Someone can be cautious with household finances and adventurous in other circumstances. Training a model to be uniformly cautious could improve a simplistic consistency score while erasing that distinction.
Other research demonstrates why evaluation must extend beyond the training objective.
In the Emergent Misalignment study published in Nature, Betley and colleagues found that fine-tuning GPT‑4o on insecure-code examples produced broader behavioural changes outside coding.4 On selected evaluation questions, the adapted model generated misaligned responses approximately 20% of the time, compared with 0% for the original model. Controls involving secure code or explicitly educational requests for insecure code did not produce similar behaviour.
This does not establish that ordinary personalisation causes those outcomes. It demonstrates that a narrow training intervention can have effects beyond its intended domain.
For a Persona, that motivates a specific experiment: if we train for familiar expression, do we also change its willingness to disagree, its handling of uncertainty, or the priorities reflected in its advice?
Testing where personalisation belongs
Rimembra’s initial design keeps the base model frozen while evaluating the capture, representation, and use of personal evidence. This establishes a baseline against which future adaptation can be assessed.
Our proposed research would then separate the contribution of better evidence organisation from the contribution of changing weights.

Within each evidence condition, comparing the base and adapted models would help estimate the additional contribution of training. Comparing evidence conditions would help establish whether the model benefits from a richer representation. The combined comparison could reveal whether adaptation mainly helps the model use that representation.
The most revealing tests would examine unfamiliar situations and the conditions that change a person’s judgment.
Suppose someone supports helping a relative financially when their own household’s essential expenses remain protected. A useful evaluation would vary that condition while keeping the rest of the scenario similar.

A Persona that always recommends helping might sound warm and consistent. A more faithful one would need to preserve the person’s qualification.
We would assess those answers along separate dimensions:
| Dimension | What the evaluation asks |
|---|---|
| Personal facts | Are claims supported by the person’s evidence? |
| Judgment | Does the answer reflect their priorities and relevant conditions? |
| Expression | Is the communication recognisable without exaggeration? |
| Uncertainty | Does the system acknowledge what it cannot establish? |
| Preference | Does the reviewer like the answer, independently of its fidelity? |
The person’s reference answers should be recorded before they see model outputs. Final testing should use scenarios withheld from training and development. Blinded comparisons and repeated runs would help distinguish systematic changes from ordinary variation.
These are proposed experiments. Their value will depend on what they demonstrate—including where adaptation provides no benefit or introduces regressions.
Preserve the system, then test the representation
A meaningful preservation strategy needs a versioned record of the model and any adapters, the personal-evidence snapshot, retrieval and context construction, instructions, and inference environment.

That record makes changes traceable. It supports investigation and recovery when behaviour shifts. It does not, by itself, establish that the resulting Persona is faithful.
The scientific challenge is to connect a system’s responses to independent evidence about the individual: their judgments, qualifications, contradictions, and uncertainty.
Owning the weights is an important form of control. Fine-tuning is a promising experimental tool. Neither substitutes for testing the complete system.
The standard is whether the person remains recognisable in the judgments the system expresses—not simply in the language it uses.
References
- Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F. and Liang, P. Lost in the Middle: How Language Models Use Long Contexts. Transactions of the ACL, 2024. arXiv:2307.03172 ↩︎
- Hu, E. J. et al. LoRA: Low-Rank Adaptation of Large Language Models. ICLR 2022. arXiv:2106.09685 ↩︎
- Abdulhai, M. et al. Consistently Simulating Human Personas with Multi-Turn Reinforcement Learning. NeurIPS 2025. arXiv:2511.00222 ↩︎
- Betley, J., Warncke, N., Sztyber-Betley, A. et al. Training large language models on narrow tasks can lead to broad misalignment. Nature, 2026. Preprint: arXiv:2502.17424 ↩︎