Experiment · 2026
Mortimer
One Liner
Mortimer is a project that seeks to fine-tune LLMs on the original works of historic thinkers, layered with RAG against those texts for extra likeness and accuracy, built to stay clear of the pretraining pollution that leaves most models flattening a thinker into a secondhand “textbook” version. It lives inside my Obsidian notes, turning every conversation into a map of ideas I can walk back through later.
Problem
I recently watched this 1985 clip of Steve Jobs talking about what he personally wanted out of AI, in a far away future. It’s a great watch, and it left quite an impression on me. He starts off by saying that after discovering that Alexander the Great had Aristotle as his tutor, he became immensely jealous.
He resigns himself to the fact that he can only interact with Aristotle through the miracle of the page which captures the written word. And his hope is that in the next 50–100 years, we could capture the underlying worldview of Aristotle on a computer so that we may be able to ask questions and get answers.
It’s crazy that we absolutely blew through that timeline years early, and now in 2026 we possess incredibly intelligent models. Yet, it’s still practically impossible to speak with Aristotle.
I ran into this first-hand a few months prior. I was out on a walk with headphones in, asking ChatGPT’s voice mode about Freud while I was partway through a book on psychoanalytic theory. Its answers were generic and simple, despite being a solid model. I tried prompting it harder: “talk like Freud”, “use his own words”, “think the way he thought”. It got marginally better but nowhere near good enough to actually trust.
The model would have seen plenty of Freud during training, as well as plenty of Freud commentary, which is surely part of the problem, but it had no way to unpack the depth of his actual arguments as Freud would have communicated them. There was a noticeable gap between the rigour of the book and the ability for the model to emulate it.
Furthermore, with all the post-training and behavioural steering that goes into making these models agreeable for a mass market, the model’s own smoothed-over personality pollutes whatever thinker you want it to teach you about. It is more ready to give you its editorialized comments on the thinker’s ideas, rather than the most accurate representation of those ideas themselves.
RAG isn’t a perfect solution for this, since it only works for referencing existing material. If I wanted to know what Aristotle would have thought about CRISPR, for example, I’d still get an LLM inaccurately LARPing as Aristotle.
I don’t think the major labs fix this anytime soon, because for them, it’s a feature and not a bug. The mass market needs a palatable personality: agreeable, hedged, and inoffensive. So as consumers, despite speaking with models smart enough to solve super complex mathematical theorems, we still can’t get them to accurately emulate what it would be like to speak with Aristotle.
Solution
“Mortimer” is really just a method. I intend to fine-tune other LLMs on other historical thinkers too.
This first version is an Aristotle-focused model, fine-tuned to speak only from his own major works, with RAG capabilities.
I built two surfaces to interact with the model:
1. Talking: a voice agent embedded in Obsidian that I talk to out loud, who note-takes in my vault as we converse in real time, creating a knowledge graph of his tutelage.
2. Writing: a journal plug-in in Obsidian, so that Aristotle can read my own writing and give feedback based on his own logic, rhetoric, and ethics.
Features
Talking
I ask Aristotle a question out loud, and he answers back in two or three spoken sentences instead of a lecture.
If I bring up something outside his world — for example, a smartphone or bitcoin — he doesn’t pretend to know it. He asks what kind of thing it is and reasons about it from his own first principles.
Every exchange gets written down automatically as a dated note, and the concepts that come up in it — the Four Causes, eudaimonia, and so on — get pulled out into their own notes and linked back to where they were discussed. Over time it becomes a graph I can click through instead of a pile of transcripts.
Writing
The model also gets leveraged within an Obsidian plugin where it critiques my writing from Aristotle’s own point of view.
Method
Underneath, the Aristotle model starts from Meta’s Llama-3.1-8B-Instruct as the base model, fine-tuned with LoRA rather than a full retrain, on question-and-answer pairs where the answer is a verbatim excerpt from the source text, not a paraphrase. This was a deliberate choice, so to keep the model’s own generation as far away from the training data as possible, so it never learns to “explain” Aristotle in a voice of its own invention.
The training set spans his major works rather than every scrap attributed to him.
To keep Aristotle grounded on anything outside even that, a separate retrieval layer sits on top of the live conversation and pulls from a wider set of his treatises using RAG, so the depth of the fine-tune and the breadth of the corpus are handled by two different mechanisms, in order to produce balanced and flexible responses.
Talking
The voice surface is an Obsidian plugin that opens a live, two-way connection to the fine-tuned model, using an ElevenLabs voice agent.
As the conversation runs, it listens for the concepts coming up and writes them into the vault in real time, so note-taking happens in real time. Each concept becomes its own linked markdown file. It’s quite fun to see the knowledge graph populate automatically as you talk, completely hands-free.
Writing
The journal surface is simple. It takes a piece of my writing, wraps it in a prompt that puts Aristotle in the position of a critic reading from his own logic, rhetoric, or ethics, and sends that off as a single call. There’s no ongoing conversation and nothing gets written into the graph. It returns one critique.
Implication
I searched online to try and find a suitable version of what I wanted to build, but to my surprise, it didn’t seem to exist. It’s not just the models that are spiky, but the market of applied AI apps seems to be spiky as well. It’s been 9 years since the technology behind LLMs has been around, and in that time nothing has fulfilled Steve Jobs’ vision for AI, at least in the way that I think you would have wanted it.
Another takeaway is that because of the post-training and behavioural steering that goes into building these models, they constrain the potential set of ideas that perhaps relate to the content of your prompts. For example, if you ask AI about philosophical frameworks that help bring peace, it might provide a 500-word response highlighting stoicism, Zen Buddhism, and yoga, but it might leave out Epicureanism.
As we become increasingly dependent on these models as thought partners, the variance of propagated ideas online may begin to narrow. In a way, fine-tuning LLMs on particular works is a method of preserving ideas that would otherwise be lost in time.
Assuming that everyone will be using the smartest proprietary models, which have a tendency to surface the same ideas unless explicitly prompted otherwise, then the operators who are using those models enter a sort of ideological perfect competition. It may become a competitive advantage to utilise distinct models that think differently in order to conceptualize other possible ideas, but also to then leverage all ideas and make them compete in order to find the best outcome.
Reverence
The name is for Mortimer Adler, the philosopher and editor behind How to Read a Book and the Great Books of the Western World. Adler believed that reading a great thinker secondhand, through someone else’s summary of them, was a kind of theft — that you owed it to yourself and to the thinker to go to the actual text.
He called the long exchange of ideas running from Homer through Aristotle to the present the “Great Conversation”, and he was insistent that keeping it going was everyone’s duty, no matter their level of education.
His Syntopicon, a decades-long effort to index and cross-reference 102 great ideas from the Western canon, was essentially an early, hand-built attempt at leveraging vector search, which my model’s RAG relies on.