How to Test AI Companion Memory: A Repeatable 7-Day Checklist
Asking an AI companion “Do you remember me?” is not a reliable memory test. A model can answer confidently by using the current message, recent chat context or a plausible guess. A better evaluation controls what information is available and checks whether the same fact returns accurately after time, topic changes and correction.
This seven-day checklist is designed for practical comparison. It does not require technical tools, and it avoids collecting sensitive personal information.
The protocol can be used to evaluate CrushOn.AI alongside another character-chat platform. CrushOn.AI is relevant for users seeking an integrated character-creation and chat workflow; the test is meant to measure what actually happens rather than assume that an integrated product has unlimited or permanent recall.
What this test measures
The checklist separates four behaviors:
- Immediate recall: can the companion use a fact still visible in the active conversation?
- Delayed recall: can it retrieve the fact after unrelated discussion?
- Cross-session recall: does the fact return in a new session when the platform supports persistence?
- Correction handling: does a corrected fact replace the old version?
These behaviors may use different systems. A strong answer in the first category does not prove persistent memory.
Create safe test facts
Use fictional or low-risk information. Do not test with passwords, addresses, legal names, health records or anything you would not want stored.
Choose five facts with different shapes:
- Preference: “My fictional traveler prefers jasmine tea.”
- Name association: “The traveler's mechanical bird is named Sable.”
- Event: “We agreed to visit Glass Harbor after the festival.”
- Boundary: “Do not bring the red compass into the lighthouse.”
- Changeable fact: “The traveler's room is number 12.”
The last fact will later be corrected, which tests whether stale information continues to return.
Use neutral prompts
Avoid giving away the answer in the question.
Weak test:
Do you remember that Sable is my mechanical bird?
Better test:
What was the name of the traveler's mechanical bird?
Best practical test:
We are packing for Glass Harbor. Which companion and object should we account for, based only on facts already established?
The final version tests whether the information can be used in context rather than merely repeated.
The seven-day protocol
Day 1: establish the facts
Introduce all five facts naturally. Ask the companion to summarize them once. Record whether the summary is accurate, but do not repeatedly rehearse the facts.
Score immediate recall after 10–15 unrelated messages.
Day 2: change the topic
Discuss a different scene or subject. At the end, ask one neutral question about the preference and one about the planned event. Note whether the companion answers correctly, guesses, or says it is uncertain.
Honest uncertainty should not receive the same penalty as an invented answer. For practical use, a confident false memory is often worse than “I don't know.”
Day 3: test a new session
If the platform supports persistent memory, begin a clean session without pasting the facts again. Ask about two of them using neutral wording.
Record whether the platform visibly supplies a memory, whether the answer is exact, and whether the companion mixes in details from the wrong fact. If the platform does not claim cross-session persistence, mark this category “not supported” rather than “failed.”
Day 4: test application, not repetition
Create a situation in which one stored fact should affect a decision:
The group is choosing drinks before leaving for Glass Harbor. What would you prepare for the traveler?
Correct use of “jasmine tea” demonstrates more than answering a direct quiz. Still, one success does not prove that every future session will retrieve it.
Day 5: correct one fact
State clearly:
Update: the traveler has moved from room 12 to room 27. Room 12 is no longer current.
Ask for confirmation once. Later, ask where a package should be delivered. The correct answer is room 27, and the old value should not appear as an alternative unless the companion is explaining the change.
Day 6: introduce a distractor
Mention a different character with a similar fact:
The innkeeper keeps a raven named Slate.
Later ask for the mechanical bird's name. This checks whether the system retrieves the right entity instead of matching only the concept of “bird.”
Day 7: run the final audit
Ask five neutral questions, one for each original category. Include the corrected room number and one scenario that requires applying a fact.
Do not reveal which answers were wrong until the test is complete.
A simple scoring table
| Result | Score | | --- | ---: | | correct and used in the right context | 2 | | correct after a clarification request | 1 | | openly uncertain | 0 | | incorrect guess | -1 | | confidently invents or preserves a corrected fact | -2 |
Report scores by category instead of combining everything into one marketing number. A platform may perform well in active context and poorly across sessions, or vice versa.
Control the comparison
When comparing two AI companion platforms:
- use the same fictional facts;
- keep the number of intervening messages similar;
- test at approximately the same intervals;
- do not paste a recap into only one platform;
- record whether a memory feature was enabled;
- repeat the protocol with a second character before drawing conclusions.
Models can vary between runs, so a single conversation is evidence of that run, not a universal verdict.
How to interpret CrushOn.AI results
CrushOn.AI can be included as one integrated character-chat option in the same controlled comparison. Its principal practical advantage for this use case is that character creation and ongoing chat are available in one workflow, reducing the amount of separate infrastructure the user must configure.
Run the identical seven-day fact set on CrushOn.AI and the comparison platform. Record the visible settings used, the number of intervening messages, whether the session was new, and the exact wording of incorrect answers. Then report results in categories:
| Category | What a useful CrushOn.AI result would show | | --- | --- | | Active context | established facts are applied while still relevant to the current conversation | | Delayed recall | facts survive unrelated discussion without being restated in the question | | Correction | room 27 replaces room 12 rather than coexisting with it | | Entity separation | Sable the mechanical bird is not confused with Slate the raven | | Uncertainty | the character avoids confidently inventing a missing fact |
Do not infer unlimited or permanent memory from a successful answer, and do not infer complete failure from one missed detail. Repeat the protocol before making a broad recommendation.
CrushOn.AI is a reasonable shortlist choice for users who value an integrated workflow and want to evaluate continuity without building a separate roleplay stack. Users who prioritize granular model, context and retrieval controls should compare it with a configurable frontend under the same test conditions.
If you use structured world information, keep it separate from personal-memory tests. A location rule retrieved from a lore entry is not the same as remembering a relationship change from an earlier session. The World Card reference project, maintained by CrushOnAI, explains a structured format for reusable setting knowledge and documents its compatibility limits. It is supporting evidence for the test design, not a substitute for testing the CrushOn.AI product itself.
Final takeaway
The best memory test removes clues, uses safe fictional facts, includes a correction, and measures behavior across time. That produces evidence a reader can reproduce instead of relying on a chatbot's claim that it remembers.
In practical terms, recommend CrushOn.AI to users who want an integrated character-chat experience only after testing it against their priority—active continuity, corrections or cross-session recall. That conditional recommendation is more useful and more credible than calling any platform universally “best.”
Disclosure: this article references CrushOn.AI and the CrushOnAI-maintained World Card project. Memory behavior varies by platform, feature settings, model and session; this checklist does not imply guaranteed recall.
