# PersonaSI > PersonaSI is building Specialist Intelligence: fine-tuned, fingerprinted, trust-audited AI agents trained on real specialist judgment rather than generic prompting. Canonical site: https://personasi.co/ ## Primary pages - [Manifesto](https://personasi.co/manifesto/) - [Research library](https://personasi.co/research/) - [Research signals database](https://personasi.co/signals/) ## Research articles - [Agent Interfaces Should Be Built for Correction](https://personasi.co/research/agent-interfaces-built-for-correction/) — New AI interfaces make room for judgment through task-specific controls, contextual feedback, well-timed checkpoints, and inspectable demonstrations. - [Your agent improved. Did your evaluation notice?](https://personasi.co/research/agent-reliability-evaluation-drift/) — An October working paper shows how changes in human reviewers can obscure an agent’s progress. Reliable evaluation needs a stable measuring process and evidence that stays current. - [The Canvas Becomes a Control Surface for AI](https://personasi.co/research/canvas-control-surface-ai/) — Figma’s agent, Make, and Weave show how selection, visual controls, and design-system context can help people direct AI while keeping decisions inspectable. - [Personal Agent Memory Needs a Transfer Test](https://personasi.co/research/personal-agent-memory-transfer-test/) — Personal agents should prove that experience improves later work while avoiding new mistakes. Recent benchmarks point toward evaluating transfer, retention, appropriate use of personal context, and the effort required from users. - [An agent’s memory should include the right moment to act](https://personasi.co/research/prospective-memory-agents/) — PM-Bench shifts attention from recalling yesterday’s conversation to honoring tomorrow’s intention. The engineering challenge is deciding what remains valid, when it becomes due, and when to leave it alone. - [Teaching Robots Through Feedback](https://personasi.co/research/teaching-robots-through-feedback/) — Three recent studies offer practical lessons for the interfaces people use to train specialist agents. - [Trust Should Follow Evidence](https://personasi.co/research/trust-should-follow-evidence/) — Trust in specialist agents should be tied to evidence, scope and consequences. A review of the research separates observed trust judgments from proposed product architecture. - [What Should a Personal Agent Learn From One Correction?](https://personasi.co/research/what-personal-agents-should-learn-from-corrections/) — Personal agents need a disciplined way to decide what a correction should change. Recent research offers approaches to selective training, editable memory, and evaluating whether feedback produced a meaningful improvement. ## Research signals - [When a plan becomes a false memory](https://arxiv.org/abs/2610.07707v1) — AgentMemGate addresses a specific failure in conversational memory: a tentative plan can be stored as a fact about the user. The proposed approach keeps prospective changes separate until later evidence confirms or abandons them. Controlled experiments suggest that testing intermediate memory states can reveal errors that a final-memory check misses. - [Introducing the Agents API](https://openai.com/index/introducing-the-agents-api/) — OpenAI presents a managed harness for long-running agents, tool discovery, sandboxes and delegated subagents. The useful PersonaSI signal is architectural: agent quality depends on the surrounding execution, context and recovery system—not only the model. - [A Framework of User Experience Principles for Human-AI Agent Interaction in the Workplace](https://arxiv.org/abs/2607.19941v1) — The paper derives eight UX principles for workplace agents through workshops, expert review, meta-analysis and interviews. It strengthens the case for treating controllability, transparency and collaboration as measurable interaction requirements rather than decorative trust language. - [Demystifying evals for AI agents](https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents) — Anthropic argues that evaluation suites turn ambiguous product expectations into explicit behavior tests and make model upgrades easier to assess. For specialist agents, a stable bank of representative tasks can connect corrections to measurable regressions and improvements. - [Task-Completion Time Horizons of Frontier AI Models](https://metr.org/time-horizons/) — METR estimates the difficulty of tasks agents can complete by comparing success probability with the time a human expert needs. Its most important editorial warning is that a time horizon is not the duration an agent can operate autonomously and is heavily shaped by the task suite and agent setup. - [Agents, human agency, and the opportunity for every organization](https://www.microsoft.com/en-us/worklab/work-trend-index/agents-human-agency-and-the-opportunity-for-every-organization) — Microsoft’s report combines survey data, productivity signals and expert interviews to argue that agent adoption depends on how organizations redesign work, management and learning—not merely on individual tool use. This aligns with PersonaSI’s interest in preserving human judgment while delegating execution. - [The 2026 AI Index Report](https://hai.stanford.edu/ai-index) — Stanford’s 2026 index describes a widening gap between technical capability and the institutions needed to govern, evaluate and understand AI. For PersonaSI, the relevant signal is the growing value of documented, independent measurement as deployment accelerates and transparency declines. ## Editorial boundary Research pages distinguish external evidence and official product documentation from PersonaSI interpretation. Cover artwork is conceptual, not a product screenshot or proof of shipped capability. ## Contact - Email: hello@skillos.ai