EPIC Helps Personal AI Remember What Matters
Abstract With the rapid emergence of personal AI agents based on Large Language Models (LLMs), implementing them on-device has become essential for privacy and responsiveness. To handle the inherently personal and context-dependent nature of real-world requests, such agents must ground their generation in device-resident personal context. However, under tight memory budgets, the core bottleneck is what to store so that retrieval remains aligned with the user. We propose EPIC (Efficient Preference-aligned Index Construction), which focuses on user preferences as a compact and stable form of personal context and integrates them throughout the RAG pipeline. EPIC selectively retains preference-relevant information from raw data and aligns retrieval toward preference-aligned contexts. Across four benchmarks covering conversations, debates, explanations, and recommendations, EPIC reduces indexing memory by 2,404 times, improves preference-following accuracy by 18.79 %p, and achieves 32.17 times lower retrieval latency over the best-performing baseline. In on-device experiments, EPIC maintains under 1 MB memory and achieves 5.21 to 29.35 ms/query latency across three platforms, while supporting streaming updates under preference drift. From previous conversations to individual preferences and constraints, personal AI agents work best when they can draw on a user's own context. However, keeping all of that information readily searchable can quickly strain the limited memory of smartphones and other personal devices. Led by Professor Taesik Gong of the Department of Computer Science and Engineering at UNIST, the team developed EPIC (Efficient Preference-aligned Index Construction), an on-device retrieval framework that selectively stores information relevant to a user's preferences. The approach sharply reduces memory use and retrieval time while improving how well AI responses reflect individual needs. EPIC is designed for retrieval-augmented generation (RAG), which allows a language model to draw on information beyond what it learned during training. On personal devices, however, indexing everything from documents and browsing activity to conversation histories can consume substantial memory. EPIC instead filters data before indexing, retaining what matters to the user while discarding what does not. For each item it retains, EPIC also creates an instruction describing when that information should be used. If a user with a seafood allergy asks for food recommendations while traveling, for example, the system can connect the request with that dietary constraint and retrieve suitable options rather than simply returning the most popular dishes. Across four benchmarks covering recommendations, debates, explanations, and conversations, EPIC reduced indexing memory by up to 2,404 times and retrieval latency by 32.17 times compared with the best-performing baseline. With Llama-3.1-8B-Instruct, preference-following accuracy improved by an average of 18.79 percentage points. Tests on a Jetson Orin Nano, MacBook Pro M4, and Galaxy Z Flip6 showed that EPIC could operate with less than 1 MB of retrieval memory, retrieving relevant information in 5.21 to 29.35 milliseconds per query. The system can also update its memory as new information arrives or user preferences change, without rebuilding the entire index. “EPIC is well suited to personalized AI that must adapt as new information arrives and user preferences change,” said Changmin Lee, the study's first author. “We plan to explore more flexible retrieval strategies, including ways to retain useful information that may be only indirectly related to a preference, while removing or compressing outdated memories.” “Rather than storing as much information as possible, this work focuses on storing what is valuable to the individual user,” said Professor Gong. “EPIC could help make personalized AI faster and more practical on everyday devices while keeping personal information on the device.” The research was presented at the 43rd International Conference on Machine Learning (ICML 2026) in Seoul. The code and data are available through the team's EPIC repository. The work was supported by the Ministry of Science and ICT and the Institute of Information & Communications Technology Planning & Evaluation (IITP), as well as the National Research Foundation of Korea. Journal Reference Changmin Lee, Jaemin Kim, and Taesik Gong, "From Volume to Value: Preference-Aligned Memory Construction for On-Device RAG," ICML '26 , (2026).
2026.10.06