zhangtongshe.com
Memory is not the end goal; evolution is. Metis, the world's first memory-native foundation model, makes AI understand you better the more you use it.
If an AI works with you every day but has to reacquaint itself with you daily, can it truly be considered intelligent?
You tell it your habits, preferences, and working style, and it responds fluently in the moment. But switch to a different window or return after a while, and it acts like a complete stranger.
In an enterprise setting, this "amnesia" becomes even more problematic.
The research report analysis methods, equipment maintenance experience, and customer service strategies that employees have accumulated over years can be invoked by AI in a single task, but they are hard to continuously consolidate through use. Every time a new task is initiated, the system still has to re-search for information, read documents, and understand the context.
Over the past few years, the industry has continuously extended context windows and integrated RAG (Retrieval-Augmented Generation) knowledge bases into large models. Models can process more information in a single pass and find answers from external materials. However, the fundamental issue remains unresolved: after a conversation and task end, can AI retain truly valuable experience?
Memory Tensor (Shanghai) Technology Co., Ltd. (referred to as "MemTensor") has focused its breakthrough on "long-term memory".
Founder and CEO Xiong Feiyu earned his bachelor's degree from Huazhong University of Science and Technology and later obtained his Ph.D. from Drexel University in the United States. He has been deeply engaged in the field of large models and algorithms for a long time, and previously led the core algorithm team building at Alibaba. In 2023, he took the lead in establishing the Large Model Center at the Shanghai Algorithm Innovation Institute. Starting from the fundamental principles of next-generation model architectures, he proactively set long-term memory as a core proposition for theoretical research and model validation.
Subsequently, he led the team to found MemTensor and gradually built a three-tier technical system comprising Memory³, MemOS, and Metis. Among them, MemOS is positioned as the industry's first memory operating system framework for large models; Metis is the world's first memory foundation model prototype.
MemTensor attempts to open the "door of memory" for AI, transitioning from "one-time intelligence" to "long-term intelligence" capable of accumulating experience and continuous learning. In Xiong Feiyu's view, the ultimate goal of memory is not merely to "remember," but to achieve "self-evolution"—enabling AI to learn while using and continuously evolve in the process of serving businesses and collaborating with humans, understanding the company better the more it is used.
Today, this exploration has stepped out of the laboratory.
The MemOS open-source project has surpassed 10,000 GitHub stars, and its cloud service monthly API calls exceed 50 million; in July this year, the memory-native foundation model Metis open-sourced three versions with 4B, 9B, and 27B parameters, achieving a total of over 10,000 downloads within 30 days of its launch on Hugging Face. MemTensor has served over 500 enterprise clients, with signed contract values reaching tens of millions of CNY.
In July and September 2026, within less than two months, the company completed two consecutive rounds of financing: the Pre-A round was co-invested by Huawei Hubble, Honor Strategic Investment, SenseTime Guoxiang, Shenzhen Capital Group, and Harmony Capital; the Pre-A+ round was led by Oriental Fortune Capital, with follow-on investment from Futeng Capital. Adding Biren Technology, which joined in the angel round, the shareholder list includes industrial ecosystem players like Huawei, Honor, and SenseTime, domestic GPU technology peers like Biren, as well as top-tier institutions such as Shenzhen Capital Group, Oriental Fortune Capital, and Harmony Capital. Developers, customers, and capital are beginning to vote for long-term memory. But as model companies, Agent enterprises, and terminal manufacturers are all building Memory capabilities, what qualifies an independent startup to define the memory capabilities of next-generation AI?
With this question in mind, Zhangtongshe conducted an interview with Xiong Feiyu. The following is the organized interview content.
PART 01: Starting from the Origin of Model Architecture, He is Convinced that "Memory" is the Next Stage
Zhangtongshe: How did you focus on the direction of "AI memory"?
Xiong Feiyu: This direction is what we saw when dissecting the underlying architecture of large models.
When you deeply dissect the current large model architecture, you will find that Transformer is essentially a stateless sequence computation engine—with every inference, it starts "from scratch." No matter how long the context window is, once the conversation is closed, the state is cleared to zero. This is not a problem that can be solved by engineering optimization, but a fundamental deficiency at the architectural level.
Around 2023, the entire industry was competing in parameter scale and context length, but our judgment is: what truly determines the upper limit of next-generation model capabilities is not doubling the parameters, but whether the model can possess continuous, evolvable memory. Without memory, a model will forever be a "one-time" system—no matter how smart, it cannot accumulate experience through use or form true long-term intelligence.
We judge that the core driving force of this wave of AGI comes from the understanding of the model itself. When you thoroughly figure out the architecture, you will realize that memory is not an add-on, but a capability that next-generation foundation models must natively possess from the very beginning of their design.
Why? Because the context window is stateless—once the window is closed, memory is cleared to zero, which cannot solve personalization and self-evolution. What memory needs to solve is the issue of whether continuous learning can occur across sessions and time dimensions.
Therefore, the first thing we did was not to write a memory plugin, but to first answer a fundamental question: what should AI memory essentially be? How does information enter, get organized, get retrieved, and get forgotten—this is the starting point of the later Memory³ hierarchical memory theory.
In our view, memory is not about attaching a small function to a large model, but a core component of the next-generation foundation architecture. If AI wants to serve individuals or enterprises in the long term, this capability must eventually grow from within the model itself.
Zhangtongshe: But now context windows are getting longer and longer, and enterprises can also integrate RAG. Why build a separate memory system?
Xiong Feiyu: Because they do not solve the same problem.
The context window can be seen as a temporary workbench. The larger the desk, the more materials can be spread out at once. But after the task is over, these materials will not automatically turn into long-term experience.
RAG is more like a data repository. The system can go in and search for whatever information is needed. But it is not responsible for determining which content is more important, which is outdated, or when it should be updated.
Long-term memory cares about whether AI can retain truly valuable things after a conversation and task end, and accurately invoke them when needed next time.
Simply put, long context and RAG mainly solve "what is seen," while MemOS solves "what to remember, how to update, and when to forget."
PART 02: From MemOS to Metis, Memory Moves Inside the Model
Zhangtongshe: So MemTensor is a company that builds AI foundation models?
Xiong Feiyu: Yes. Our entry point is long-term memory, but what we ultimately want to explore is the next-generation AI foundation architecture.
For AI to form long-term intelligence, it requires a complete memory system: how information enters, where it is stored, when it is invoked, and how it is updated and forgotten. Centered around these issues, we have built a three-tier system comprising Memory³, MemOS, and Metis.
Zhangtongshe: Can you explain this in a more accessible way?
Xiong Feiyu: Memory³ is like a blueprint, first defining what AI memory is, and how it is generated, organized, retrieved, and updated.
MemOS is like an operating system, responsible for managing AI's memory. It uniformly schedules which information is worth saving, where to store it, when to invoke it, and what to do when it expires.
Metis takes a step further, enabling the model itself to possess memory. We call it a memory-native foundation model, which is also the world's first memory foundation model prototype.
The entire path is actually very clear: Memory³ first defines the problem, MemOS validates it in real-world scenarios, and Metis then enters the model itself, going deeper into the underlying layers.
Zhangtongshe: MemOS already has an open-source ecosystem and real-world invocations. Why still invest in Metis?
Xiong Feiyu: Because this approach has a ceiling.
Current large models themselves lack long-term memory capabilities. MemOS is equivalent to adding a memory management system outside the model to first make up for the missing capabilities. But memory and reasoning are still separated, and it will ultimately be limited by the capabilities of the underlying model.
I often say within the company that many external systems will eventually and gradually return inside the model. Memory is the same. True long-term memory cannot rely on external plugins forever; the model must have memory from the very beginning of its design. This is the problem Metis aims to solve.
Zhangtongshe: What is the difference between a model having its own memory and the current situation?
Xiong Feiyu: Current large models are somewhat like people who suffer from amnesia every time they wake up. They can read a large amount of materials and complete immediate tasks, but these experiences are hard to truly retain.
Metis designs a persistent memory state inside the model. Historical information, after being compressed, can be stored in the internal memory matrix. During inference, the model can proactively find relevant memories and also write new experiences.
It also distinguishes between long-term memory and working memory. Some information is only used for processing immediate tasks, while some experiences need to be preserved for the long term.
More importantly, Metis can update memory during inference without needing to be retrained every time, giving the model the opportunity to truly "learn while using."—By continuously writing, updating, and consolidating memory during ongoing interactions with users and enterprises, it allows "continuous evolution" to truly happen.
Currently, Metis has open-sourced three versions with parameter scales of 4B, 9B, and 27B, with the code, model weights, and training evaluation framework fully released. Within 30 days of launch, downloads on Hugging Face exceeded 10,000. What we are validating is that as the model scale continues to expand, the Scaling Law of this native memory architecture continues to hold true.
This is a question that no one in the entire industry has systematically answered yet.
PART 03: Behind Over 500 Customers, An AI Memory Business
Zhangtongshe: There are already over 500 enterprise customers. What exactly are they using long-term memory for?
Xiong Feiyu: The industries vary, but the core needs are very similar—retaining scattered experiences for continuous use by AI. At the enterprise level, this is organizational-level self-evolution: precipitating knowledge scattered in personal experiences into organizational memory, so the AI understands the company better the more it is used.
In the financial sector, we helped China Merchants Securities build a financial analysis agent, precipitating research report sorting methods, market review frameworks, and customer service strategies into organizational memory, transforming personal experience into institutional assets.
In the industrial sector, we cooperated with China Haisum Engineering to incorporate drawings, equipment time-series logs, and operation and maintenance experience into an industrial memory system. When equipment problems occur, the system can proactively match relevant experiences and decision-making references, eliminating the need for staff to re-search for information or consult experts.
In the edge and smart hardware sectors, we collaborate with Honor, Lenovo, and Haier Xiaoyou to enable cross-device memory synchronization capabilities for smartphones and smart home devices, while also addressing data boundary issues between the edge and the cloud. In the direction of gaming and emotional companionship, memory is becoming the core of product differentiation—the relationship between users and AI characters can accumulate over time, rather than starting from scratch with every interaction.
Zhangtongshe: The open-source project has popularity, but commercial revenue is another matter. How do you bridge the two?
Xiong Feiyu: Our commercialization is divided into three layers: open-source to build a developer ecosystem, cloud services to meet paid demands in production environments, and private deployment to serve data-sensitive customers in finance, industry, etc.
We will not rely on a large number of customized projects, but rather focus on doing a good job with standardized products, and then let ecosystem partners complete the "last mile" of industry adaptation. Only when products can be replicated can this truly scale up.
Zhangtongshe: Now model companies, Agent companies, and terminal manufacturers are all building their own Memory capabilities. Is an independent third party still needed?
Xiong Feiyu: Precisely because everyone is doing it, a neutral memory layer is even more needed.
First is memory sovereignty. If data and experience accumulated by enterprises over the long term are locked in a certain platform, changing the model is equivalent to moving "home" all over again. We believe that memory should belong to the customers.
Second is cross-platform capability. In reality, enterprises often use multiple models simultaneously, and users also use AI on different devices. Memory needs to flow across models and terminals, and cannot be dependent on a single platform.
Another point is technical depth. Many companies can build Memory at the application layer, but to answer "why the model itself cannot remember," one must delve into the underlying architecture.
We went from the Memory³ theory to the MemOS system, and then into the Metis model, following a full-stack route. The fact that enterprises such as Huawei, Honor, and Lenovo have become our ecosystem partners also shows that the market needs such a neutral memory capability.
PART 04: Why Shanghai?
Zhangtongshe: From technical validation to entrepreneurial landing, what key support has Shanghai provided for you?
Xiong Feiyu: The roots of our technology are right here in Shanghai.
In July 2023, the Shanghai Algorithm Innovation Institute took the lead in establishing the Large Model Center. The Memory³ hierarchical memory theory and the first-generation memory foundation model were validated for nearly a year before moving towards independent development. The team also gradually took shape during this process.
Shanghai also provided a complete transformation chain. The institute provided early-stage computing power and research teams, Shanghai Jiao Tong University provided theoretical support, Futeng Capital under Shanghai State-owned Capital Investment participated in our angel round investment, and enterprises in finance, industry, communications, and smart hardware provided real-world scenarios.
For startups in underlying technologies, this is crucial. Because you cannot just write papers, nor can you just tell business stories. Theory, models, products, capital, and customers must be linked one after another. In Shanghai, this chain can be connected.
Zhangtongshe: In the next five years, where do you hope MemTensor will be positioned?
Xiong Feiyu: I hope MemTensor becomes the company that defines the next generation of foundation models through memory.
Technically, Metis aims to become the reference implementation of memory-native foundation models. We propose a new paradigm question and also provide our own prototype answer as a Chinese team, standing on the same starting line as overseas counterparts. Long-term memory is not an external plugin, but a core capability that must grow into the model from the very beginning of its design—this is the true opportunity at the foundation model level.
We also hope that MemTensor can become a calling card for Shanghai in the exploration of next-generation foundation models. Shanghai has large model companies and computing centers; it should also have teams delving deep into the underlying layers of models, exploring new architectural paradigms starting from fundamental principles.
AGI will not arrive suddenly. It will grow layer by layer.
Reasoning capability is one layer, and memory capability is another. What MemTensor aims to do is to lead AI from "one-time intelligence" to "long-term intelligence." And above memory lies continuous learning and self-evolution—this is precisely the most essential difference between next-generation foundation models and today's.
This is something we are doing in Shanghai, starting from Shanghai.
Text | Nan Yi Layout | Guo Beibei