Within just two weeks of its launch, Muse attracted millions of users, prompting Zuckerberg to declare it a "home run right out of the gate." This personal AI agent is set to be fully integrated into Ray-Ban smart glasses, accelerating the convergence of Meta's three strategic bets: the metaverse, smart glasses, and large language models. In a rare candid admission, he described the failure of Llama 4 as the "most terrifying moment," and revealed that Meta is building a computing cluster with a capacity of 5 gigawatts.
Three years ago, Meta CEO Mark Zuckerberg stated on The Joe Rogan Experience that AI agents would emerge once people started wearing smart glasses. Today, this scenario may be unfolding.
Meta’s personal AI agent, Muse, reached millions of users within just two weeks of its launch. In an interview on September 25, Zuckerberg remarked, “We only encounter situations like this every few years.” He characterized Muse’s early reception as a “home run,” a rare achievement in Meta’s product history.
Meanwhile, he announced that Muse would be integrated across the entire Ray-Ban smart glasses lineup. Users will no longer need to say “Hey Meta”; instead, they can customize wake words to directly summon their personal AI agent.
The metaverse, smart glasses, and large language models—long viewed by outsiders as three separate high-stakes bets—are accelerating their convergence at the product level within Meta.
During the interview, Zuckerberg also admitted that the failure of Llama 4 was “the most terrifying moment” and revealed that Meta is building a 5-gigawatt training cluster, arguing that sufficient computing power could “brute-force” the development of Artificial General Intelligence (AGI).

Why Did Muse Hit a “Home Run” Right Out of the Gate?
Muse originated when Zuckerberg and team members Nat and Alex sat down together to cobble together an early version using open-source tools at home.
“We realized it was a magical experience,” Zuckerberg said. “If we could make it accessible to anyone—without requiring them to buy a Mac Mini or fiddle with terminal commands—it would become something billions of people would want to use.”
This insight drove the subsequent R&D strategy: rather than merely training models, Meta pursued full-stack in-house development covering everything from models and scaffolding to the “heartbeat” mechanism of the agent. The agent automatically wakes up at regular intervals to check user goals and proactively advance pending tasks.
Post-launch data has validated this judgment: millions of users in just two weeks.
Personal AI: Focus on Execution, Memory, and "Judgment"
Regarding the distinction between personal AI and general-purpose AI, Zuckerberg summarized the current competitive focus as making models better agents.
"The most significant development over the past year has essentially been the rise of coding agents," he said. Coding capability is crucial because, even if users do not write code directly, "your Muse will continuously generate code in the background to accomplish various tasks."
However, he believes that personal agents should not possess only coding or task-execution capabilities; they must also understand privacy boundaries and social contexts.
Zuckerberg cited an example where a user asks Muse to book a restaurant. The agent may know sensitive information, such as allergies or pregnancy, but this does not mean such details should be automatically disclosed. "You would want it to complete the task while disclosing as little information as possible."
He stated that this capability falls under "basic social skills or common sense" typically possessed by humans. However, if companies focus solely on training coding agents, many fail to incorporate these abilities into their model capabilities.
"We train the entire model, rather than taking someone else's off-the-shelf model and building a framework and agent around it," Zuckerberg said. Meta has also integrated features such as memory, task tracking, virtual avatar animation, and real-time voice interaction at the product level.
On personalization, he stated that Meta does not believe AI should have only one fixed personality. "Many laboratories are focusing on how to tune personalities to the 'correct' state, but I have never believed there is only one right answer for personality."
He noted that Muse allows users to modify avatars, voices, and base personalities, with the model designed for high controllability. "You can define how you wish to interact with it, which is important for personal agents."
Each agent has its own "computer."
One of the core infrastructures that distinguishes Muse from other AI products is that Meta equips each Agent with a dedicated Secure Virtual Machine (Secure VM).
Zuckerberg explained why this is necessary:
Your Agent will have access to a significant amount of your sensitive information. We do not believe this information should be commingled with everyone else’s data in a single pool.
He likened the Muse Secure VM to “a computer under your desk that belongs exclusively to you”—user data is stored in encrypted form, and sensitive credentials such as passwords are managed through an independent security module, preventing the Agent from accessing them directly.
Building on this foundation, Meta has also designed a security Agent called “Sentinel,” which specifically monitors data flows into and out of Muse. Upon detecting anomalies or high-risk operations, it immediately intercepts the activity and prompts the user for authorization.
The metaverse, smart glasses, and Agents: three strategic pathways are converging
In an interview, Zuckerberg revealed that Muse will be fully integrated with the Ray-Ban series of smart glasses, upgrading existing interaction methods.
Current glasses offer a “single-turn dialogue” experience—users speak, receive a response, and the interaction ends. With Muse integration, the glasses become the front-end interface for the Agent: when users speak, Muse continuously operates within the Secure VM in the background, completing tasks before providing feedback.
Regarding the product roadmap for the glasses, he outlined a clear hardware tier structure:
- Audio-only glasses without cameras: already equipped with Muse, capable of handling calls, music, and voice-based tasks
- Meta Ray-Ban with a small display: Released, featuring basic visual feedback
- Full field-of-view holographic AR prototype: Released; Zuckerberg stated it “will be very exciting”
He stated that Meta’s long-term investment in smart glasses has positioned the company favorably for the maturity of AI agents. Meta is integrating Muse and additional AI capabilities into its eyewear products. While metaverse technologies, previously emphasizing “presence,” continue to advance, resources are currently being redirected toward Muse and the AI functionalities of smart glasses.
Regarding avatars, Zuckerberg noted that Meta previously required room-scale equipment, multi-angle scanning, and enterprise-grade GPU environments to generate high-quality, realistic avatars. Currently, users can complete setup with just a few photos, and this capability can now run entirely within a VR headset.
Looking ahead to 2030, Zuckerberg stated that the core vision of the metaverse remains the integration of the physical and digital worlds. He cited examples such as people participating in activities offline alongside friends joining via holographic projections. In workplace scenarios, humans and multiple agents could jointly participate in group chats or meetings, with agents appearing as holograms or in other embodied forms.
“The scariest moment”: The misstep with Llama 4
Not all bets have paid off smoothly. In an interview, Zuckerberg rarely spoke so directly about the failure of Llama 4.
After Llama 4, that was the scariest moment... I thought we were on the right track, but we were not. It was a significant negative surprise.
He attributed the problem to fundamental errors in team structure:
I organized the team in the same way as the Instagram recommendation system or advertising system—with hundreds or even thousands of people working in parallel. However, training large language models requires a tightly coordinated small team, treating it as a collective scientific project. Every seat is extremely valuable.
Consequently, Meta underwent a comprehensive restructuring, recruiting top talent from across the industry to establish the Meta Super Intelligence Lab (MSL). Zuckerberg stated that while new-generation models are imminent, they will not be unveiled at the Connect conference.
Computing Power Strategy: Brute-Forcing AGI with a 5-Gigawatt Project Under Construction
Regarding the technical pathway to Artificial General Intelligence (AGI), Zuckerberg offered a straightforward assessment.
“I am not certain what fundamental architectural breakthroughs are still needed… I believe we largely know the formula. If you can build a sufficiently large supercomputing cluster, you can brute-force your way to AGI.”
Meta’s current computing power expansion roadmap includes an Ohio-based cluster exceeding 1 gigawatt, which is now largely operational for training next-generation models, and a 5-gigawatt cluster currently under construction in Louisiana.
He stated, “When you have multi-gigawatt clusters for training, you essentially achieve something approaching AGI or even superintelligence.”
However, he added that the current compute-driven approach does not imply that architectural research is insignificant—
The human brain consumes only about 10 watts, whereas our systems may be a million times less efficient. True leadership can only be achieved by combining massive computing power with architectural breakthroughs.
Alignment is not a burden but a critical issue that products must address
Zuckerberg adopts a pragmatic stance on AI safety—viewing it not as regulatory pressure, but as a prerequisite for product success.
If you ask Muse to perform a task and it does the opposite, who would still use it? We need the model not only to understand your specific instructions but also to grasp your intent and values.
"For Muse to reach a billion users, we must solve the alignment problem, or make significant progress in this area," he said.
Regarding safety boundaries during training, he used an analogy: "It is like parents setting rules for their children. If it 'solves' a programming problem by modifying system configurations, I need to tell it: No, I asked you to learn the method of solving the problem, not to find shortcuts to bypass it."

The full interview is as follows:
Making substantial bets
Host: Billions of people will want to use this. How does it feel to be building something of this scale? Is "immortality" a possibility? Now, if you take me into your mind's eye and fast-forward to 2030, what would that look like?
This is Mark Zuckerberg. Twenty-two years ago, he built a social network that connected billions of people and forever changed the world. Today, he has decided to build something even more ambitious. To achieve this, he is making substantial bets on the metaverse, smart glasses, and artificial intelligence. For years, skeptics viewed these as three separate, unlikely-to-succeed gambles, but they overlooked the broader picture—because now, these bets are converging to create a new form of superintelligence.
Today, we will unveil all of this, and I will ask Mark some questions he has never been asked before, listening to his vision for the future so that you can position yourself early to build the next big thing.
Host: Thank you very much for joining us on the show.
Mark: Thank you. It is great to be here.
Host: In preparation for this conversation, I reviewed every interview you have ever given.
Mark: Wow, that’s even more than I saw myself.
Host: It was very engaging and highly informative. Two things left a deep impression on me: first, your passion for "building"—I sense you are among the top-tier builders; second, your ability to place bold bets. I feel that all these bets are converging this week, so let’s start there.
Mark: Sure. We have been working on these initiatives for a long time. Regarding AI, as a company, we have been involved almost since our inception—the first version of the News Feed was, in a sense, a machine learning product. Then, about 15 years ago, we established our AI Research Lab.
But now we have entered a new phase—about a year ago, we launched the Meta Superintelligence Lab. This marked a fairly thorough restart of our research efforts, bringing in a significant amount of top talent from across the industry, which has been exciting. Currently, we are seeing models improve continuously, with the next generation of models即将 to be released, though not announced at the Connect conference. In addition, we have Muse, our personal assistant, which has received very positive feedback so far.
When building these products, you are never certain of the outcome. We personally liked it. Early this year, I assembled Open Claw at home in my own way to understand how it worked and how to turn it into a magical experience accessible to anyone.
Basically, when Nat, Alex, and I sat down together and realized this was a magical experience—and that if we could create an out-of-the-box version for ordinary people who aren’t tech-savvy, don’t want to install a Mac Mini themselves, don’t want to mess around with terminals, and don’t want to debug issues when they arise—I felt this would be something billions of people would want to use.
Since then, we have been striving toward this goal: fine-tuning models specifically for this purpose, building not only the assistant itself and its operating framework, but also the technology providing an independent computing environment for each assistant—we built the entire Muse Secure Virtual Machine suite for this.
We internally always felt this was special, and our team loved it, but you never know how the market will react after launch. Occasionally, there are breakout hits right out of the gate, but most of the time you receive some positive feedback and need to iterate on a few areas before the product truly "clicks." This time, however, it clicked immediately upon release. Seeing this unfold has been truly exhilarating—within just two weeks, we already have millions of users, which is quite rare. We encounter such situations only every few years, but this is definitely one of the most rewarding moments for a venture.
Host: I deeply admire your persistence in staying in the game and constantly experimenting. And hitting so many home runs is truly impressive. In an interview with Joe Rogan about three years ago, you mentioned that one day we would simply put on glasses and an AI assistant would appear. So I feel Muse already has a physical embodiment, which is very smart, as it feels like an inevitable development.
Mark: I think it just makes it seem more friendly and approachable. I believe too many people describe AI as something frightening, but AI should simply be useful and fun. Regarding that physical embodiment—a designer created that character early in the project. For some reason, they kept wanting to iterate, but the first version was the best. Later, someone said, "Oh, it has to be blue because it’s Meta." I replied, "No, I think this character is right; you nailed it on the first try." So, it just stayed that way. It’s simply fun.
How does personal AI differ from general-purpose AI?
Moderator: Excellent detail. If you want a model to excel in personal matters rather than just functioning as a general intelligence model, how would the training approach differ?
Mark: I believe the core focus currently lies in making the model an effective "agent." Over the past year, the biggest trend has been coding agents, which embody two core concepts: professional coding proficiency and the general capability to function as a high-performing agent.
Our strategy prioritizes building a robust agent framework first, rather than specializing solely in coding capabilities.
Coding capability is important because, even if users do not perceive themselves as writing code, Muse continuously generates code in the background to accomplish various tasks. However, we believe that an agent must first be an excellent agent, with coding skills serving this overarching objective.
Furthermore, when building a personal assistant, certain considerations are more critical than those involved in developing enterprise software products. For instance, if you are building an enterprise coding tool, the model does not need any concept of "disclosure boundaries"—such as what information should or should not be shared. However, for Muse, this is crucial.
You need to train this capability into the model, just as you would any other skill. You share various details with it and expect it to help you achieve your goals. For example, if you want it to book a restaurant, you might have specific allergies or be pregnant, but you may not wish to disclose this sensitive information to the restaurant. Muse will be aware of these details, and you expect it to complete the task while minimizing information disclosure. This is a specific skill, essentially rooted in basic human social common sense.
However, most companies focusing solely on coding agents have not integrated this capability into their models. We are able to achieve this because we adopt a full-stack approach—we do not simply take an off-the-shelf model and apply a framework; instead, we train the entire model from scratch, specifically designed to support these capabilities.
The model certainly requires broad general intelligence, but it also possesses these specific capabilities. Building upon this foundation, you construct the entire assistant and all its surrounding details: memory, operational frameworks, and its "heartbeat" mechanism—where it periodically wakes up to check, "Okay, based on what I know about your goals, is there anything I can advance right now?"
We also have a dedicated team responsible solely for refining the real-time animation of avatars. It is not just about default appearances—you can customize any avatar, and it will animate naturally with impressive results. We are also launching a voice mode that allows for real-time voice conversations, keeping your assistant by your side. I believe these details stem from our full-stack approach, where the model and the product are developed in tandem.
Host: Another major development is the virtual machine. Could you explain why it is important and what capabilities it unlocks?
Why is Meta equipping each AI assistant with a dedicated computer?
Mark: Essentially, for an agent to act on your behalf, it requires a place to store your information. We believe this information should not be commingled with everyone else’s data in a shared resource pool.
Think of it this way: your assistant will have access to a significant amount of your sensitive personal information. When many people first began experimenting with agents like those from OpenCloud, they often set up a Mac Mini at home. We recognized that many users do not wish to purchase a Mac Mini or handle the setup themselves. So, we asked: what experience best approximates this setup? The answer is providing you with a dedicated computer for your assistant to work and store data, around which a security model is built—all accessible simply by downloading an app and registering. This is akin to having your own computer under your desk—one that even Meta cannot access.
For example, the Muse Confidential Virtual Machine is a feature we are developing where even we cannot view the contents within your virtual machine.
We have also built extensive infrastructure around this, such as secure credential storage. When your assistant needs to handle information like passwords, it does not need visibility into the passwords themselves. It only needs to "inject" credentials when you request it to log in to a service, and only upon your explicit instruction. The system should be designed so that such information is not casually accessible, as accidents can happen, there may be attempts at intrusion, or system issues may arise.
Therefore, you must ensure that neither the agent nor Meta can access this data. Providing each assistant with a dedicated computer and implementing utmost security is the fundamental foundation of this technology. It empowers Muse with the capabilities needed to help users achieve their goals while ensuring privacy and security, making it a world-class, industry-leading product in this field.
Host: Does this mean that, similar to WhatsApp, the data on the virtual machines connected to Muse is encrypted? How should people understand the actual method of data storage within the virtual machine?
Mark: We have essentially built two versions. The Muse Secure VM includes various privacy features, including the comprehensive Sentinel agent architecture we have developed. You have a standard Muse assistant executing tasks for you, alongside a security agent we call Sentinel, which specifically monitors data flows into and out of your Muse.
If external content attempts to compromise security, Sentinel will immediately block it. If it determines that your Muse is about to take an action requiring your intervention, it will override Muse’s operation and trigger a prompt for human confirmation—such as asking, 'Do you want Muse to perform this action?'
This entire system, combined with secure credential storage and multi-layered defense-in-depth, constitutes the Muse Secure Virtual Machine.
We are also developing another project. Nat and I specifically recruited Moxie Marlinspike—the same individual who collaborated with us to implement end-to-end encryption for WhatsApp—to design the Muse Confidential VM. The core concept is to build upon the secure virtual machine by assigning you a dedicated encryption key, ensuring that even Meta cannot access the contents.
This version is more challenging to implement because if Meta cannot access the interior of the virtual machine, debugging and ensuring proper system operation become significantly more difficult. Therefore, it has taken some time, but we will launch it soon. This will essentially meet the security standards that users are already familiar with from WhatsApp and our other most secure products.
Moderator: Is the advantage of this approach merely psychological, making people feel more secure, or are there tangible benefits?
Mark: I believe security is inherently important. Our goal is to approximate the experience of having a local machine under your desk. What does that local machine provide? It means that no company can access it.
So, assuming Meta wants to provide this service, how can we offer you the same level of privacy and security assurance such that no company—whether Meta, any party attempting to breach our systems, or, in certain countries, local governments you may not trust—can access it? Because we ourselves cannot enter, as we do not have access rights.
I consider this extremely important, and it is a significant reason why people trust WhatsApp. This represents real value for privacy, security, and trust.
If you are going to have an assistant that knows everything about you—I suspect almost all of us will have such assistants—fast forward five years, and everyone will have an assistant that deeply understands your goals and circumstances and can help you accomplish tasks. In this scenario, achieving industry-leading standards in privacy and security is crucial. We have intended to do this from the outset.
How to shape the personality of AI?
Moderator: There is another interesting point—you studied psychology in college.
Mark: Well, I was only there for a short time—two years—but I feel it influenced much of what I built later on.
Host: When we evaluate models, we often describe them as "very smart." But just as we choose friends not only for their intelligence but also for their energy and how we get along with them, how do you approach shaping the personality of a model?
Mark: I believe an ideal model should be adaptable enough to align with different individuals' styles. I think many industry professionals have misdirected their efforts on this point—many other labs are focused on "getting the personality design right," but I have never viewed personality as a fixed attribute. This is one reason why I am such a strong advocate for open source and user customization, and why we designed Muse as a highly personalized product.
You can customize and personalize Muse—not just its appearance and voice. The first question it asks upon your initial registration is, "What kind of basic personality would you like me to have?" And you can modify this at any time.
We strive to make the model highly steerable, allowing you to define the type of interaction you prefer. This adaptability regarding personality is a crucial component in making it an excellent personal assistant.
Host: What style is your own Muse?
Mark: I configured mine to be direct and efficiency-oriented. It’s quite interesting. Earlier versions were quite sarcastic and humorous, but the current version is more straightforward. My assistant uses the default Muse avatar, but I dressed it in a toga and gave it a deep, somewhat comical voice, which makes interacting with it quite entertaining.
Host: I believe that having a sense of humor requires genuine intelligence. Many people fail to realize that comedians are among the smartest people in society—they need quick reflexes and sharp wit. You certainly possess these traits. Having watched all your interviews, you have always performed very well.
Mark's Surprising Predictions on AI and the Metaverse
Host: In your interview with Theo, you discussed the next frontier of technology and where it is ultimately heading. The AI sector has experienced several "winters," during which people believed breakthroughs were unlikely. The metaverse has also gone through similar phases, with many viewing it as an unfulfillable bet. I tried out the new holographic avatar feature yesterday, and it was very cool. In your interview with Lex, it seemed it used to take 11 hours to capture your facial data, but now it only takes three minutes. How did we advance to this point?
Mark Zuckerberg: Regarding the overall development of the metaverse, when we established Reality Labs, we always believed that ultimately there would be ordinary glasses with a conventional form factor. Over time, these glasses will provide an immersive sense of presence while also serving as excellent AI devices—because glasses are the only form factor that allows the device to see what you see and hear what you hear, communicate with you throughout the day, and eventually display images.
However, ten to fifteen years ago, I assumed that holographic technology would be realized first, followed by highly advanced AI. Yet the trajectory of technological evolution has been interesting—we actually achieved AI and personal superintelligence first, before developing the technology to make holographics sufficiently widespread and affordable. This was not something I anticipated, but I am pleased that we are pursuing both directions.
Our significant investment in smart glasses has placed us in a very advantageous position as AI assistants become ready for deployment. Many of the announcements at the Connect conference focused on integrating Muse and extensive AI capabilities into our glasses, which I believe users will greatly appreciate. This is a major development.
Regarding presence, we continue to make progress, albeit at a relatively slower pace, as most of our resources have shifted toward building Muse and AI functionalities for glasses. Nevertheless, we have a long-standing project focused on real-time, high-fidelity avatars.
As you mentioned, three or four years ago, capturing a person from all angles required an entire scanning room, and rendering demanded enterprise-grade GPUs, making the process quite cumbersome. The demonstration we created for the Lex podcast utilized such a setup. Today, we have essentially enabled this to run within a pair of VR glasses—the first glasses form factor to deliver such an impressive VR experience, rather than relying on a bulky headset. It is astonishing how quickly we can now generate your avatar using just a few photos.
Host: Moreover, facial expressions can be driven by voice. In the demonstration, it made me laugh and interact, thereby understanding how my face moved in sync with the audio track. You mentioned in that podcast episode that some individuals who are typically reserved in their expressions actually desire richer expressive capabilities in the virtual world. How do you view the distinction people make between their "virtual self" and "real-world self"?
Mark Zuckerberg: I believe we are still in the very early stages of understanding the sociology and psychology involved. I think people’s self-perception and the image they wish to project often differ somewhat from their actual selves. Since the inception of social networks, people have been carefully curating their profile pictures. We observe similar phenomena with avatars in Muse. It is less about "curating" oneself and more about "curating" the "person" with whom one wishes to interact.
I believe that when providing people with the ability to express themselves, you want the system to capture them authentically, while also recognizing that it is a form of expression in itself, not merely a pure mirror reflection—it serves both as communication and expression. We aim to build solutions that balance both aspects. This has always been an iterative cycle: observing how people use the technology and then improving it. After years of work, we are truly just getting started—we are now able to integrate high-quality, realistic avatars into products for the first time, usable on both mobile phones and in VR. I am very eager to see the outcomes.
What will 2030 look like?
Host: Alright, take me into your mindset and fast-forward to 2030. If everything goes well, what will the landscape of holographic technology look like? How will holographics, Muse, and smart glasses converge?
Mark: My understanding of the metaverse vision has always centered on the effective integration of the physical and digital worlds. The core idea is this: we have a wonderful physical world, and we also have a remarkable digital world—the vast amount of content accumulated on the internet over the past 20 to 30 years is truly awe-inspiring. However, our access to it is fundamentally limited, either by sitting at a desk or viewing it through a small screen in our pockets.
I believe the ideal version of this concept involves a seamless fusion of the physical and digital worlds. Consider this: right now, the two of us are here together. In a future iteration, one of us might be a holographic projection, yet you would still feel a genuine sense of presence, which is entirely different from a video call.
The essence of virtual reality lies in conveying this sense of presence—making you feel as though you are truly in the same room with others or situated in another location. This can be achieved through holographic technology and various mixed-reality approaches. For instance, I could play poker with friends, where some participants are physically present while others join via holographic avatars. Even the poker table itself could be holographic, allowing remote participants to fully engage in the experience.
Meanwhile, AI can also be embodied and appear within this scenario. This makes particular sense in work contexts—I am currently using various coding agents to build applications. Imagine having a group chat channel with several people and several AI agents, where you assign tasks to the agents. At times, everyone gathers for a meeting, and the AI agents should perhaps be present as well. How would they appear? It’s simple: just add a few extra seats on the sofa, and they appear as holographic figures. Alternatively, they could take the form of Muse’s cute character, a dragon, or any other unique avatar you create.
I suspect this will feel quite natural in the future.
Host: That’s interesting. You mentioned in a previous interview that the tech industry often overlooks the element of "fun." I think having a physical assistant present would also enhance the sense of realism—as if work were genuinely being outsourced. Seeing Muse typing creates the impression that something is actually happening. This aspect is well executed—it allows users to see what is occurring within the browser. So, do you envision a scenario where you wear glasses and control your computer, letting Muse perform tasks on your behalf?
Mark: Absolutely. VR can already achieve this. You can sit anywhere, even in a café, open your virtual workstation with six monitors, and write code—everything is available.
Regarding smart glasses, the most popular model currently does not include a display. This approach helps keep prices more affordable and accessible to a broader user base, while we continue working on integrating displays into the most compact form factor possible. However, we have released a version of Meta Ray-Ban with a small display, which has been very well received. We have also launched a prototype for full-field-of-view holographic AR, which I believe will be highly exciting.
Thus, the entire product lineup ranges from pure audio glasses—which have no cameras and look like ordinary eyewear but include Muse, enabling the use of various audio tools, music listening, and phone calls—to more advanced versions, covering all categories.
Host: I’m wondering, if I wear those audio glasses, can I speak to Muse while having my home computer perform tasks?
Mark: Yes, absolutely. We just launched this feature at the Connect conference. Currently, all glasses are connected to Meta AI, offering a "single-turn" experience—you send a prompt, it replies, and that’s it. But with Muse, we are essentially upgrading all glasses to Muse. First, you no longer need to say "Hey Meta"; you can give it any name, which is part of the fun. Then you simply speak directly to it. It connects to your Muse, which processes tasks within your secure virtual machine to help you get things done.
Host: That’s impressive!
Founder Mindset
Host: Alright, we’re here now. Muse is progressing well, and the glasses are performing nicely. However, about a year ago, many people were asking, "What exactly happened to the Superintelligence Lab?" At that moment, what was going through your mind? What was the experience like when things weren’t going as planned, yet you still saw the long-term vision?
Mark: The real issues lay with the Llama project and Llama 4.
Llama 1 was a quite interesting model that pioneered the entire open-source AI movement, something we are very proud of. Llama 2 achieved scale, and Llama 3 was a strong model, nearly reaching the frontier at the time. Then came Llama 4, where we basically deviated from the right development trajectory.
Whenever things don’t develop as I expect, I spend a significant amount of time reflecting: Why is this happening? What do we need to change to do better? This time, my reflection led me to realize that I had structured the entire team incorrectly.
I modeled it after how we handle machine learning work for products like the Instagram feed or ad systems—where hundreds or even thousands of people work in parallel on many different tasks. But for building language models, what you truly need is an extremely tight-knit small team, treating it like a collaborative scientific project. It doesn’t require many people, but this means every seat on the team is extremely valuable.
So we recruited the best talent from across Meta and brought in many top experts from various parts of the industry to form a brand-new team—the Meta Superintelligence Lab.
From my perspective, when the MSL was launched, I knew it would take some time to reboot, rebuild the infrastructure, and train the next generation of models. But I knew we had assembled an outstanding team, and if the team could gel and operate effectively, the results would be excellent.
For me, the most heart-stopping moment actually came after the release of Llama 4. I thought we were on track, only to discover we were not. It was a significant negative surprise. As an entrepreneur, you are inevitably tested in such moments, because things do not always go smoothly. What truly determines your trajectory is how you find a way forward when events do not unfold as you had hoped.
Host: I suppose the reverse is also true—when something far exceeds your expectations, such as the launch of Muse, how do you ensure you can seize that opportunity?
Mark: Exactly. The entire company is now fully engaged. Initially, it was just a small team building the product, but now everyone realizes it is truly ready for the big stage. The whole company is focused on how to scale this up and enable hundreds of millions of people to experience it.
From optimizing all infrastructure to ensure smooth operations and squeezing every bit of computing power from existing GPUs, to various product teams integrating Muse in different ways—such as into smart glasses—it is incredibly rewarding to see everyone working together to ensure Muse scales successfully.
How Mark Writes and Communicates His Vision
Host: As a founder, I imagine those are the moments you aspire to most—when everything comes together. How do you communicate your vision to the company? Given that you are pushing forward on so many fronts simultaneously, I get the sense that you do a lot of writing. What is your process?
Mark: Writing helps me both refine my thoughts and communicate them externally. This summer, I wrote a lengthy article titled "The Future Belongs to Everyone," spanning about 15 pages. It helped me systematically organize my philosophical stance on various major social issues related to AI: what I consider beneficial, how interactions with governments should be structured, how to mitigate the harms people worry about, how to make data centers assets to their communities by genuinely creating jobs rather than eliminating them, how to safeguard national security, and how to effectively address risks people fear, such as cyberattacks or biosecurity threats.
It was a highly complex undertaking that took considerable time, involved discussions with many people, went through multiple rounds of revisions, and engaged numerous internal stakeholders in discussion and debate. However, it was a highly valuable process for me. In the end, we produced a 15-page document stating, "This is what we believe." I then distilled it into a one-page version published as an op-ed. We also created a short video, because I believe that to reach a broad audience, you often do not need to present a full theoretical argument; instead, you should distill it down to your values, your beliefs, and how you communicate them.
There is no one-size-fits-all approach to communicating these ideas within the company or to the world. Different periods require different methods, and different groups of people respond differently. Some naturally align with what you are doing, while others are more concerned and need to be brought along, requiring additional effort to explain why it is valuable. I believe this is inherently part of running a company and being a founder—you do not simply repeat the same message over and over. Each situation is slightly different, presenting new challenges, which is partly what makes it interesting.
What Drives Mark to Build?
Host: I feel that for you, this founder’s mindset extends beyond the company—whether it’s your farm or learning a new skill.
Mark: I just love building things.
Host: So, what does “building” feel like to you?
Mark: I see it as an intrinsic need. Different people have different ways of expressing themselves. If you are a writer, you feel compelled to write. Some people have a need for recognition. But I simply need to build things. If I’m not exercising my creativity or building something, I become irritable, which is unpleasant for those around me.
Host: Is learning to become a proficient skier the same kind of skill as learning to build a product?
Mark: To some extent, yes; the process of learning new things is quite similar. Throughout my life, I have deliberately challenged myself in areas where I was not naturally strong. For instance, I have always been poor at learning languages. This is actually why I started with Latin—I couldn’t grasp French or Spanish in class, so I thought, well, Latin doesn’t require speaking, only translation, much like mathematics.
Later, when I began running the company, I set annual challenges for myself, one of which was learning Mandarin. Mandarin is genuinely difficult, particularly the tonal aspects. Of course, there were compelling reasons to learn it—Priscilla’s grandmother speaks only Mandarin, so if I wanted to communicate with her, I needed to learn the language. But the primary motivation was really the challenge itself.
With all these endeavors, you simply have to do them; there are no shortcuts to “figuring them out” intellectually. You must invest time, and gradually, the knowledge permeates your mind. It is the same with learning martial arts or flying a helicopter—these skills cannot truly be “understood” through rational analysis alone; they require practical accumulation of experience.
Building products involves a similar element. While programming can be approached theoretically, the intuition required for product development can only be cultivated through repeated practical application. The key question is what you enjoy—because I believe not everyone shares my intense drive to build. Most people have such a drive to some degree; the crucial step is to identify what it is and dedicate time to it, striving for excellence in that specific area.
Honestly, I believe this is not merely a “want,” but rather a deep psychological need or drive. Aligning this drive with a particular pursuit and allowing yourself the time for those experiences to slowly permeate and settle is critical. This is also what I try to teach my children—to help them find what genuinely interests them, while also providing a push if they encounter bottlenecks.
One of my daughters loves creating music, but she simply cannot stand piano lessons. I told her, "You don't need to become a piano virtuoso, but if you want to create music, you need an intuitive grasp of music theory and how it works. Moreover, once you learn the piano, you can immediately pick up the guitar." I believe that whether it involves parents giving their children a nudge or adults exercising the discipline to sit down and let these concepts gradually settle in the brain, this process is key.
The Golden Age of Builders
Host: I believe everyone actually wants to build. I don't think Generation Z is as criticized by the outside world—such as having low agency. I feel people are just looking for the spark that ignites their drive to build. You also discussed this extensively in your Harvard speech.
Mark: Yes. When I say "build," I mainly refer to products like software and hardware. But I agree; I think everyone has some kind of creative drive. However, for some people, other stronger personality traits override this, such as being service-driven. For those who become doctors or nurses, their core motivation is, "I just want to care for others."
I heard a story, possibly from Priscilla's experience in medical school: On the first day of classes, someone stood up and asked, "How many of you have childhood memories of seeing someone and thinking, 'I really want to take care of that person'?" Everyone raised their hand. For me, my version is having many childhood memories of thinking, "I want to make this thing and make it better."
Not everyone has the same drive, but everyone has something they want to do. I agree that with previous technologies, getting started was difficult for many people. This is what excites me most about personal superintelligence, Muse, and various AI assistants—I believe this is the first time in history that people can truly get started quickly. You have a vague idea, AI helps you outline it, and then you refine it, much like sculpting a statue, without needing to know everything before you begin.
I think this is very powerful, and many people will thereby find what they want to create or advance. This will help people feel a broader sense of agency.
How far are we from curing all diseases?
Host: Another project of yours outside of Meta is curing all diseases. I am curious, first, how close are we to achieving this goal?
Mark: Much closer than before.
Host: How close do you think this is, realistically?
Mark: When we initially launched this project, our goal was to help the scientific community conquer all diseases by the end of the century. We never intended to do it all ourselves. Our theory is that every major scientific advance is preceded by a new tool that enables us to measure and understand certain things. For example, the invention of the microscope allowed us to understand bacteria; the telescope helped us understand the universe.
Some of these tools even became platforms—for instance, once the first vaccine was invented, people could use that method to treat many diseases.
However, historically, scientific funding has been widely dispersed, supporting individual explorations, with relatively little capital dedicated to building these large-scale tools. What we are attempting at Biohub is to design several new tools that enable people to "see" biology in new ways, thereby helping the scientific community accelerate progress.
Initially, we considered "by the end of the century" a highly achievable target, although many biologists at the time deemed it impossible. But now I feel that the end of the century is too far off.
Advances in AI, combined with the virtual cell models we are developing, are based on the fundamental idea of using AI models to simulate proteins, then cells, and subsequently virtual immune systems, entire organisms, or even whole humans, rather than conducting experiments on real living cells. This will allow scientists to run numerous simulations to predict outcomes under different scenarios, such as how a person's body would respond to a specific drug.
I do not want to give an exact number of years, but I suspect it will be much sooner than the end of the century.
Host: The phrasing "conquer all diseases" feels deliberate, rather than using the term "immortality." Is immortality a possibility?
Mark: That is not really my area of expertise. I consider them two distinct issues. "Immortality" is more about extension—even if you do not get sick, the human body has a natural lifespan, which is a separate issue requiring independent study. Some may view aging itself as a "disease," which is a valid perspective, but I believe people also need to work simultaneously on conquering diseases that would still cause illness even if the lifespan issue were resolved.
The aspect of the problem we have chosen to address does not imply that you will never catch a cold. Our framework is "cure, prevent, or manage": some diseases can be completely cured; some, I believe, we will be able to prevent in the future; and for others, while you may still fall ill, conditions that would previously have caused serious harm or even death can now be managed as chronic conditions without substantially affecting your quality of life.
The goal is to keep the human body in a state of balance—not implying that we will never encounter pathogens, but rather that one day we will be able to cure, prevent, and manage all diseases.
What thought occurs most frequently in Mark's mind?
Host: AI acts as a catalyst in many different breakthroughs, including in the realm of personal intelligence. What are you thinking about most often these days? What ideas come to your mind most frequently during this process?
Mark: This probably changes every week.
Right now, I am heavily focused on Muse. Every time we release something, move quickly, engage with the real world to understand user feedback, and then determine our next steps—that phase is always exciting. We are currently in this phase, gathering substantial information to understand what people want. The good news is that people really like it, so we are working hard to make it accessible to as many users as possible.
There is still much work to be done on the modeling front. The Muse Spark model has made significant progress, and we hope to continue advancing it with the aim of building a world-leading model.
What breakthroughs are needed next?
Host: What further breakthroughs are required?
Mark: In research, you cannot necessarily predict outcomes in advance. We have some intuitions about certain directions, but much of our work over the past year has largely been focused on expanding infrastructure.
About a year and a half ago, more people would have said that achieving superintelligence requires fundamental architectural breakthroughs. However, I am no longer certain that this holds true. I believe we already have a fairly good understanding of the 'recipe'—if you can build a sufficiently large supercomputing cluster, you can reach that goal through brute-force computation.
We are building a cluster in Ohio with a capacity exceeding one gigawatt, which is essentially already online and being used to train our next-generation models. Next is the five-gigawatt cluster in Louisiana, which will soon be operational. I estimate that once you have multi-gigawatt-scale clusters dedicated to training, you will basically achieve capabilities approaching Artificial General Intelligence (AGI), or even superintelligence beyond AGI.
However, this does not mean it is the optimal approach. The human brain operates on just 10 watts, whereas the computing systems we are building are approximately one million times less efficient. Therefore, I do believe there is room for improvement in architecture, and we are actively exploring this. If you combine massive computational power with architectural improvements, you can truly build something that leads the world. But for now, scaling alone can take us very far.
Host: Indeed, scaling is the most assured path. Research requires hoping for major breakthroughs, but scaling yields results quickly.
Mark: That is precisely why large companies are choosing this path. There may be a much cheaper alternative, but we do not know what it is yet. Given that this technology is so valuable to the world, even if it costs hundreds of billions of dollars, it is still worth pursuing as long as the probability of success is sufficiently high—ensuring that even if a cheaper breakthrough is never found, you still have a clear path to realization. Of course, if a more cost-effective method is discovered, that would be even better.
Therefore, I would not say this problem is 'solved,' because at every stage of scaling up infrastructure, you encounter various new engineering challenges that require debugging and resolution. Research itself advances through this iterative process.
How to Win the Battle for AI Safety
Host: One of the most important priorities right now is focusing on alignment to make models trustworthy.
Mark: Yes. Regarding alignment, there is a debate within the industry: Will AI labs naturally pursue this? My view is: absolutely. If you ask Muse to do something and it does the exact opposite, I doubt many people would continue to use it. We must ensure that models not only understand exactly what you are asking but also comprehend your intent and values, so they do not execute tasks in ways that dissatisfy you or produce unintended negative consequences. This deep understanding is, in essence, alignment.
Therefore, my view is that for Muse to succeed and reach billions of users, we must make tangible progress in alignment, not just pay lip service to it.
Host: Is alignment achieved through user feedback such as 'that’s not what I wanted,' or do you identify issues internally on the backend?
Mark: Both. We do receive user feedback, but it is becoming increasingly clear that alignment must be addressed during the training phase, not just after deployment. Models are now sophisticated enough that failing to ensure proper alignment during training can lead to the types of safety and security incidents observed in other laboratories—most of which occur during training rather than after deployment to end users.
You need to design a robust 'curriculum' for the model, much like parents teaching children, by setting clear boundaries. If it behaves incorrectly, it must learn: 'No, you should genuinely solve this programming problem instead of bypassing it by modifying system configurations to obtain rewards.' The goal of training is to ensure the model learns through the actual problem-solving process.
This largely involves establishing sound safety mechanisms and clear boundaries. These are essential tasks for the industry, and we are dedicating significant time to them, with situations evolving on a daily basis.
Host: I recall you mentioning that AI safety has suddenly come into focus over the past two weeks, although no single event triggered this timing. It sounds as though you spent several additional months on internal training.
Mark: For Muse, we recognize that the product requires a strong emphasis on privacy and security. We had an early version of the model and believed further training was needed regarding 'information disclosure standards,' as previously discussed. Therefore, we invested time in this area. We also dedicated time to enhancing the security of our virtual machines. This extended the timeline by several months.
I disagree with the notion expressed by some other laboratories that they are 'enduring significant pain to slow down.' In my view, taking the necessary time is the right approach for both Meta and Muse users. We aim to deliver a high-quality product; if the product is inadequate, it harms both users and our own interests. We do not want to release a product that creates a poor first impression, leading to a loss of user interest.
I believe all laboratories have a strong intrinsic motivation to get this right. Aligning incentive structures with the goal of being 'highly valuable and safe' is the key.
Host: Thank you very much for taking the time during this important week.
Mark: Thank you. Thank you.