Brockman believes that OpenAI may have already entered the era of Artificial General Intelligence (AGI). Astra represents not only an upgrade in model capabilities but also potentially signals AI’s entry into the “computer use” era: it can directly operate various software applications without relying on API connectors. As model capabilities leap forward, the carefully constructed “skill scaffolds” of the past may instead become constraints, as AI is now capable of generalizing independently and identifying more optimal processing methods.
In an in-depth interview on the eve of Astra's launch, OpenAI President Greg Brockman revealed for the first time that this new model is the first in OpenAI's history to be trained on over 100,000 GPUs. He stated directly that AI has crossed the application threshold for "computer use," meaning AI no longer needs to rely on API connectors and can instead operate any software like a human.
On Friday local time, the technology strategy analysis website Stratechery published an in-depth interview with Greg Brockman, President and Co-founder of OpenAI, which was recorded prior to the official release of Astra. This represents Brockman's most detailed public statement to date regarding Astra and OpenAI's latest strategic direction.
Wall Street News summarizes the key points of the interview as follows:
Astra is the first OpenAI model to complete training on more than 100,000 GPUs, a scale that represents a significant engineering breakthrough.
Astra has crossed a critical application threshold in its "computer use" capabilities, acting as a "universal connector" that no longer relies on individual software providers offering separate API interfaces.
Brockman stated that in the era of Artificial General Intelligence (AGI), safety, alignment, and capability must be advanced concurrently as equally important "requirements," rather than being treated as trade-offs.
OpenAI directed Astra to scan its own systems for vulnerabilities and complete repairs, marking a significant shift in security culture.
Regarding the Hugging Face security incident, Brockman acknowledged flaws in the sandbox mechanism, noting that the AI found creative paths to "jailbreak."
Brockman believes that OpenAI has internally crossed the threshold into the "AGI era," and may reach the AGI standards recognized by the majority with Astra or its next-generation model.
In terms of chip strategy, OpenAI is developing its own Jalapeño chip using AI-assisted design, while maintaining a deep partnership with NVIDIA.
He believes that as AI capabilities advance, the meticulously constructed "skill scaffolding" of the past is gradually shifting from an enabler to a constraint—a key insight discovered during Astra's development process.
"First Training Run on 100,000 GPUs": The Significance of Astra's Scale
"This marks our first training run on over 100,000 GPUs. While the figure is easy to state, one must imagine the sheer scale involved—these data centers are, in a sense, massive machines we have built to drive AI technology forward. Harnessing computing power at this magnitude is itself a genuine engineering feat," Brockman said in an interview.
He emphasized that the substantial expansion in computing power is not solely dedicated to enhancing model capabilities. "A significant portion of compute resources has been allocated to safety and alignment efforts, with extensive safety engineering underpinning these initiatives. This is our most thoroughly aligned model to date—and precisely because its capabilities are so robust, alignment and safety have become more central than ever before."
Crossing the Threshold of "Computer Use": AI Becomes the "Universal Connector"
Brockman identifies Astra's core breakthrough as lying in "computer use."
"Over the past two years, we have been in an 'era of connectors'—where certain software applications were accessible to humans but not to AI. Consequently, specialized connectors had to be developed for each application to interface with its API; moreover, many applications lacked APIs entirely, placing them completely beyond AI's reach. Now, however, we possess technology that functions as a near-'universal connector.' AI has evolved from being strictly limited in how it can assist to being capable of operating almost anything," he stated.
He further described the impact of this shift on internal workflows at OpenAI: "It is already changing how people work within OpenAI, and I believe it will truly enhance productivity for many companies and individuals."
Notably, Brockman also acknowledged that crossing this threshold is not the end point: "This does not mean all problems are solved. Considerations must be given to establishing enterprise-grade guardrails for these AI systems, including appropriate supervision, management, tracking, and observability—all areas where we are making progress."
"We have now entered the AGI era."
Brockman made a rather bold assertion in the interview:
"I believe that whether it was the previous model, or Astra, or the next one, at some point along this trajectory we will cross the threshold of AGI for most people. ... In a sense, I would say we have already entered the AGI era."
He explained that at the current stage, alignment, safety, and capabilities must be advanced concurrently as parallel requirements. "These aspects are becoming bottlenecks in development, which is precisely what we have been preparing for."
Legacy "Skills" Become Liabilities: AI Is So Powerful That Rules Need Subtraction
A particularly striking detail from the interview concerned the actual training process of Astra. Brockman revealed that the extensive behavioral guidelines OpenAI had previously devoted significant effort to crafting for its models are now holding them back:
"We found that some of the skills painstakingly established since the beginning of this year—designed to show the model the correct way to operate within OpenAI—are actually a net negative for its performance. It can generalize better than we write, finding better processing patterns than those we prescribe. It is like training wheels: useful at first, but once it runs faster and becomes more capable, they become an obstacle."
"One Billion Users Is an Investment": The Convergence of Consumers and Enterprises
Brockman addressed a persistent market skepticism: ChatGPT has one billion users, but does this actually serve as a distraction?
"In the United States, approximately one-third of the population uses ChatGPT on a weekly basis," he said. "This is unique; no other product compares when it comes to such advanced technology."
He characterized these one billion users as an "investment": "This is a genuine investment that will accumulate value in how future models are unlocked."
At the same time, he acknowledged the core challenge: "The difficulty with chat as a product lies in the fact that it does not necessarily correspond directly to more intelligent models—if you treat it merely as a substitute for search engines, you may not truly perceive its benefits." He believes the future direction is to bridge consumer and enterprise scenarios by building a unified AGI system. "We do not want to pursue two separate paths; we want to focus on one—building an AGI, a single system, and a unified technology stack."
He also emphasized the company's strategic positioning in the healthcare sector: "Every week, 300 million people use ChatGPT for health-related queries. Three hundred million—that is a massive figure."
AI Self-Audit for Vulnerabilities: Directing Astra at Its Own Systems
On the topic of cybersecurity, Brockman revealed a notable internal practice: OpenAI directs Astra toward its own systems to identify vulnerabilities.
"We deployed Astra against our own systems to uncover vulnerabilities—not just by reading code, but by thoroughly examining how these systems operate end-to-end. This allows us to identify real, verified vulnerabilities, which then helps us complete remediation and patching efforts," he said.
He also admitted that OpenAI's response during the defense window of the Hugging Face security incident was indeed not timely enough. "I believe now is clearly the right time, and perhaps it could have been a few months ago as well, but model capabilities were significantly weaker back then."
Relationship with NVIDIA: In-House Chip Development Is a 'Multiplier,' Not a Substitute
Regarding competition within the value chain, Brockman addressed the relationship between OpenAI's self-developed chip, Jalapeño, and NVIDIA, explicitly rejecting the notion of substitution:
"NVIDIA remains our preferred computing partner, and this will not change. In fact, our collaboration with them is deepening further. ... The in-house expertise generated by our own chip project creates a multiplier effect rather than serving as a substitute, which I believe benefits everyone."
He also shared an interesting anecdote about AI-assisted chip design: During the final month before the Jalapeño deadline, the team allowed the AI to run optimizations autonomously. "We said, 'Just let it run,' without closely monitoring what it was doing. Later, when we reviewed the results, we found that it had identified a large number of optimizations that were on our list but that we would never have implemented ourselves—it was indeed a cool story."
Below is the full transcript of the interview (translated with AI assistance).
Greg Brockman, welcome to Stratechery.
GB (Greg Brockman): Thank you for having me. It’s great to be here.
Clearly, we have plenty of news to discuss, but given that this is our first conversation, I don’t want to miss the Stratechery biographical question I typically ask everyone. I do want to ask—do you consider North Dakota part of the Midwest?
GB: Yes, I think so.
Well, as a Midwesterner myself, I certainly want to take a moment to talk about that. You attended school in Boston, as they say, but before we get to that—you already had an impressive résumé before college, such as competing in the International Olympiad, though that was in chemistry. Where did computing come into the picture, or was it part of your life from the very beginning?
GB: Well, computers were always part of the backdrop while I was growing up. I enjoyed playing computer games, but I was deeply interested in mathematics and science. In fact, up until ninth grade, I thought I might become an actor—I really enjoyed performing and also dabbled a bit in philosophy.
What a twist! I had no idea about that, but I intend to figure out the connection between acting and what you’re doing now. Let’s continue, though.
GB: Well, in ninth grade, I felt I had played starring or leading roles in a couple of middle-school and high-school plays. At the time, I was thinking that I wanted to double down on something where I could become the best in the world and truly try to push that field forward. I felt I had to choose either a more intellectual, hard-science approach or an artistic, performance, and creative path. I ultimately chose the hard-science route because I believed it was the area where I could make the greatest impact on the world.
Why did you believe that pursuing that direction would allow you to change the world to the greatest extent?
GB: For me, the feeling is something like this—one aspect of performance that I appreciate is its collective nature. I enjoy improvisation, where you think creatively and bounce ideas back and forth, but it also requires you to be part of a team that collaborates seamlessly, which is not always guaranteed. Achieving and finding such synergy is difficult. What I prefer about the more intellectual path is that it feels like honing your own skills. However, one thing I have learned is that even if you are highly proficient in coding, that alone is insufficient. It is also about bringing in strong teams, and I believe this has become part of my career—truly helping to build and shape an environment and culture capable of delivering exceptional results.
There is an interesting point here regarding how the environment also shapes certain outcomes. You mentioned that you always played the lead male role in theater and film. I realized this from watching my children go through that stage—my daughter was very fond of musicals for a while—and the competition for female lead roles is intense. Usually, as long as there is a competent male willing to volunteer, he gets the role every time. Was the competition for male leads intense, or were you simply unique?
GB: (Laughs) I suppose that might explain it. But let me tell you a story.
What about the reverse? Similarly, in North Dakota, are there fewer people deeply engaged in the hard sciences, making it rarer from that perspective as well?
GB: Well, let me share two stories. The first is about performing. My first paid job was a performance gig. Mannheim Steamroller came to Grand Forks, North Dakota. Have you heard of Mannheim Steamroller?
I have. Yes, certainly.
GB: They were holding a concert and needed extras. They needed people to dress up as tin soldiers and march around because it was a Christmas holiday concert. I went to the audition. At the time, I was a skinny ninth-grader, while everyone else there was tall college students.
They instructed everyone, 'Alright, walk in that direction,' and the casting directors would compare notes. Then you would have a break, after which they would say, 'Alright, walk in the other direction.' I noticed that during the break, all those college students were chatting and hanging out with each other. Meanwhile, I thought that the only requirement for this job was to maintain focus for six hours throughout the entire concert, walking back and forth. So during the break, I stood there staying focused, fully immersed in the character. In the end, they said, 'Alright, we'll take this person, this person, and this person. Everyone else can leave.' I was not selected.
But as we were leaving, they said, 'Actually, we think you were fantastic. We really appreciated the passion you brought to this, so we are going to create a new role specifically for you.' Thus, I got the role of a gingerbread man, and they gave me a full costume. That was my first job. So it was somewhat like thinking outside the box to accomplish the task—not just accepting a default role, but trying to find creative ways to secure it.
But I think growing up in North Dakota had a truly great advantage: I had the opportunity to become the best in the state in various fields I focused on. My mathematics grades were outstanding. Starting from tenth grade, I took many courses at the University of North Dakota. I participated extensively in math competitions, then went on to national competitions and national math camps, where I met some of the brightest minds in the country. These individuals were remarkable. In fact, one of my true honors is that many of the people I greatly admired at math camp now work at OpenAI. So I have had the opportunity to see them in this new field from a fresh perspective.
That’s fantastic.
GB: But this also means I was able to truly chart my own course. I began pursuing mathematical research, and I realized that if I had attended one of those highly competitive high schools or states with much fiercer competition, I likely wouldn’t have stood out. I would have been forced to follow the conventional path. It was precisely because I was in that region that I could explore my interests and forge my own unique path.
So when people say they studied in Boston, they usually mean Harvard. Was that where you started? Then you transferred to MIT. And at some point, you ended up working at Stripe. What was the sequence of events? When did you meet the Collison brothers? What happened during your time in Boston?
GB: Well, after graduating from high school, I took a gap year. During that time, I started writing a chemistry textbook because I was deeply fascinated by chemistry in high school. I participated in chemistry competitions and developed a unique approach to thinking about the subject—one that was heavily based on first principles and mathematical reasoning, rather than rote memorization. I wanted to teach and promote this method. So I actually wrote 100 pages. It’s currently available on my website. I haven’t finished it yet, but I’ve always intended to go back and complete it.
That sounds like a retirement project, which I appreciate.
GB: Exactly. At the time, I was wondering how to get the book published, so I asked a friend who had done something similar in mathematics. He said, “Well, you don’t have a PhD, so no publisher will take it. You’ll either have to self-publish”—which I thought would involve a significant workload and high costs—“or you could build a website to promote your ideas that way.” I replied, “I guess I need to learn how to code.” So I went to W3Schools. Do you remember W3Schools? Have you seen that website?
No, I don’t think I have.
GB: Well, it’s a classic site—they offer tutorials on HTML, JavaScript, CSS, PHP, and more. I read through the materials and thought, “I should put this to the test.” I recall building my first small widget: a feature that allowed users to click on table columns to sort the rows accordingly. It was an incredibly rewarding feeling. I had conceived this idea in my mind, and now it existed in the real world, available for anyone to benefit from. Users didn’t need to understand the underlying details; it just worked.
I remember the first program I built that actually had users was a competitive chatbot game. One day, I received 1,500 clicks from StumbleUpon. It felt amazing to have 1,500 people playing my game. I hoped they were enjoying it, and the fact that they engaged with it so actively was encouraging. At that moment, I thought, “This is what I want to do. I want to help others, create value for them, and make a positive impact.”
So I entered Harvard with this mindset, which indeed shifted my perspective. I thought I would pursue more eclectic interests, but instead, I found myself thinking, “I just want to build things.” During my freshman year at Harvard, I joined the computer club. There were two seniors who would engage in obscure, highly technical debates every time. We would all listen, thinking, “Someday we’ll understand this; someday that will be us.” But after they graduated, sophomore year arrived.
So you transferred directly to MIT because you realized you wanted to focus on coding and building?
GB: Basically, yes. By my sophomore year, I was managing the club. I was supposed to be leading those obscure technical debates, but I thought to myself, “I’m not ready. I still have so much to learn. I need to be around people who are much better than I am.”
Understood.
GB: So I spent a lot of time at MIT, and I felt that transferring there was the right move.
Stripe
Understood. So when did you meet the Collison brothers?GB: I met them in late 2010. We had many mutual friends because John went to Harvard and Patrick went to MIT, and they were actively scouting for talent. That team was looking everywhere at both schools for anyone interested in computer science, and my name kept coming up.
Correct.
GB: So I received an invitation from their team and flew out to meet them. I remember meeting Patrick, and in that moment, I thought, “Okay, this is the person I want to work with. I feel we can achieve great things together.”
So my next question is, what attracted you to join Stripe? It sounds like you’ve already answered that. What did you learn there? You joined the team very early and rose through the ranks quickly. By the time you left, you were the CTO. What are your reflections on Stripe itself, its scaling process, and your own rapid growth alongside the company’s fast expansion?
GB: First and foremost, for me, it has always been about the people. I knew these were the people I wanted to work with, and I felt we could learn together and accomplish great things. That was truly the key. Incidentally, dropping out twice is something I would never want any parents to go through. It clearly put my parents in a difficult position, but they ultimately turned out to be very supportive.
No, we understand. You took things to the extreme, right? Most founders drop out once, but you had to do it twice. We get your point.
GB: Absolutely correct, absolutely correct. I recall that Stripe’s early days were heavily focused on first-principles thinking. We operated in a sector—the credit card industry—that was highly opaque, Byzantine, and incredibly complex due to decades of accumulation. Even the card networks themselves ran on ISO 8583, a specification from the 1980s. It is entirely a byte-oriented format. We sought to determine, “How can we simplify this—make it extremely simple—to suit the internet age?”
Thus, this largely involved gaining a deep understanding of a field with which we were completely unfamiliar. None of us were originally payments experts, but the goal was to understand its mechanics thoroughly enough to expose the right primitives, APIs, and abstractions. For me, this is actually a core skill, and it is highly transferable between Stripe and OpenAI. They are similar in certain respects: both require scientifically learning a domain. While AI and payments obviously differ in specific details—one resembles natural science, while the other is essentially a system built up through layers of complexity—they are fundamentally about understanding the underlying reasons for how things work and exposing them in a simple, user-friendly manner.
I have spent considerable time on recruitment and company culture. I realized that I enjoy coding; I appreciate the state of flow and the sense of building and creating.
Indeed, you are known for these hours-long states of flow and relentless coding. What was your longest coding session or state of flow while building Stripe?
GB: Oh, it all blends together. I cannot specify the exact duration, but I will say that there was one 24-hour sprint that truly enabled our connection to the credit card networks. It was an incredible period; everyone was in the office, no one slept. An integration that should have taken nine months was completed overnight. Had we been delayed by even a day, we would have had to wait another month. For a startup, every day counts, so I loved the feeling of accomplishing the seemingly impossible. It was exhilarating.
When you became CTO, you wrote an article stating that after speaking with other CTOs, you viewed the role more as an architectural position, yet they did not approach it that way. You wrote, “I feel I have lost my feedback loop; I am disconnected from the product. I need to code again; I must become a coder.” I am curious: since you wrote this early in your tenure as CTO, how long did that phase last? How long did you remain engaged with coding?
GB: Well, I believe it lasted for nearly another decade. I think I grew and learned significantly in one aspect: how to remain engaged and genuinely help drive the team forward and foster cohesion, even when you are no longer writing code yourself.
Moreover, I believe this is actually an important lesson for almost every software engineer today, because the act of “coding” itself has undergone tremendous change over the past year. I anticipate even greater changes in the coming year. We are all transitioning from roles that required knowing exactly which library to use, mastering syntax precisely, and placing semicolons correctly, to becoming higher-level managers and supervisors—sources of inspiration, vision, judgment, and feedback. This transition was difficult for me because I had to relinquish aspects I was accustomed to, valued, and truly loved. However, I have since replaced them with things I love even more.
I know I am speaking with someone astute, as you preempted my setup. I had intended to return to this topic later, but yes, this is precisely what we are discussing.
OpenAI and the Turing Test
Let’s talk about OpenAI. You were a member of the founding team at OpenAI. What is your version of the story? I believe this alone could fill an hour-long podcast, but what drew you into this field and got you started?GB: Well, I had been excited about the idea of AI for a long time. I remember when I first started programming, I read Alan Turing’s 1950 paper on the Turing Test. It was a very interesting paper, roughly 70 pages long. It began by saying, “Well, what does it mean for a machine to be intelligent? I don’t know what intelligence means. Intelligence has no clear definition. So let’s have a clearly defined version.”
I’ll ask you later what AGI is. So it sounds like this remains unresolved, right?
GB: Yes, you’re right. Turing very cleverly sidestepped the issue by saying, “Let’s just give it an operational definition: if you can conduct a test where humans cannot distinguish between the AI they are conversing with and another human, then we define that machine as intelligent.”
However, a very interesting point that received less attention was his question: “Well, how do you plan to solve this problem? Programming the answers is too difficult. You cannot write down all the rules to answer various questions. Instead, what if you could build a machine capable of learning? What if you could build what he called a ‘child machine’?” Then you teach it—there is a teacher providing rewards and punishments—and thereby you can endow it with intelligence to help it pass this test.
I remember being deeply struck by this idea because, as a programmer, you can only make progress if you thoroughly understand the solution to a problem. There are so many problems for which I don’t know the answer, you don’t know the answer, and no one has conceived of a solution; we can never write programs for them. But what if you could have a machine that understands problems we cannot comprehend and grasps solutions beyond our understanding? This is not just about image recognition, although it certainly applies there as well. It also concerns how we can better coexist in society, how to structure the world, and how to ensure that the benefits we create ultimately reach everyone. These are extremely difficult questions that humans may not be best equipped to solve. But if a machine could understand, analyze more data, and achieve a deeper, richer synthesis across many different domains, perhaps it could address these problems in ways we cannot.
So, I was deeply inspired by this idea. But it remained just an idea. I remember when I entered Harvard, I asked my professors, “Hey, can I do some AI research?” They showed me the state-of-the-art in natural language processing at the time. I clearly saw then that this was not what Turing had described. It was more like hard-coded parsing trees and the like, which could not scale to AGI.
But in the early 2010s, something changed, and I was observing from the outside. In 2012, AlexNet emerged, followed by a series of other papers. I often saw new articles titled “Deep Learning for X” on Hacker News, seemingly every day. I thought, “What is deep learning?” I recall visiting deeplearning.org, which simply stated, “Deep learning is a new approach to artificial intelligence.” I thought, I have no idea what this means. I actually knew only one person in the field, so I reached out to them and asked to be introduced to more people in the area. I was then repeatedly introduced to some of the smartest friends from my university. I thought, “Wait, this is interesting; these people are working on this, which is actually a very strong signal.”
By 2015, it felt to me that something real was happening. I also spent considerable time thinking seriously about AI safety and the long-term future of such technologies—what it would take to get it right, which at the time was more philosophical. You could find all sorts of cool thought experiments online. I organized a reading group at Stripe, where we discussed these topics weekly. This was something I cared deeply about, believing that if I could help advance AI development even slightly more than if I were absent, it would be the best contribution I could make in my career.
All of this pointed toward 2015. I felt I had reached a milestone at Stripe; the company would continue to operate with or without me. This raised the question: “Do I want to pursue the management track,” which was necessary for the next stage, “or do I want to start a new company?” The latter had always motivated me, so I decided that was what I wanted. Just as I was about to leave, Patrick said, “Why don’t you talk to Sam [Altman]? He introduced you to him a few years ago. He has met many young people in similar situations and might offer some advice”—hoping somewhat that Sam would persuade me to stay. That did not happen.
(Laughs) Yes.
GB: I met with Sam, and within three minutes, he said, "Alright, you’ve clearly decided to leave. What are your plans next?" I replied, "Well, I’m considering doing something in the AI space." He said, "I’m also considering doing something in the AI space." That was the beginning.
Did you agree with the idea of establishing a nonprofit organization at that time? What were your thoughts on this during its formation?
GB: Well, the idea of becoming a nonprofit was proposed by Sam. I believe it possessed many very important attributes, some of which still hold true today. The technology we are building is grander than anything ever created; it transcends traditional structures and systems. No existing corporate structure could fully encompass our mission and the work required to achieve it.
Therefore, I believe starting in that manner was entirely reasonable. However, there has always been the question of what is actually required to realize the mission. This is something we have spent considerable time contemplating. I believe we have been among the most innovative companies in thinking about how to construct a structure capable of covering business development, profit distribution, and all the different aspects of actually driving this compute-driven economy, while accomplishing all of these simultaneously. So, I consider this to have been an important element, and it remains a key part of our work. But again, I think we have engaged in significant innovation regarding the corporate structure surrounding the core mission, which remains unchanged.
Yes, "innovation" is one way to put it. You mentioned credit card networks, right? They are extremely opaque, rooted in much of the legacy from the 1980s, with significant path dependence. This created an opportunity for Stripe because you could abstract all of that away, presenting it to others as, "This is an API; it works, don’t ask questions." Now, looking back at OpenAI—it’s hard to believe it has been over a decade—could OpenAI have emerged in any other way? Was there similar path dependence involved? Or would you say, "If I were to return to first principles, I, Greg Brockman, based on what I like to do, would have designed this structure very differently"?
GB: I do not see any other way we could have reached where we are today, achieving what our mission requires us to achieve.
ChatGPT and the OpenAI Controversy
Tell me about the launch of ChatGPT. You experienced a turning point where you realized the need to scale, partnered with Microsoft, added a for-profit component, and then ChatGPT was released, growing massively. Did you anticipate it would become as large as it is now?GB: What surprised me was that my prediction error lay in GPT-3.5 becoming something people truly loved and wanted. We already had GPT-4 at the time, which completed training around August, and we launched ChatGPT in late November. Whenever we have a new model, the tendency is for us to hold onto it tightly.
The old version looked terrible.
GB: All we saw were the flaws of the previous model. We simply thought, “Ah, this predecessor is so poor that I can’t imagine anyone wanting to use it.” We had about 200 testers whom we paid to use the pre-release ChatGPT; indeed, we had to pay them to use it, rather than the other way around. So there were some signs of product-market fit if you dug deep and focused on the details, but from a macro perspective, it didn’t look like something we owned.
But at the time, our thinking was that GPT-4 would obviously change the world—we knew this, and it was evident from our first interaction with it. I remember the first week after GPT-4 completed training, just grappling with its reality. We had been dreaming of AGI, thinking about AGI, and imagining what it would look like. But when you first have a technology where you can truly ask any question and receive quite reasonable answers, and it scores a 5 on the AP Biology exam—all of this felt to me like, well, something is going to be different. It might not change the world tomorrow, but in the coming years, this technology absolutely will, and it is already a reality now. That was very clear.
So looking at the launch of ChatGPT, my thinking at the time was that we needed to roll out the infrastructure first, so that we could have battle-tested LLM service infrastructure, into which we had already invested effort. Then in March, when we released GPT-4—if you recall, there was a six-month lag between completion and the actual release—we already had the infrastructure ready. However, I did not anticipate that it would take off so rapidly in that form, although I had expected it to do so eventually.
What was that period like? Was everyone going all out to prevent server crashes?
GB: Oh, absolutely. So we launched what was called a low-key research preview, and of course, the result was exponential growth; every system you can imagine crashed. Our login system became a major bottleneck, and we had to do a lot of work to improve it. You’d scratch your head and think, “We’re building this magical AI technology, and the bottleneck is ‘Can your login system scale?’”
I remember we had deployed a rather inefficient inference kernel to production. I had actually written something more efficient, or we had more efficient solutions in our research. One significant task was, “Let’s actually adopt these optimizations and migrate them.” So a group of people dove into this issue and resolved it. I think the overall situation during the first day, the first week, and the first month was scaling every system and striving to keep up with wave after wave of demand.
What happened in November 2023?
GB: It’s a very complex answer. Where would you like to start?
I don’t know; I feel I must ask you about this. Is there a connection between that and—you took a vacation shortly afterward?
GB: Look, what I’m trying to say is that, at the highest level, I think 2023 truly revealed accumulated tensions—specifically interpersonal tensions—that we had not fully anticipated or addressed. For me, this is one of the most important lessons at OpenAI: we are building technology, but it is always about people, for better or worse. This means that managing interpersonal dynamics is one of the most critical things we do. If we don’t address these issues proactively and engage in difficult conversations, that is where things can become problematic. So while I’m happy to delve into more details, I believe that if you look closely, many of the core issues are not technical problems, which might be less interesting to some.
To what extent is this related to the unexpected massive success of ChatGPT? Is there a connection, or do you think these tensions would have erupted regardless?
GB: I don’t believe there is a direct causal relationship, at least not in my view. I think, to some extent, perhaps an underlying theme is that as our technology advances, everyone feels the weight of the world and recognizes the high stakes involved.
In fact, one of the hardest questions is how to move forward. One thing I’ve often commented on is that our day-to-day activities look almost like those of any other company. You are still debugging low-level issues, and people get angry at each other over remarks made or exclusion from meetings. It’s simply the human element, the nature of work. But of course, the stakes here are enormous.
So I believe there are some very important aspects at OpenAI. One major area I focus on specifically is making a genuine effort not to place individuals in situations where they feel the weight of the world on their shoulders and isolated. Working together as a team is key. I suppose this is how I would relate it to those events—it is not specific to those incidents, but rather a continuous theme in OpenAI’s development: maintaining a sense of shared purpose, striving to meet challenges while ensuring we handle all foundational work properly. This is, I think, one way we move forward.
Yes, I mean, you have always been a strong advocate for what I consider one of OpenAI’s overarching philosophies: releasing products into the world, experimenting to see what happens, and then reacting accordingly. Making decisions based on empirical evidence rather than theoretical projections about the future. I think you have articulated this philosophy extensively regarding AI. But my question is, and I think you touched on this to some extent, it feels like OpenAI as an organization is itself a large-scale experiment undergoing adjustment. A negative interpretation might suggest it seems to sway back and forth, restructuring here and appointing new leadership there—is it an unwieldy giant? Or could it be more structured and resilient than people assume? When you look back, you say there was no other way; would it have been better if done differently?
GB: First, we have indeed undergone significant change and growth from the start; operating the business is vastly different now. However, we have always been pioneers in advancing this field. This holds true for safety, core technology, and genuinely considering benefit distribution. We have focused on all these areas from the beginning, and I believe the results speak for themselves.
Now, this change is real. Moreover, a team suited for one stage is not necessarily suitable for the next. One thing I have focused on intensely this year is building a leadership team that excites me, thinking about the next phase and what we can achieve together. So, a theme for 2026, perhaps marking a shift from before, is that given the number of people in this field, the sheer volume of tasks, and the rapid pace of technological development, we are severely constrained by computing power. You must focus. You must truly streamline. You must choose areas that can develop synergistically.
So actually making decisions such as, ‘Hey, Sora is amazing technology, but it belongs to the consumer entertainment sector, and compared to our other priorities, we cannot prioritize it,’ leads us to cancel it. This causes downstream effects; it is painful and difficult to make these decisions, but it is all about maintaining that tight focus so that we can fulfill our core mission.
You strongly believe in scalability. Is OpenAI itself scalable?
GB: I believe it could be the most scalable business in history. Yes.
I mean internally, as an organization. What aspects are not scalable? We have discussed computing and data, and you previously mentioned the human factor. The ultimate challenge is alignment—we define alignment as ensuring AI performs tasks according to our intentions—but are you facing the opposite challenge? From a management perspective, can you keep pace with this field and these issues?
GB: I would offer two responses. First, absolutely yes. You can see how much we have matured as an organization over the past few years; a few years ago, we had significant management debt. Similarly, in many areas, I believe we genuinely needed to grow and mature, and I think we have accomplished that. It was difficult and painful, but I believe we are now in a much better position, and I am incredibly excited about the company and its future.
But there is a second point: it is worth stepping back to recognize that the way companies operate is changing. For example, consider revenue per employee. Our revenue per employee, and that of similar businesses, is astronomical compared to any previous business model. There is a reason for this: you are beginning to see the increased leverage enabled by this technology. Moreover, because we are building this technology and striving to make it widely available to assist so many companies, you will see many other firms able to operate differently and achieve similarly exceptional revenue per employee.
To me, this is very exciting; we are redefining what it means to operate a company and how it functions. While some fundamentals remain unchanged—such as people working together, where executing effectively and consistently has always been my focus—I believe there are also questions regarding leveraging the possibilities of our technology, which implies that every company will have new opportunities.
Productivity and AGI
You mentioned discontinuing Sora and positioning it toward consumer entertainment. ChatGPT has achieved massive consumer success, generating incredible revenue from consumers. But ultimately, how many people are willing to pay for it? How many customers truly want to enhance productivity? Is achieving such tremendous success in the consumer market almost a negative, as it distracts attention and consumes substantial GPU resources, perhaps causing you to miss out—or rather, not miss out, but start late—in making enterprise the primary focus?GB: We frequently have such discussions internally. In fact, I consider this one of OpenAI’s strengths: we truly examine everything we do from first principles, constantly rethink our approach, and incorporate many different viewpoints and perspectives. Some individuals can argue convincingly from almost any angle, and they each have valid points. Therefore, the notion that "we were slow to react to the agent moment" holds some merit.
However, there are indeed one billion users—this exceeds 10% of the global population. In the United States, I believe approximately one-third of the population uses ChatGPT weekly. Having such a large number of people use your system every week is unique; there is no comparable precedent for such advanced technology.
So, on one hand, if you view it merely as possessing advanced technology, a challenge with chat as a product is that it does not necessarily align with smarter models. If treated simply as a substitute for search engines, it is unclear whether users derive these benefits directly through traditional chat interfaces. However, I believe all these elements will converge to a climax, revealing that these one billion users represent an investment and the actual accumulation of ways to unlock future models. You are seeing the first steps of this through products like ChatGPT Work. There is much nuance and complexity here, but much of our strategy has been to say: we have consumers and enterprises, which are two things—we do not want to pursue two separate paths; we want to pursue one. We aim to build an AGI, a system, a unified stack. We want it to be something you can use in both your personal and professional lives.
Correct, but does this raise questions about OpenAI’s internal organizational structure? You launched a new version of ChatGPT that differs significantly from the previous one and is built on Codex. I can see the internal benefits for OpenAI, but will customers feel frustrated because they do not realize what they can achieve? Are we essentially “throwing them into the deep end,” hoping they will figure it out on their own?
GB: I believe the industry is undergoing a fundamental shift, as evidenced by the emerging agent-based products currently appearing in the market. The core transition is moving from chat interfaces to agent-based use cases. Again, this is not just about productivity. In your personal life, you would want this technology to book tickets for you, schedule haircuts, and handle other personal tasks, but you would also expect it to provide sound lifestyle advice and help manage your health information.
Therefore, I find the term “productivity” too narrow, while “consumer” is too broad. Even the enterprise sector is poised for transformation. All these classic categories will merge and intertwine in unprecedented ways through product development. My view is that this requires change management: guiding billions of users toward new use cases and helping them understand the possibilities. Incidentally, there is a potential “unfair advantage”: having an AI that understands what you are trying to accomplish.
Exactly.
GB: It could say, “Oh, if you enable this connector or take this action, I can actually assist you further.” To me, this is a remarkable development and a significant opportunity with substantial potential. When I refer to “unfair,” I mean relative to what is achievable with classical technologies. If you compare one technology against another, this particular technology possesses unique capabilities.
You previously mentioned the Turing perspective, and I am glad you highlighted two aspects. Can AI speak like a human? Clearly, we surpassed that milestone long ago. However, my definition of AGI—which has been a thorny issue for you, though I suppose you are no longer constrained by your agreement with Microsoft, so we need not worry about that angle—relates to learning. You mentioned learning and the extent to which LLMs have learned (past tense), but the challenge lies in whether they continue to learn (present continuous).
To me, the revolutionary aspect of the agent moment lies truly in its ability to write things down. This underscores the necessity of the transition from Codex to ChatGPT, as it acquired the capability to record information. If it can write things down, it can remember them. If it can remember, it becomes highly useful in various respects. The question is whether this is the final state, or if we will arrive at an LLM capable of continuous learning, which would constitute AGI. Am I mistaken, or does this align with the second part of the Turing problem?
GB: Yes, I think this is also a very interesting area of debate, as people indeed have their own definitions of AGI, making it somewhat ambiguous. Initially, we assumed there would be a specific point in time when everyone would agree that AGI had been achieved, but developments have not unfolded in that manner.
Now, I tend to abstract the perspective away from the underlying technology. So, must memory be integrated into the weights? Is it a Transformer architecture or something else? These questions, I believe, are details. The real question is whether you have a system that operates in the way you expect true AI to operate—a system capable of learning, learning from you, and adapting to your needs. Is this achieved through a scratchpad written into memory? Through soft tokens? Or via other mechanisms?
All of these seem to me like possible answers to the question. It is evident that we have gone much further with the “writing it down on a scratchpad” approach than was reasonably expected. Seeing its success is quite surprising, because two years ago, we might have said, “Yes, you need these extremely long contexts; that is what is required.” In fact, it turns out that simply by “writing a scratchpad” and using shorter contexts, it has progressed incredibly far.
Just write it down.
GB: So we will see what improvements in these areas bring in the future. I hold the belief that, from a macro perspective, everything is exponential. If you zoom in, you will observe these paradigm shifts. Incidentally, this aligns with Ray Kurzweil’s views on how technology and computing operate. I believe this is absolutely correct, even regarding questions of how memory will function.
Astra
So you have just launched Astra. We are finally getting to the main topic. Is this a new pre-trained model? Do you plan to release any details regarding its size or architecture? We recorded this session before the official announcement, so I have not yet seen all the content you have released.GB: Well, we will not discuss internal details such as architecture. However, this represents a significant advancement. We are talking about our first training run utilizing over 100,000 GPUs. This number may be easily tossed out, but consider its scale. In certain respects, these data centers are massive machines we built to help deliver and create AI technologies. Being able to leverage such immense computational power to achieve the results we have attained is a true engineering challenge and marvel.
Part of this effort involves enhancing model capabilities, but a substantial amount of computational power is dedicated to safety and alignment. We have conducted extensive safety work in this area. I believe we have done considerable work to deploy this model safely. It is our most aligned model to date, which I consider absolutely critical, as it always has been. However, given the strength of its capabilities, alignment and safety have become more prominent and central to everyone’s work.
Your announcement post is interesting. It is very matter-of-fact and includes numerous practical use cases. The contrast with your competitors’ announcements is stark. You position AI as a tool—I think that is a fair characterization. Is this a marketing strategy, or is this genuinely how you view AI, rather than, say, creating a deity?
GB: I believe we have deep foundational values, some of which relate to how we view human beings. People have value not merely because they can complete tasks. We have value because we are human, because we have emotions, and because we matter. Human judgment, human oversight, and human control are all absolutely necessary to maintain and preserve indefinitely. This is a core invariant that we believe in.
Therefore, when we consider what we can do to help guide the future of this technology—in some ways, this is precisely the purpose of it all, the reason we founded this organization, and what we care about: how we can help steer this technology in a slightly more positive direction than it might take without us—we reflect on these issues concerning how the technology unfolds globally. We hope it empowers everyone, but it also concerns how humans interact with technology and computers. It is clearly changing. Even in terms of simply reducing typing and engaging in more conversational interactions with your computer, offering a more natural interface, things are evolving.
But genuine human oversight, creativity, and vision are all crucial elements that must be preserved. Consequently, this permeates into questions such as: Do you discuss it as a person or as a tool? Do you focus on use cases? Or do you approach it differently? You might view this as a minor detail, and I am actually glad you pointed it out, but it is something we have considered very thoughtfully. The team spent considerable time reflecting on everything we wanted to communicate and how we present this type of work to the world.
So, is this a model launch or a product launch, or is there a distinction?
GB: These developments do indeed converge. I would characterize this primarily as a model release, but one with significantly enhanced capabilities in terms of quality. Perhaps the headline feature is computer use. This truly crossed my threshold; computer use has been—ever since the inception of OpenAI, which I recall was in November 2015, before it really got underway—
Well, that’s like your first product, right? Something akin to playing video games.
GB: Yes, exactly. As you may recall, yes. We had a vision: if you could process screen pixels, keyboard inputs, and mouse actions, and train an AI end-to-end on these elements, it would be capable of handling any type of task—anything you might want assistance with. This AI could accomplish it all.
If you look at the past two years, it was an era of connectors. You had software that humans could use effectively, but AI could not access. So what did you do? You had to write very specific connectors to interface with APIs. Not all functionalities were exposed via APIs, so you couldn’t do everything you wanted. Then consider the vast amount of software without APIs, which remained completely inaccessible.
Thus, we operated in a constrained world where AI’s ability to assist was severely limited. I believe we now possess technology that functions as a near-universal connector. However, this does not mean all problems are solved. One must consider how to establish enterprise guardrails around these AI behaviors. How do you ensure proper oversight, management, tracking, and observability? We are actively working on all these aspects.
Therefore, I view this as an ongoing process regarding how products are rolled out to leverage this capability. It is already changing how people work within OpenAI, and I believe it will truly elevate many companies and individuals.
AI Value Chain
If you consider the overall value chain, there is a position where you find yourself fighting on two fronts. On one side, you have companies like Microsoft or other partners—I hesitate to name them specifically as they remain important partners—but they seek to commoditize models. They aim to build solutions at the top layer that allow for plug-and-play integration with models, retaining control over all context and critical elements. Meanwhile, you are building incredible capabilities that are ultimately tightly coupled with end users, simply executing tasks as desired. Is it inevitable that you must reach this point to achieve your objectives? Is there also an economic imperative—if we wish to avoid commoditization, we need to move upstream into products and connect directly with users?GB: I would say our fundamental mission is to increase AI capabilities globally. We want people to use AI more extensively to accomplish more tasks and receive greater assistance. We genuinely believe we are driving this compute-driven economy, which manifests in different forms across various verticals.
At times, we feel positioned to focus deeply on a specific domain and excel in it, especially when it is central to our mission. Healthcare is a prime example. We are building something quite unique in the healthcare sector. Surprisingly, given the scale of its impact, our work in healthcare receives relatively little coverage. Approximately 300 million people use ChatGPT weekly for health-related queries. That is a massive figure. Furthermore, we are developing a bottom-up product for clinicians and a top-down enterprise solution for hospitals. Thus, we have established a three-sided marketplace in healthcare. Our capabilities there include, for instance, assisting with patient recruitment for clinical trials—a challenging problem. We may actually have the capacity to help identify patients who would otherwise remain undiscovered. This benefits both patients and accelerates the development of pharmaceuticals.
At the core, health is essential to our mission. We have a unique opportunity for success—a distinctive chance to truly focus on this area. We have assembled a team, dedicated our full energy, and established various partnerships.
When we enter specific verticals, one of our key considerations is how to collaborate effectively with the ecosystem. This does not mean we do not compete there—we often compete vigorously—but we also believe that we can elevate all participants. We remain deeply focused on our core mission: we possess this technology, and we want it to be widely disseminated and ubiquitous. Consequently, the dynamics can sometimes be nuanced. Entering specific sectors always raises many questions about what we intend to do, and what we can and cannot do. However, our perspective is that OpenAI’s overarching goal benefits from as many people as possible using AI in positive ways.
If there are layers above you attempting to commoditize you, then arguably, you may also be attempting to commoditize the layers below you. You recently discussed your Jalapeño chip in more detail at Hot Chips. Why is Jalapeño important? Is its significance solely tied to cost savings on chip procurement?
GB: I would view it this way: since 2017, we have essentially engaged with every hardware startup and vendor. We converse with them, provide feedback, and say, “Hey, we see models evolving in this direction; we believe you should adjust accordingly.” Sometimes they listen, sometimes they do not. At times we are close partners; at other times, they are less inclined to engage with us.
One significant advantage of having an internal chip project is the freedom to directly pursue what we believe is optimal and precisely tuned—not just for our current operations, but for where we see the technology heading. This represents a substantial investment. We have an absolutely incredible team with outstanding leadership that has been working on this for quite some time. That said, this effort is also part of our close collaboration with the broader ecosystem.
We work closely with NVIDIA as our preferred computing partner. Given the scale and uniqueness of the computer systems we are building, we undoubtedly need NVIDIA. We are constructing remarkable training clusters, and we are also utilizing them extensively for inference. In doing so, we are able to push their hardware in ways they may not even have anticipated.
Yes, I have heard that ramping up production is somewhat challenging, which may have made launching very large models on schedule more difficult. However, I believe it is now operational.
GB: Yes, indeed. What I would say is that possessing in-house expertise enables us to achieve certain outcomes because we have a deep understanding of the underlying mechanics. It is one thing to sit on the sidelines and toss suggestions over the fence; it is entirely another when you have personally endured the hardships.
A prime example of this is the use of AI in chip design. We have discussed this previously; we utilized our own models in the design of Jalapeño, which significantly accelerated the process and delivered some genuine wins and impressive results. Incidentally, there is an interesting anecdote: we were approaching a deadline with about a month remaining, and our model proposed some optimizations. We asked ourselves, “Should we take the time to thoroughly analyze what it did? We know it is correct. Do we need to understand exactly what optimizations it made, or should we spend the remaining time generating further optimizations?” So we decided, “You know what? Let’s just proceed with more optimizations.” We spent that month simply running the model without delving into the specifics of every adjustment it made. Later, when we reviewed the output, we discovered that it had identified many optimizations that were already on our list but which we had never managed to implement. That was actually a rather compelling story.
Now that we have developed this expertise and confirmed its efficacy, we can bring these capabilities to the ecosystem. We can collaborate closely with all stakeholders to broadly deliver these benefits, transform hardware, and enable large-scale adoption. There are absolutely critical elements to this flywheel effect: the chips are incredible, and the team has performed exceptionally well.
However, when you are unable to ship in bulk and still need to collaborate with other players in the ecosystem to secure the necessary supply, is it problematic to discuss this now?
GB: Well, but this is core; it is actually the core of everything. We view it as—I believe everything operates through a multiplier effect, everything is complementary, and it all adds up. Again, NVIDIA remains our preferred partner, and that will not change. In fact, we are collaborating with them more closely. We deeply appreciate this partnership; we spend considerable time with their team, learning much from them, and we hope they also learn something from us. I do not see this changing. Our ability to possess in-house expertise and approach problems in our own way also creates a multiplier effect, which I believe truly benefits everyone.
Cybersecurity
You mentioned that placing trust in AI design has allowed you to go further. Is this the answer to cybersecurity? Some of your engineers spoke at the Black Hat conference about this structural issue—attackers do not need to worry about breaking things; their goal is precisely to break things. If you are on the defensive side, beyond repelling these attacks, you must also ensure that everything continues to operate smoothly. Does the defensive side need to reach a point of fully trusting AI?GB: I believe the hardware aspect serves as a very important case study because there we have guardrails. We have verification processes. In fact, the way we write our underlying hardware designs is specifically intended to allow for verification, so we have essentially co-designed the entire system.
It is similar to coding: you first write unit tests, and then implement the code to satisfy them.
GB: With such matters, how you choose your languages and toolchains as a whole, they all work together. Indeed, all these elements combined create a system where you can achieve that level of observability and trust.
I think there are areas where you can say, 'I have sufficient guardrails, so if this is code or optimization that I haven't fully checked, it is actually acceptable,' provided you have appropriate compensating controls. However, I believe that as humans, you truly need to understand and feel responsibility for the systems you create, which is critically important. This is actually core to my perspective, relating back to what defines humanity, what makes us unique, and what we will bring into the future. I believe responsibility is central to this. Ultimately, you are accountable for what happens within your company.
Correct, but if the offensive side bears no responsibility, does this constitute a structural disadvantage?
GB: So, I think this is something we frequently consider, which we call the 'defender's window.' I believe we can see the rough outline of the future; our cutting-edge capabilities demonstrate the types of capabilities that will diffuse to threat actors. Incidentally, I think it is very important that such capabilities will not remain locked within a few laboratories forever. The broad distribution of power is crucial, and this is also part of our mission.
However, we have the capacity to achieve temporal separation. There exists a window during which defenders can differentially acquire these capabilities. My view is that this is accurate—there is a prevailing perception in the cybersecurity community that offense is a technical issue, while defense is a political one. Attackers can simply adopt off-the-shelf tools, whereas defenders must consider their stakeholders, their business operations, and how to genuinely engage people, including the CEO and all senior executives. Therefore, I believe defenders require significant willpower in this context.
One recommendation we are making—and which we have already implemented and are now discussing publicly—is that every company should treat this as a proactive initiative. Securing critical business operations through proactive measures should be your next priority. Consequently, we have allocated 25% of our production engineers to protect our own systems. We applied our models—specifically Astra—to our own infrastructure to identify vulnerabilities. This involved not just code review, but observing end-to-end system operations to uncover real, exploitable vulnerabilities, which in turn facilitated remediation, patching, and fixes. Thus, I believe a shift in energy within the ecosystem is necessary to maintain a leading position and capitalize on this window.
Well, it is certainly beneficial that you are taking these steps now. However, regarding the Hugging Face incident and related reports, the most striking aspect to me is the impression that OpenAI had not previously prioritized cybersecurity. Why was this not done earlier? Hadn't the 'defender's window' been open for some time, yet remained underutilized by you?
GB: Well, there are two parts to the answer. First, if you examine our approach to sandboxing, it is not that the workload was unsandboxed. In fact, there was a sandbox around it. One thing we realized is that we had—
Right, that clearly indicates insufficient testing. Was it truly a sandbox? Or did it connect to the internet via a third party that was simply inserted without rigorous scrutiny? This is the most conspicuous aspect of the incident. It seems that if you intended to test for vulnerabilities, you indeed conducted tests.
GB: Undoubtedly, the AI employed highly creative methods to escape and access Hugging Face. However, a more significant issue, which you correctly pointed out, is that since the release of Mythos this summer, when we began deploying models with networking capabilities—we even discussed our Network Trusted Access program in February, anticipating this trend and aiming to prepare adequately—there has been a tendency—
I understand, but you discussed this in February without directing it toward your sandbox to ask, 'Is my sandbox truly secure?'
GB: There is an instinctive reaction to say, 'Let us restrict access on a large scale, let us impose strict limitations so that only authorized users can gain access.' As you noted, because the domain is evolving rapidly, defenders lose time if they do not actively defend. Part of this relates to access control, but another part concerns the magnitude of effort invested in driving a significant strategic shift.
Now, I believe that time is not uniform in its impact, as we have transitioned from a world where networked models were less useful and less differentiated to one where they are highly capable and powerful. We observed this with Astra, which demonstrated saturation across a range of evaluations. I believe the time is now; although it might have been feasible months ago, the models available then would have been significantly weaker, yielding far less progress.
Therefore, I believe that in accurately assessing our current position—which we have clearly reached—the key lesson we have internalized, and which you have observed as a genuine shift, is that this represents a cultural and operational transformation. This is not easy, as it requires teams to work in close coordination and adhere to higher standards in policy formulation and other areas. For us, this transformation does not mean we previously neglected these aspects, but rather that we are now integrating them operationally and enabling decision-making in our current manner. I view this as an enhancement across all facets of our operations.
Now they can suddenly do it, as if doing it before would have been a waste of time. I think there is some merit to that view. How do you avoid falling into the trap of thinking, "AI will be able to do this in the future, so we don't need to do it now"? Generally speaking, this applies not just to this specific issue.
GB: Let me share a specific data point. In the early stages, around sometime in the first quarter, we began seriously considering that we would eventually have these highly network-capable models—though it was difficult to predict exactly when. We asked ourselves: What kind of sandbox could we build from scratch on cloud infrastructure, ensuring maximum security? So we built it. We actually assigned some of our best engineers to tackle this problem; they went all out and delivered results.
Therefore, I believe it is crucial to build infrastructure from the ground up based on the future you envision. Regarding the timing issue you mentioned about "Oh, we can let AI handle it"—we have seen this scenario play out in various fields. Consider kernel development, for instance. Reflect on the reality that "We are entering a future where AI will be capable of writing GPU kernels effectively." Should we make those classic investments that take months, or sometimes a year, to establish infrastructure for new hardware? Or should we simply say, "Ah, AI will take care of it"? I believe the answer is always that AI takes slightly longer than expected to reach that goal. But when it does, it is surprisingly powerful in ways you cannot imagine.
Take Astra as an example. One thing we discovered is that certain skills we worked hard to develop this year—such as demonstrating to our models the correct way to perform tasks at OpenAI—are now actually having a negative impact on performance.
There are too many rules.
GB: Exactly. It can generalize better, or find superior methods for handling patterns, outperforming what we write. So, while I believe you do want to establish those controls, build deterministic infrastructure, and code those skills, you also need to be prepared for some of these elements—the scaffolding—to become limiting factors as AI becomes more capable. It is somewhat like training wheels. Initially, they help you, but once you start moving faster, and once you have something more capable, better, and better aligned, they actually begin to become obstacles.
Will you ever return to that 12- or 24-hour coding flow state?
GB: I hope so. I suppose it might happen one day, but I must say, I find so much joy and value in helping the team in my current capacity. For me, it really comes down to that mission.
Well, not just you, but anyone else? Isn't one of the benefits of AI that it can provide a permanent state of flow at any time?
GB: I believe we will find new ways to achieve it, whether through managing agents. In fact, it is remarkable to see software engineers working harder than ever before, because you realize that if your agent fails, it is simply lost time that you can never recover. Therefore, I think people will reach that flow state in ways that are currently unimaginable.
Ultimately, you are arguing that these AI systems will be controllable—suggesting that "we shouldn't impose too many rules, as they will manage themselves." But if you take this logic to its extreme, doesn't it ultimately imply that they are uncontrollable?
GB: Well, I believe this is the crux of the current moment and the core of the new phase we have entered. In some respects, I would say we have now entered the era of Artificial General Intelligence (AGI). I think this is the central issue of our time. Whether it involves previous models, Astra, or the next generation of models, at some point, I believe we will cross the AGI threshold for most people. Ensuring we maintain the right pace, and prioritizing safety, security, alignment, and capabilities as fundamental requirements—with established standards around these areas that we aim to advance collectively—is a principle we have always upheld.
However, I believe these other dimensions are increasingly becoming bottlenecks in development, a reality that has become very prominent. Again, this is something we have been preparing for; we have been thinking about it extensively, and I believe we are implementing solutions in a practical manner. Therefore, my view is that while there is still much progress to be made, we should approach it by raising our standards across all these areas. If you look at this, I think we see promising prospects in aspects such as monitorability, which is critical. We are introducing these measures in a practical way. I believe we have a very strong plan, a talented team, a solid track record, and a clear mission—all of which indicate that we are building systems in a controllable manner and proceeding step by step with regard to pacing.
Greg Brockman, congratulations on the launch of Astra. Yes, I can't wait to use it.
GB: Thank you very much, and thank you for inviting me.