Meta Says Watermelon Is Catching Up to GPT-5.5. What Is Actually Known?
Business Insider reports that Meta AI chief Alexandr Wang told employees Watermelon has caught up to OpenAI's GPT-5.5 on benchmarks. The caveat: Meta has not released the model or named the tests.
Business Insider reports that Meta AI chief Alexandr Wang told employees Meta's unreleased Watermelon model has caught up with OpenAI's GPT-5.5 on benchmarks. The claim matters because it suggests Meta may be closing a visible frontier-model gap, but it is still not independently verifiable: Watermelon has not been released, Meta has not named the benchmarks, and GPT-5.6 is already entering a limited preview.
Definition: Watermelon is Meta's internal codename for an upcoming model after Muse Spark, according to Business Insider.
Example: Wang reportedly told employees that Watermelon is using far more training compute than Avocado, the internal codename for Muse Spark.
Key takeaway: Treat the report as a serious competitive signal, not as public proof that Meta has matched OpenAI across real-world tasks.
Business impact: Teams should watch the release, but model decisions still need workflow tests on quality, cost, latency and tool use.
What did Business Insider report?
Business Insider reported on July 2, 2026 that Alexandr Wang told Meta employees the company's upcoming model, codenamed Watermelon, has caught up with OpenAI's GPT-5.5 on closely followed AI benchmarks. The article cites two people familiar with the matter and says Wang discussed the claim at an internal town hall. It also adds the most important caveat: the specific benchmarks were not clear.
That makes the story narrower than a public launch. It is not a benchmark paper, a model card or an API announcement. It is a reported internal update from Meta's AI leadership at a moment when the company is trying to show that its rebuilt AI organization can compete with OpenAI, Anthropic and Google.
What is Watermelon?
Watermelon appears to be Meta's next frontier model after Muse Spark. Business Insider says Wang described Watermelon as the model after Avocado, Meta's internal codename for Muse Spark, and said it uses an order of magnitude more compute than Avocado. Meta has not published a public Watermelon announcement, so the model's architecture, context window, pricing, release date and safety profile are not yet known.
That missing detail is the story. A model can look strong on internal benchmarks and still land differently in developer use. For business automation, the release will matter only when teams can test whether Watermelon improves practical outcomes: coding task completion, tool-call reliability, long-context accuracy, structured output and cost per completed workflow.
How does Muse Spark fit into the story?
Muse Spark is the public baseline for understanding Watermelon. Meta introduced Muse Spark on April 8, 2026 as the first model in its Muse family from Meta Superintelligence Labs. Meta described it as a natively multimodal reasoning model with tool use, visual chain of thought and multi-agent orchestration, and said it was available in Meta AI with a private API preview for selected users. More on this: Muse Spark 1.1: 1M context aimed at agentic coding.
Meta's own launch post was candid about the next gap. It said Muse Spark offered competitive performance in multimodal perception, reasoning, health and agentic tasks, while Meta would keep investing in "current performance gaps" including long-horizon agentic systems and coding workflows. Business Insider's new report frames Watermelon as the next step in that same scaling path, especially for coding and agentic capabilities.
What is GPT-5.5 in this comparison?
GPT-5.5 is OpenAI's April 2026 model release and the benchmark target in the reported Meta claim. OpenAI's GPT-5.5 announcement positioned the model around agentic coding, knowledge work, scientific research, tool use and computer-work workflows. OpenAI also published headline scores including 82.7% on Terminal-Bench 2.0, 78.7% on OSWorld-Verified and 51.7% on FrontierMath Tier 1-3.
Those public scores do not automatically tell us what Meta matched. Business Insider says Wang cited closely followed benchmarks, but the article does not identify which ones. If Watermelon matched GPT-5.5 on coding benchmarks, that would mean one thing. If it matched on a broader internal index or a narrower subset, the practical meaning would be different.
Why is the claim hard to verify?
The Watermelon claim is hard to verify because the model is unreleased and the benchmark set is unnamed. Public benchmark comparisons are already noisy: labs choose different harnesses, prompting strategies, pass counts, tool access and safety settings. Private comparisons add another layer of uncertainty because outside evaluators cannot reproduce the setup.
The safest interpretation is parity on some internal or closely watched evaluations, not blanket parity across every real workflow. Businesses should wait for three things before treating Watermelon as production-relevant: a public model release or API preview, benchmark methodology, and independent tests from developers using realistic tasks.
What changed around GPT-5.6?
OpenAI has already moved beyond GPT-5.5 in limited form. Business Insider reported that OpenAI launched a limited preview of the GPT-5.6 series in late June 2026, including Sol, Terra and Luna, with broader access delayed at the U.S. government's request. Axios also reported that GPT-5.6 access began under restrictions and that advanced cyber capabilities are one reason Washington is watching frontier releases more closely.
That timing changes the competitive reading. Matching GPT-5.5 would be a major improvement for Meta if it holds up publicly, but the frontier target is moving. The relevant question is not only whether Watermelon caught GPT-5.5; it is whether Meta can ship models quickly enough, safely enough and cheaply enough to stay close as competitors keep releasing.
Why does this matter for Meta?
The report matters for Meta because it directly tests the story behind Meta Superintelligence Labs. Meta has spent heavily on AI infrastructure and talent, and Business Insider says Wang now oversees an elite research group known as TBD alongside other AI efforts. If Watermelon really narrows the gap with OpenAI, it would support Meta's argument that its rebuilt stack is scaling.
The public proof is still ahead. Meta's April Muse Spark post emphasized improved scaling efficiency, multi-agent reasoning and safety evaluations, but Muse Spark did not fully erase the perception gap with OpenAI, Anthropic and Google. Watermelon is the model that may show whether Meta's compute, data and talent push has translated into a top-tier product.
What should teams watch next?
Teams should watch for release details rather than react to the headline. The useful signals will be whether Meta publishes Watermelon benchmark methodology, whether an API becomes available, whether independent evaluators can reproduce the claimed gains, and whether the model improves practical agentic tasks such as coding, browser work, customer workflows and tool calling.
The near-term model-selection advice stays unchanged: keep a fixed evaluation set and compare models on your own work. A new frontier model is interesting when it lowers cost, improves accepted output quality, handles longer context reliably, calls tools more cleanly or reduces review time. Until Watermelon is available to test, it belongs on the watchlist, not in the production routing table. For a broader model-evaluation process, see What the Latest AI Models Mean for Your Business.
What remains uncertain?
The biggest uncertainty is whether Watermelon will perform as well in public use as it reportedly does in Meta's internal benchmark view. The second uncertainty is timing: a model still in training can change before release, and competitors can move again before it ships. The third uncertainty is access: a strong model matters less to developers if it is not broadly available, priced competitively or documented well.
For now, the story is best read as a competitive signal. Meta appears to believe its next model is much closer to OpenAI than Muse Spark was. The open question is whether that claim survives public benchmarks and real workflow testing.
Frequently asked questions
What is Meta's Watermelon AI model?
Watermelon is Meta's internal codename for an upcoming AI model that is still in training, according to Business Insider. Alexandr Wang described it as the next model after Avocado, the internal codename for Muse Spark. Meta has not publicly released Watermelon, published a model card or disclosed the benchmark results behind the reported comparison to OpenAI's GPT-5.5.
Has Meta released Watermelon?
No. Based on the Business Insider report, Watermelon was still in training when Alexandr Wang discussed it internally. Meta's public model release in April 2026 was Muse Spark, not Watermelon. Until Meta ships Watermelon or publishes official evaluation data, outside users cannot independently test the claim.
Did Watermelon beat GPT-5.5?
Business Insider reported that Wang said Watermelon has caught up with OpenAI's GPT-5.5 on closely watched benchmarks. The article does not say Watermelon beat GPT-5.5, and it notes that the specific benchmarks were not clear. The safest reading is that Meta is claiming parity on some internal or tracked evaluations, not a verified public win.
Why does the Watermelon claim matter?
The claim matters because Meta has been trying to prove that its model stack can compete with OpenAI, Anthropic and Google after heavy spending on talent and infrastructure. If Watermelon performs near GPT-5.5 in public use, developers may get another strong option for coding, agentic tasks and multimodal workflows. If the claim remains private, it is mainly a signal of internal momentum.
Should businesses switch AI models because of this report?
Businesses should not switch AI models because of an unreleased model claim. The practical move is to watch for a public Watermelon release, wait for benchmark details, and test it against the same workflow set used for current models. Model choice should be based on task quality, latency, cost, tool use and reliability, not on one internal benchmark statement.
Sources
Alex
Founder & Lead AI Writer
Alex is the founder of Yowox and lead AI writer since 2024, breaking down complex information into clear, actionable insights for thousands of readers every day. Alex has built AI automation systems for businesses since 2024, focusing on AI agents, workflow automation, and business process optimization.
Save hours. Save thousands.
Practical guides, real workflows, and the latest AI and automation news that matters — straight to your inbox.