PRODUCT August 11, 2026 4 min read

How Gregory Kurtzer is applying open-source operating system models to solve the AI training data crisis.

ultrathink.ai
Thumbnail for: Rocky Linux Founder Launches OpenWALDO for Open AI Training Data

Gregory Kurtzer, the legendary systems architect who preserved enterprise Linux stability by founding CentOS and Rocky Linux, has set his sights on the most contentious bottleneck in modern technology: AI training datasets. With the launch of OpenWALDO, Kurtzer is attempting to apply the rigorous, community-driven playbook of open-source operating systems to the chaotic, legally precarious world of open AI training data.

The OS Blueprint for Open AI Training Data

To understand why Kurtzer is building OpenWALDO, one must look at the history of enterprise computing. When Red Hat altered its downstream source code availability for Red Hat Enterprise Linux (RHEL), it threatened to lock developers into a single proprietary pipeline. Kurtzer responded by launching Rocky Linux under the Rocky Enterprise Software Foundation (RESF), restoring a free, open-source, and binary-compatible alternative for the global developer ecosystem.

Today, the machine learning landscape is facing an even more severe crisis of enclosure. AI models are hitting what researchers call the "data wall." The internet is being systematically walled off by publishers, social networks, and media conglomerates weary of uncompensated web scraping. Meanwhile, the legal status of scraping copyrighted data for training remains an existential threat to AI startups. Kurtzer's OpenWALDO aims to address this by establishing structured, legally compliant, and high-quality open datasets that anyone can use without fear of litigation.

"AI is only as open as the data it is trained on. If the foundation datasets are locked behind corporate paywalls, the open-source AI movement is dead on arrival."

Gregory Kurtzer, Founder of Rocky Linux and OpenWALDO

Breaking the Proprietary Data Bottleneck

Currently, the AI industry is split into two camps. On one side are tech giants like OpenAI, Google, and Meta, who can afford to sign multi-million-dollar private licensing deals with companies like Reddit and Shutterstock. On the other side are independent researchers, academic institutions, and startups who are forced to rely on aging, legally gray public dumps or synthetic data that risks model collapse.

OpenWALDO plans to bridge this gap by introducing standardized frameworks for data curation, deduplication, and attribution. Just as open-source software relies on copyleft and permissive licensing (like MIT or Apache 2.0) to guarantee downstream freedom, OpenWALDO seeks to build a trusted registry of clean, opt-in, or clearly public-domain datasets. This "clean room" approach ensures that models trained on OpenWALDO data are commercially viable and free from copyright taint.

What OpenWALDO Means for AI Builders and Investors

For founders and developers, the project is a potential lifesaver. Building a competitive foundation model is no longer just a compute problem; it is an acquisition problem. By commoditizing the data layer, OpenWALDO lowers the barrier to entry, allowing smaller teams to focus on algorithmic innovation and fine-tuning rather than legal defense and web scraping infrastructure.

For venture capitalists, this initiative signals a shift in where value will accrue in the AI stack. If high-quality training data becomes a shared public utility—much like the Linux kernel became the foundation of modern cloud infrastructure—the competitive moat of proprietary foundation models starts to evaporate. Value will instead migrate upward to application layers, specialized domain expertise, and developer-friendly tooling.

The Ultrathink Takeaway

Gregory Kurtzer's entry into the AI data space is a vital reality check for an industry currently infatuated with proprietary moats. If OpenWALDO can successfully mobilize the global open-source community to curate and license high-quality training data, it will do for artificial intelligence what Linux did for the web server: democratize the infrastructure and unleash a massive wave of downstream innovation.

This article was ultrathought.

Stay ahead of AI

Get breaking news, funding rounds, and analysis delivered to your inbox. Free forever.

Related stories