The next time you find yourself on EDGAR, consider using my SECSift reader instead!

It turns SEC filings into clean, readable documents. Search every filing on EDGAR, get alerts on what changed, let AI highlight what matters and gray out the rest, and compare against previous filings.

Not familiar with the reader? Here’s a 3-minute interactive demo (viewable on your computer) and another video I made walking through the entire tool.

🌾 Welcome to StableBread’s Newsletter!

Practically everyone knows about OpenAI’s ChatGPT today, but not many are familiar with InstructGPT, what OpenAI calls a sibling model to ChatGPT.

To train the model to follow instructions, OpenAI hired “a team of about 40 contractors” to write sample answers and rank the model’s outputs.

Even with far more capable models today, OpenAI and other top AI labs still pay experts to write and grade the problems their models get wrong. For example, in October 2025, OpenAI was paying more than 100 former investment bankers $150/hour to build financial models for its AI.

These top AI labs also pay for practice environments, simulated software where a model attempts a task and gets scored on the result.

And that’s where data labeling companies play a role. They recruit experts in fields where models still make mistakes, like coding, finance, and medicine, and turn what those experts know into graded examples and practice environments for the labs to train on.

Take a look at the valuations of the largest AI data labeling companies:

AI Data Labeling Valuations (Scale AI, Mercor, Bloomberg, Forbes, TechCrunch; October 2026)

Clearly, investors expect AI labs to keep paying for expert data.

One of the skills the labs need help with is reading malware, the malicious software hackers use to break into or lock up computers.

Labs like OpenAI and Anthropic now offer their top models to companies defending against cyberattacks, through programs like Daybreak and Project Glasswing.

A model can only help stop an attack if it can tell what the malware does.

Malware often arrives as a binary, a compiled program in machine code, and many are obfuscated, or scrambled to hide their purpose.

Security analysts test suspicious programs on an isolated computer, called a sandbox, to see what the programs do. But some malware detects the sandbox and stays dormant, waiting for a real computer.

Picus Security, a security testing company, analyzed 1,084,718 malware samples for its 2026 report on the techniques hackers use most. Sandbox evasion grew faster than any other technique, rising to the fourth most common.

In August 2026, researchers led by Columbia University released SRE-Bench, a test of AI models on 262 binaries compiled from 19 programs they wrote from scratch, and concluded that reverse engineering, or figuring out what a binary does, “remains largely unsolved.”

According to Vals AI’s SRE-Bench leaderboard, OpenAI’s GPT-6 Astra claims the top spot, fully solving 56.9% of the binaries vs. 30.5% for GPT-5.6-sol, the top model when SRE-Bench came out on August 11.

But GPT-6 Astra still can’t fully solve the other 43.1%. When the Columbia University researchers updated their SRE-Bench results on October 1, they found that with its cyber safeguards off and no spending limit, GPT-6 Astra usually got a binary right within four tries but still couldn’t reliably tell which answer was correct.

Data labeling companies are now recruiting malware experts to train the labs’ models. Scale AI, for example, is hiring a product manager to build a cybersecurity product line, with “malware and binary analysis” among the skills it covers.

Mercor also pays security experts $200-250/hour to evaluate AI-generated analyses, preferring people with reverse engineering or binary analysis backgrounds.

So why am I writing about AI data labeling?

Because data labeling companies’ demand for malware skills has now reached a ransomware nanocap I wrote up on September 27, 2026.

The company’s software copies the encryption keys ransomware creates during an attack, so its engineers can build a decryptor (a program that reverses the encryption) and recover a victim’s files without paying the ransom.

Its CEO said in a new interview that data labeling companies have asked the company to help label malware, and that most of its engineering focus is now on training its own AI model to read malware binaries.

The CEO also expects the third of its four AI patents to be approved within two months.

I’m considering the interview a material change to the thesis, which requires an updated analysis.

So is a company the market values at under $10M on track to be acquired by, or partner with, an AI lab or data labeling company? Or is binary analysis another AI plan that misses its targets, like the AI decryptor builder it has pitched since 2023?

Below, I provide an update on where the pending patents stand, then focus most of the write-up on what the data labeling companies want and who’s likely behind them, what the binary analysis model changes for investors, and what could keep it from paying off. Finally, I conclude with the milestones I’m tracking and whether I’m adding to my position.

Like my original write-up, this is an illiquid nanocap, and with a sizeable list, the name and the rest of this update only go to paid subscribers.

logo

Subscribe to StableBread's Research!

Full access to this write-up and every future one. Cancel anytime.

View Plans

Subscribe to unlock:

  • 50+ write-ups every year
  • Buy/sell/hold decisions
  • Periodic industry, market, and thematic research
  • Deeper analysis and valuation models
  • Full archive of all past research

Reply

Avatar

or to participate

Keep Reading