Back to Insights
AI Strategy • Infrastructure • IP RiskApril 2026·8 min read

They Trained the Machine on You. Now They're Selling It Back.

Companies are quietly reclaiming private development infrastructure — and it's not a technical decision. It's a sovereignty play.

Aaron Neufeld · Stratusight · April 2026·Stratusight AI Strategy Series
They Trained the Machine on You. Now They're Selling It Back.

There's a pattern showing up across enterprise IT shops, mid-market software teams, and serious independent operators right now. They're spinning up self-hosted Git infrastructure. They're pulling AI inference behind the firewall. They're auditing which third-party platforms have been quietly ingesting their workflows, their codebases, their ideas.

Most people are calling it a compliance move. Some are calling it paranoia. I'm calling it what it is: the market is waking up to the fact that it spent the last decade handing its intellectual DNA to platforms that used it to build products now competing with the people who created that value in the first place. That's not a conspiracy theory. That's just the business model, finally made legible.

WHAT ACTUALLY HAPPENED TO YOUR KNOWLEDGE

Think about what you've put into public-facing platforms over the last ten years. GitHub repositories. LinkedIn posts walking through how you solve problems. Stack Overflow answers explaining your reasoning. Slack integrations piping your team's decision-making into cloud infrastructure you don't control. Blog posts. Forum threads. Code reviews. Pull request comments explaining why you built something the way you did.

Every one of those artifacts is training data. Not hypothetically — actually. Large language models trained on internet-scale corpora didn't draw a careful line around your professional output and say 'not this one, this belongs to someone.' They ingested it along with everything else. The patterns in how you structure a solution, the vocabulary you use to describe architectural tradeoffs, the heuristics you've built from years of field experience — that's in the model now.

The question was never 'can AI generate code?' The question is whether it can generate your code — shaped by your patterns, your reasoning, your hard-won domain knowledge. The answer is increasingly yes.

GitHub Copilot was trained on public GitHub repositories. Microsoft acquired GitHub in 2018. Copilot runs on OpenAI infrastructure. The same OpenAI that Microsoft has invested billions into. The chain of custody on your public commits runs directly into one of the most commercially aggressive AI ecosystems on the planet. Those are just the documented facts. Connect them however you like.

WHO ACTUALLY CONTROLS THE DATA

Here's the part that doesn't get talked about enough — not because it's hidden, but because it's buried in terms of service that nobody reads until something goes wrong. When you push code to a platform, post on a professional network, or connect a productivity tool to a cloud service, you're entering a licensing relationship. You retain nominal ownership of your content in most cases. What you give up is far more valuable than ownership: you grant a broad, royalty-free, sublicensable licence to use that content for the platform's purposes. And 'the platform's purposes' has quietly expanded over the last five years to include model training.

The critical detail isn't who owns the data. It's who controls what gets done with it. Those are different questions with very different answers.

Platform Controller What That Means for Your IP
GitHub (public repos) Microsoft / OpenAI pipeline Your commits trained Copilot. No opt-out existed at time of ingestion.
LinkedIn Microsoft Posts and engagement patterns used for AI product development. Opt-out added in 2024 — after training data was already collected.
Stack Overflow Prosus / OpenAI partnership Corpus licenced to OpenAI. Community content became commercial training data without contributor consent.
Slack (free/standard) Salesforce Message data may train platform ML features. Enterprise Grid plans offer stricter controls — most orgs aren't on them.
Google Workspace Google / DeepMind Usage patterns inform product AI. Gemini integrations process document content unless explicitly restricted at admin level.

THE SHIFT: BUYING BACK THE SPACE

What's changed in the last 18 months is that the output quality has crossed a threshold. AI-generated code isn't just syntactically plausible anymore — it's domain-aware, contextually coherent, and in specialized areas, genuinely difficult to distinguish from expert output. That's the inflection point that's changing enterprise behavior. If a model has seen enough of your architectural style, your preferred patterns, your industry-specific implementation choices — it doesn't need you specifically anymore. It needs someone who can prompt it with the right context. The gap between 'your expertise' and 'the model's capability in your domain' is narrowing. Fast.

WHAT 'BUYING BACK SPACE' LOOKS LIKE IN PRACTICE

Self-hosted Git platforms like Gitea and Forgejo are seeing significant adoption growth among teams who want version control off the public cloud. Air-gapped AI inference — running models locally or within private VPCs — is becoming standard for any organization handling sensitive IP. The move isn't anti-AI. It's about controlling what trains what, and making sure the answer to 'who benefits from this knowledge?' includes the people who generated it.

Smart organizations aren't abandoning AI tooling. They're restructuring the relationship. Private repositories, local inference engines, walled-off development environments that keep proprietary logic from leaking into shared model training pipelines. They're drawing a line between what they consume from AI and what they contribute to it — and making sure that line runs in their favour.

THE KNOWLEDGE MOAT IS REAL, AND IT'S DRAINING

For decades, the competitive advantage of a skilled technologist or a high-performing engineering team was the accumulated knowledge in their heads and their codebase. That moat was structural. It took years to build and couldn't be easily replicated. Public platforms changed the economics of that moat without anyone noticing, because the extraction was gradual and the returns felt positive. You got visibility. You got community. You got tools that made your work easier. The trade seemed reasonable. You weren't paying with money — you were paying with knowledge, and knowledge felt infinite.

It isn't infinite. Specifically, the value of your knowledge — your particular synthesis of experience, domain context, and problem-solving approach — is finite and irreplaceable. And once it's in the training data, it's not coming back out.

You weren't the customer. You were the dataset. The product is what gets built after you've been fully processed.

WHAT YOU SHOULD ACTUALLY BE DOING

This isn't a case for shutting down your LinkedIn or pulling every public repo. Public presence still has real value — for credibility, for business development, for thought leadership that drives actual pipeline. The play is intentionality, not retreat.

Separate your signal from your source code. Share your perspective, your frameworks, your conclusions — the synthesized output. Keep the underlying implementation logic, the proprietary decision trees, the hard-won domain heuristics in controlled environments.

Be thoughtful about which AI tools you use for what. Tools that run inference locally or in your own infrastructure don't feed a shared training pipeline the same way cloud-based assistants do.

Audit your toolchain the way you'd audit a vendor contract. What data does this platform ingest? What are the training terms? Who owns the derivative models? These aren't theoretical questions anymore. They're procurement hygiene.

If you're building a team, or advising one: the conversation about self-hosted infrastructure, private AI, and knowledge governance belongs in the architecture review, not the compliance checklist. By the time it hits compliance, the valuable patterns are already out the door.

THE COMPANIES THAT MOVE FIRST WIN THE SECOND HALF

The organizations restructuring their infrastructure now aren't doing it because they're scared of AI. They're doing it because they understand that the next competitive advantage isn't access to AI capability — everyone will have access to AI capability. The advantage will be proprietary context: private training data, domain-specific fine-tuning, institutional knowledge that hasn't been commoditized yet.

The companies that spent the last decade sharing everything publicly are going to spend the next decade rebuilding the moat they didn't know they were draining. The ones paying attention now get a head start.

"The machine was trained on you. The question is whether you're going to let it keep running on your tab."

At Stratusight, we work with Canadian leaders in technology and government who are thinking seriously about AI sovereignty, infrastructure governance, and what it means to compete when capability is commoditized. If this resonates, we'd welcome the conversation.

Share this article

Related Insights

Beyond the Pilot
AI Strategy

Beyond the Pilot

Building AI strategies that actually stick. How to move beyond proof-of-concept to scalable adoption that transforms how your organization thinks and delivers.

Read More
Working in Uncertainty: The New Operating Condition
Strategic Perspectives

Working in Uncertainty: The New Operating Condition

The most effective leaders don't wait for clarity — they create it. Here's how to build decision-making frameworks that thrive in ambiguity.

Read More
The New Manipulation
Threat Intelligence

The New Manipulation

Understanding coordinated inauthentic behaviour in the age of AI — and why every Canadian leader needs threat literacy now.

Read More