Self-improving AI systems
AI systems that autonomously improve through experience
The self-improvement challenge
Today’s AI models are largely static. They can be fine-tuned, prompted, and equipped with tools, but they don’t reliably get better simply by doing their tasks. Each deployment starts from the same fixed capability as the last.
We believe the next frontier is AI systems that learn from experience: systems that observe outcomes, remember what they’ve learned, refine their own reasoning, and adapt to new tasks with minimal human intervention. Getting there requires advances across the entire stack, including reinforcement learning, reasoning, memory, verification, planning, and agent architecture.
Cybersecurity is the ideal domain to study this problem. Threats evolve continuously, environments differ across organizations, and success is objectively measurable — an exploit succeeds or it doesn’t, a detection is correct or it isn’t. That gives every interaction a clean signal to learn from, at a pace few other domains can match.
At SPARC, we’re building the algorithms that turn this signal into lasting capability through RL post-training on accumulated experience and through architectures that adapt at runtime. We’re also building the infrastructure that lets organizations run this on their own systems using open-weight models, so AI can continuously adapt to each environment while staying within existing security and safety policies.
The next scaling law may come not from larger models or more data, but from intelligence compounded through experience.
Research focus
Reinforcement learning environments
Build realistic cyber environments for training and evaluating reasoning agents — digital twins that provide rich interaction, objective rewards, and long-horizon tasks suited to post-training research.
Long-horizon reasoning
Real-world cyber operations unfold over hours or days, not minutes. Build persistent reasoning systems that maintain memory, recover from failure, and make coherent decisions across long-running investigations.
Self-evolution
Design algorithms that let AI systems improve their own components — learning better memory organization, synthesizing new verifiers, refining planning strategies, and adapting their workflows as new tasks emerge.
Continual Learning
Cyber threats never stop evolving. Develop training methods that let models and AI systems continuously acquire new capabilities from experience without catastrophic forgetting.
Let’s build AI that improves the way experts do — through experience.