Governed agent control planes
Move authority out of model outputs and into deterministic policy, typed state, checkpoints, evidence, recovery, and explicit stopping rules.
I conduct independent, applied research at the intersection of AI systems, software engineering, and agent governance. My work focuses on understanding how modern AI-enabled tools actually behave in production—not in demos, benchmarks, or abstractions.
The emphasis is practical: studying real systems, extracting reusable patterns, and publishing findings as open technical notes and artifacts.
The program studies how to make agentic systems more governable, verifiable, inspectable, and useful in real workflows.
Move authority out of model outputs and into deterministic policy, typed state, checkpoints, evidence, recovery, and explicit stopping rules.
Study how identity, delegated capability, information flow, and execution history constrain whole agent trajectories—not just individual tool calls.
Measure whether durable specifications, hidden verification, operational checks, and evidence bundles improve generated software reliability.
Treat sources as untrusted input, claims as canonical units, and graphs, search indexes, and reports as rebuildable projections.
Understand which repository, workflow, and team conditions make autonomous delegation safe, useful, and sustainable.
Explore MCP interoperability, durable execution, harness adaptation, prompt-stack promotion, and the quality–cost–latency frontier.
These public repositories are research instruments: experiments, benchmarks, and reference systems. Findings are published as evidence becomes available.
Model-specific agent harness profiles and their effect on task success, transfer, and cost.
Implementation-independent verification for AI-generated software and durable failure assets.
Claim–evidence–decision bundles for governing changes made by coding agents.
Synthesizing deterministic controls from agent failures without ceding policy to the model.
From per-call permissions to authorization of an entire computation trajectory.
Causal and provenance graphs for reproducible agent-security analysis.
Filesystem-first, evidence-backed knowledge compilation.
Traceable deliberation through structured claims, evidence, challenge, and revision.
Evidence-linked training and evaluation for forward-deployed engineering.
A control plane for repository readiness, delegated execution, verification, and bounded autofix.
Typed, deterministic, restart-safe agent workflows.
Reference-indexed execution for long-running, multi-document research.
Reproducible prompt-style evaluation with deterministic grading, cost, latency, and variance.
When a persona should become a skill, agent, or harness.
First-class identity, delegated authority, and authorization for agents.
Agentic Workflows is a research project focused on building a usable map of agentic workflows across the ecosystem.
The work centers on compiling an ontology of workflow types, extracting recurring patterns, and grounding those abstractions in concrete real-world instances. The objective is to make agentic systems easier to compare, reason about, and design intentionally.
Core questions include:
The goal is a practical field guide for understanding how agentic workflows are actually structured, not just how they are described in product language.
System Prompts Forensics is an ongoing research project dedicated to analyzing how contemporary AI tools structure their system prompts.
System prompts are treated here as first-class system components: mechanisms that encode authority, scope, permissions, constraints, and control over agent behavior. Rather than viewing them as opaque or proprietary text, the project examines them the same way one would examine APIs, protocols, or execution models.
The research is comparative and grounded in real tooling, including developer assistants and agent-based systems.
Core questions include:
The goal is not optimization for a single model, but clarity about how agent control actually works and how it can be designed deliberately.
→ system-prompts-forensics.rmax.ai
This work follows a few simple principles:
Research outputs typically include:
All material is written to be useful to practicing engineers designing or operating agent-enabled systems.
This is not academic research, benchmarking work, or product marketing. It is an engineering-driven effort to make agent-first development more legible, governable, and repeatable.
If you are working on AI tooling, agent systems, or developer platforms and want to discuss patterns, failures, or governance models, you can reach me via the links on this site.