rMax.ai
Enterprise AI Engineering

Research

I conduct independent, applied research at the intersection of AI systems, software engineering, and agent governance. My work focuses on understanding how modern AI-enabled tools actually behave in production—not in demos, benchmarks, or abstractions.

The emphasis is practical: studying real systems, extracting reusable patterns, and publishing findings as open technical notes and artifacts.

Research directions

The program studies how to make agentic systems more governable, verifiable, inspectable, and useful in real workflows.

Governed agent control planes

Move authority out of model outputs and into deterministic policy, typed state, checkpoints, evidence, recovery, and explicit stopping rules.

Computation authorization and provenance

Study how identity, delegated capability, information flow, and execution history constrain whole agent trajectories—not just individual tool calls.

Evidence-first software delivery

Measure whether durable specifications, hidden verification, operational checks, and evidence bundles improve generated software reliability.

Research and knowledge compilers

Treat sources as untrusted input, claims as canonical units, and graphs, search indexes, and reports as rebuildable projections.

FDE and organizational readiness

Understand which repository, workflow, and team conditions make autonomous delegation safe, useful, and sustainable.

Protocols, runtimes, and model fit

Explore MCP interoperability, durable execution, harness adaptation, prompt-stack promotion, and the quality–cost–latency frontier.

Active research programs

These public repositories are research instruments: experiments, benchmarks, and reference systems. Findings are published as evidence becomes available.

Benchmark

HarnessFit

Model-specific agent harness profiles and their effect on task success, transfer, and cost.

Benchmark

Regenerable Software Lab

Implementation-independent verification for AI-generated software and durable failure assets.

Assurance

Evidence-First Harness

Claim–evidence–decision bundles for governing changes made by coding agents.

Adaptive systems

Enterprise AutoHarness Lab

Synthesizing deterministic controls from agent failures without ceding policy to the model.

Knowledge systems

Kiln

Filesystem-first, evidence-backed knowledge compilation.

Knowledge systems

DebateLab

Traceable deliberation through structured claims, evidence, challenge, and revision.

Org systems

Agentic Engineering Org Lab

A control plane for repository readiness, delegated execution, verification, and bounded autofix.

Evaluation

PromptBench

Reproducible prompt-style evaluation with deterministic grading, cost, latency, and variance.

Evaluation

PromptStackBench

When a persona should become a skill, agent, or harness.

Identity

Agent Identity Lab

First-class identity, delegated authority, and authorization for agents.


Agentic Workflows

Agentic Workflows is a research project focused on building a usable map of agentic workflows across the ecosystem.

The work centers on compiling an ontology of workflow types, extracting recurring patterns, and grounding those abstractions in concrete real-world instances. The objective is to make agentic systems easier to compare, reason about, and design intentionally.

Core questions include:

  • Which workflow primitives recur across agentic systems?
  • How do patterns differ between orchestration, delegation, memory, and verification flows?
  • What concrete implementations reveal the gap between theory and operational reality?
  • Which ontology is stable enough to support analysis across tools and domains?

The goal is a practical field guide for understanding how agentic workflows are actually structured, not just how they are described in product language.

agentic-workflows.rmax.ai


System Prompts Forensics

System Prompts Forensics is an ongoing research project dedicated to analyzing how contemporary AI tools structure their system prompts.

System prompts are treated here as first-class system components: mechanisms that encode authority, scope, permissions, constraints, and control over agent behavior. Rather than viewing them as opaque or proprietary text, the project examines them the same way one would examine APIs, protocols, or execution models.

The research is comparative and grounded in real tooling, including developer assistants and agent-based systems.

Core questions include:

  • How is authority delegated between humans, tools, and agents?
  • How are constraints and safety boundaries encoded?
  • What governance patterns recur across different tools?
  • Which prompt primitives are reusable across systems?

The goal is not optimization for a single model, but clarity about how agent control actually works and how it can be designed deliberately.

system-prompts-forensics.rmax.ai


Research Principles

This work follows a few simple principles:

  • Applied over theoretical — grounded in real systems and artifacts
  • Comparative over anecdotal — patterns emerge across tools, not opinions
  • Open by default — notes, reports, and taxonomies are published publicly
  • Engineering-first — prompts are analyzed as system design, not copywriting

Outputs

Research outputs typically include:

  • Technical notes and essays
  • Comparative reports
  • Prompt taxonomies and primitives
  • Open repositories with artifacts and documentation

All material is written to be useful to practicing engineers designing or operating agent-enabled systems.


Scope and Non-Goals

This is not academic research, benchmarking work, or product marketing. It is an engineering-driven effort to make agent-first development more legible, governable, and repeatable.


Contact

If you are working on AI tooling, agent systems, or developer platforms and want to discuss patterns, failures, or governance models, you can reach me via the links on this site.