Skip to content
Logo MandaMANDA.
AI Engineering Lab / Build log

AI ENGINEERING LAB

Building useful agents and showing the evidence: architecture, tools, safeguards, deployment and limitations of systems designed with Claude and OpenAI.

Independent lab. Manda is not affiliated with, certified by, or sponsored by Anthropic or OpenAI.

ClaudeShipped and documented

Claude / Anthropic

Claude Code is used as an engineering environment on real products: repository understanding, implementation, review, tests and Git-based delivery. MCP and n8n extend that work through bounded tools and human approval.

  • Claude Code
  • MCP
  • Git + tests
  • Human approval
OpenAIPublic research track

OpenAI / Codex

The next public track applies the same discipline to OpenAI: typed tools, structured outputs, traces, evaluation sets and action control. The benchmark will publish failures, cost and latency — not only a happy-path demo.

  • Agents & tools
  • Structured outputs
  • Evals
  • Cost & latency

Publishing standard

A demo is not production evidence.

Every new lab study must make these six elements verifiable.

  1. 01

    Problem

    A bounded business task with an observable outcome.

  2. 02

    Architecture

    Model, data, tools, permissions and stop conditions.

  3. 03

    Evaluations

    Normal, incomplete, conflicting and adversarial cases.

  4. 04

    Safeguards

    Sensitive actions stay in draft or require approval.

  5. 05

    Operations

    Logs, cost, latency, recovery and model changes.

  6. 06

    Limits

    Known failures and decisions that must remain human.

Available evidence

Products and systems already built

Product built with Claude Code

Factumation

An invoicing application connected to business automation: assisted design, Next.js integration, controls and continuous delivery.

  • Claude Code
  • Next.js
  • n8n
  • Webhooks
View case study

Documented MCP experiment

Claude-operated Webflow

A CMS operated through two MCP servers, covering publishing, SEO and explicitly documented API limitations.

  • Claude
  • MCP
  • Webflow API
  • SEO
View case study

Multi-provider architecture

ZeroClaw

A Rust agent infrastructure compatible with several providers, including OpenAI and Anthropic, designed for multiple tools and channels.

  • Rust
  • OpenAI
  • Anthropic
  • Tool use
View case study

Next protocol

Same task, same data, two models.

The next benchmark will compare Claude and OpenAI on a bounded research agent. Planned measures include task success, source quality, tool calls, cost, latency and human recovery.

  1. 01Create a versioned scenario dataset
  2. 02Run both variants with the same tools
  3. 03Publish results, redacted traces and failure cases

Hiring or working on a demanding use case?

I can walk through the architecture, code and trade-offs behind these systems — in English or French, remotely for teams in Europe and the United States.

Discuss the system