Four strands since April 2026, one stack underneath
I joined BCG X, the build-and-design arm of Boston Consulting Group, at the Milan office in April 2026, as a visiting forward deployed engineer: I sit with the client rather than behind a delivery team. The four strands share their engineering: multi-agent automation on Google Cloud and Vertex AI, with LLM-based generation on top.
Dark Factory factors out five parts and leaves three per agent
Five parts repeat almost unchanged from one agent to the next, and rebuilding them per engagement is where the time goes:
- tool surface
- guardrails
- evaluation harness
- deployment path
- observability that names the step that failed and why
What is left per agent is the task, the tools it may touch, and the checks that say it did the job. An agent leaves the factory deployed, with its evaluation attached.
Evaluation is half the build on two of the four strands
Brand voice in fashion and luxury is an asset with legal and commercial weight, so the guardrails and the evaluation harness take as much of the work as the generation does. On fraud prediction I built five models on one problem and compared them. The comparison showed what the winner was winning against, and the cases where the ranking between models flips.
An internal Tech Hour for senior BCG X engineers
I run it on agentic AI: how agents are built, where they break, and how to evaluate them.