Before the resume, before the metrics, what I actually believe.
I build AI systems for domains where getting it wrong has consequences, immigration eligibility, security operations, regulated decisions. My focus is narrow and deliberate: constrain what a model is allowed to do with deterministic code, not with instructions it can choose to ignore.
Most AI engineering treats the model as the trusted core of the system and hopes a good prompt is enough. That works until a regulator asks how a decision was made, or an edge case slips past a system with no floor beneath it. The gap isn't capability. It's that almost nobody builds the constraint layer that makes capability safe to deploy.
I've built this three times now: a regulatory boundary that runs as a pipeline stage at HealthCrew.ai, a policy engine at Kroneus that holds under full model compromise, and a graph-constrained retrieval system that cut hallucination 94.5% against a naive baseline for my MSc dissertation. Same thesis, three domains.
Constraint over trust
The model is a component to be constrained, not a system to be trusted. Every regulatory boundary, every eligibility decision, every tool call runs through deterministic code first, the model narrates or reasons, it doesn't decide alone.
Evaluation is the product
A golden-set evaluation gates deployment, it doesn't just describe it. At HealthCrew.ai, eight metric families block any release below threshold before a person ever sees the output. An eval that reports without stopping anything isn't an evaluation, it's a dashboard.
Report the uncomfortable number
False negatives, refusal rates, the times a system correctly said no and the times it shouldn't have, these numbers are more useful than an accuracy score, and almost nobody publishes them. I try to.
Operate it, don't just design it
I built the proof of concept that cleared every exit gate for a database migration, then recommended against it after the real bottleneck turned out to be GPU inference, not the database, at 28 times the cost. Most people show the adoption. Showing the refusal is rarer, and more useful.