— MARCELUS FERNANDES · TECHNICAL KNOWLEDGE SUMMARY
Model behaviour
SPECIFY · OBSERVE · CHANGE
Evals and testing
MEASURE · CALIBRATE · GATE
Agent workflows
ORCHESTRATE · CONSTRAIN · RECOVER
Product leadership
DIRECT · TRANSLATE · DEVELOP
Product behaviour emerges from the model, prompt, state, history, retrieval, tools, routing, policies, channel and evaluation loop acting together.
I currently lead AI initiatives at AB InBev and work hands-on in Python and TypeScript through AI-assisted engineering. My ownership spans behavioural boundaries, agent architecture, the evaluation programme, retrieval quality, safety, observability requirements, and the operating model around the team.
14+ years in product and design are the substrate. I have led a 26-person design organisation, worked as a Principal and Staff designer, shipped zero-to-one products, and built operating systems. The tools changed. The responsibility did not: make complex systems legible enough to shape with intent.
State, context, memory, tool contracts, agent loops, retrieval, routing, evals, judges, safety, provider semantics, observability and iteration.
MY PRIMARY HANDS-ON WORKING LAYER
Application and operations
Channels, interfaces, integrations, product flows, deployment constraints, support operations and business outcomes.
BUILT AND LED, BUT NOT THE DIFFERENTIATING CLAIM
01 · TECHNICAL KNOWLEDGE
Eight connected areas form my working model of an AI product. The boundaries between them matter more than the labels: most failures emerge where one subsystem hands behaviour to another.
01
What should the system do, and what history teaches it to do next?
02
What must the model see now, what should survive, and what must be forgotten?
03
How do we know behaviour improved, and whether the instrument itself is trustworthy?
04
Which steps require model judgement, which require code, and how does work retain lineage?
05
How do we maximise useful recall without letting ranking, business logic or generation corrupt relevance?
06
What must the model never decide alone, and where must policy become architecture?
07
What changes when the model, provider or inference surface changes?
08
Can the team locate a failure, explain a cost change and protect the data needed to learn?
02 · TECHNICAL JUDGEMENT
Tools change quickly. These are the principles I use to decide what belongs in prompts, code, data, policy and team practice.
01 Prompts guide. Architecture controls.
If a rule matters, it must exist in authorisation, routing, validation or an invariant. A prompt-only constraint is a preference the model may ignore.
02 State is part of the experience.
What a system remembers, forgets and presents as precedent changes behaviour as much as the visible response.
03 Evaluate the evaluator.
A judge is a measuring instrument. Before it gates a launch, its variance, bias, schema, rationale-verdict consistency and hard classes need their own evaluation.
04 Use probability for judgement, code for invariants.
Semantic quality often needs a model. Credential absence, tool permission, schema validity, price evidence and route downgrade should not be scored. They should be enforced.
05 Preserve failed attempts.
Negative results, regressions and rejected mechanisms are part of the knowledge system. Removing them makes a team repeat expensive loops and overestimate certainty.
06 Complexity must earn its place.
More agents, routes, memories or judges create more failure surfaces. They stay only when an ablation, operational constraint or clear product boundary justifies them.
03 · EVIDENCE INDEX
SYSTEM
TECHNICAL TERRITORY
INSPECTABLE ARTIFACTS
Conversational agents (NDA)
Production agent architecture, state, context, tool traces, behaviour campaigns, evals, retrieval, safety, model routing and observability requirements.
Graph and state contracts · eval datasets · judge schemas · documented campaigns · safety regressions · retrieval architecture
Lohra
Inference runtime, provider semantics, opaque reasoning-state continuity, prompt snapshots, compaction lineage, orchestration, workflow DSL and capability security.
Agent loop · provider adapters · session lineage · typed operators · sandbox paths · extensive invariant suite
Laura
Longitudinal behaviour, canonical identity, surface policy, episodic and durable memory, context rotation, relationship-specific behaviour and response timing.
Behaviour audit · identity contract · policy map · memory jobs · runtime histories · pacing tests
Synthetic Users / PHB
Intermediate behavioural representation, explicit trace, context-to-state propagation, experimental design, criteria fixed before the runs, baselines and falsification.
Session traces · adversarial audits · negative and partial results · documented limitations
Upstream Agents (NDA)
Context engineering, cognitive decomposition, external memory, named inputs, model routing, deterministic validation and human gates.
Pipeline DAG · agent contracts · templates · validators · architecture self-audit
04 · SCOPE AND TRAJECTORY
01 I have not yet trained or fine-tuned a model.
My hands-on work so far sits in model behaviour, inference and the harness. Training and post-training are my next frontier, not a boundary I intend to keep. I already work with datasets, failure classes and evaluation, which gives me a concrete bridge into training architecture, data mixtures, optimisation and reward design.
02 I implement and ship code, with AI and with teams.
I implemented and evolved an AI copilot in production, including its first end-to-end prototype, search engine, catalogue enrichment and application-level evals. I use AI-assisted engineering extensively and work with engineers, but I read, review, change and ship code. My claim is hands-on system direction and implementation, not lone-wolf authorship.
03 I separate knowledge from maturity.
Production, pilot and side-project evidence are not equivalent. I can explain which mechanisms have met real traffic, which are tested invariants and which remain research hypotheses.
05 · PUBLIC WORK
Public artifacts carry the technical thesis outside private systems and employer context.