Software engineer with 15+ years of experience — on building the governance and evaluation infrastructure that turns AI adoption from a compliance risk into a measurable return.
Verizon's 2026 Data Breach Investigations Report has flagged a governance gap that's growing fast: employee use of unapproved "shadow AI" tools tripled in a single year, from 15% to 45% of the workforce. The report marks shadow AI as an elevated — and still-expanding — data-exfiltration risk. Employees get real productivity gains from AI, adoption spreads faster than anyone can track, and nobody can say with confidence what data went where or what it's costing. For a company running hundreds of engineers, that's an open door — and we wanted to speak with a professional who has already handled closing it.
Vissarion Chakvetadze, a software engineer with 15+ years of experience, specializes in designing AI-powered systems that help organizations address exactly this kind of problem. His work includes architecting LLM-powered agents, retrieval-augmented generation (RAG) pipelines, and reporting systems, backed by Langfuse observability and DeepEval's LLM-as-a-judge framework, so that AI output is checked by a second model rather than trusted blindly.
We spoke with him about building AI infrastructure that scales without losing control of data, cost, or trust.
Shadow AI use has swept the industry, tripling in a single year. How does this actually play out inside a real engineering organization?
It looked like a lot of companies that grew fast without planning for this: a few people using AI tools on their own, no shared process. That's fine for five people. It's dangerous for large-scale organizations, because the risks don't surface right away — someone pastes something they shouldn't into a chatbot, and you might not find out for months. When you work with sensitive user data, sending private information to a third-party model isn't a minor risk but a hard constraint. That makes data privacy a fundamental requirement when building AI systems.
Many companies think giving everyone an unlimited ChatGPT subscription is enough to adopt AI at scale. Why does that break down once you scale past a handful of engineers?
Visibility breaks first. Once you give a large engineering group unrestricted access to a general-purpose model, you lose the ability to answer basic questions, such as what data is going out, what's coming back, and whether it's actually correct. That can't be checked manually at scale, so the first thing that fails is trust in the output itself.
Budget is the second problem. An open subscription has no sense of what's worth spending on. People default to the most capable model for tasks a cheaper one would handle fine, and nobody notices until the cost adds up. And the third is consistency. Every engineer develops their own way of prompting, their own workarounds. That's harmless for personal use, but if we're shipping something to real users, we need to know it behaves the same way every time, not just that it worked once.
Production AI systems raise three recurring challenges: visibility, cost control, and consistency. What does the infrastructure for addressing them consist of, and what was the hardest problem to solve?
If you want to deploy AI systems at scale, you need more than the models themselves; you need the right infrastructure around them. That infrastructure typically consists of three key layers. Retrieval connects agents to internal knowledge. Centralized prompt management gives engineers access to shared, tested prompts instead of starting from scratch. And observability provides visibility into how the system is performing, helping teams understand what's happening across the entire AI stack rather than treating it as a black box.
The hardest part, however, is evaluation. Prompt management and retrieval can largely be addressed through architecture and iteration, but evaluating AI-generated language is fundamentally different. You're not just checking whether an answer is right or wrong — two responses can both appear correct while differing in ways that materially affect their quality or usefulness. This is where LLM-as-a-judge approaches become particularly valuable: a second model can evaluate the first, while observability tooling traces what was generated and provides insight into how the result was produced. But this introduces another layer of complexity, because the judge model has its own biases and blind spots.
Ultimately, the real challenge is building the evaluation and confidence mechanisms needed to know when those outputs can actually be trusted at scale — which matters far more than getting models to generate good outputs in the first place.
You've also judged international hackathons, including the TechGenius Hackathon hosted by Uzbekistan's Ministry of Higher Education. What does evaluating hackathon demos teach you about evaluating production AI systems?
You can always polish a demo, a pitch, or a set of slides, but that's not what people ultimately need. A polished presentation can look impressive; the real test is how the system handles edge cases and behaves outside the exact scenario it was designed or tested for. That's the same discipline I rely on when evaluating our own AI agents: a working demo isn't the same as a working system.
That said, the format doesn't transfer. As a judge, I'm making a subjective call once, against a handful of submissions, with time to think it through. A production evaluation pipeline has to make that same kind of judgment automatically, consistently, and at a volume no human could keep up with. The instinct for what "good" looks like is the same, but turning that instinct into something a second model can apply reliably every time is a completely different engineering problem.
What does AI ROI actually look like when you're accountable for both the engineering and the budget?
ROI here isn't a single number. It's a ratio that has to hold across every part of the system, not just the impressive parts. A model that produces great output but costs more than the value it creates is a liability with good marketing. That means tracking cost per task alongside quality per task, while being deliberate about which model handles which job. The most capable model isn't automatically the right choice if a cheaper one gets the same result.
Compliance is less about ROI and more about a hard boundary that shapes everything else. When you can't send private data to third-party models, it limits which tools are even on the table before cost enters the conversation. And since it may be impossible to manually check output at large scale, verification has to be built into the cost model too. An evaluation layer that catches bad output before it reaches a user is what actually makes the cheaper, faster option safe to use in the first place. So when I'm accountable for both sides, the real measure of success isn't how much the AI initiative saves on paper. It's whether we can trust the numbers we're reporting, because we've verified them systematically rather than assumed they're right.
Where do you see enterprise AI governance heading next, and what's the one lesson other engineering leaders should take from it?
Governance is going to stop being optional and start being table stakes, the same way security or CI/CD did. Companies that treat AI as something individuals figure out on their own are going to keep showing up in reports like the one that prompted this conversation. The ones that treat it as infrastructure, with real ownership and real verification, are going to be the ones actually getting value out of it a year from now.
If there's one lesson, it's patience. If you rush this to claim an early win, you end up with something nobody trusts and everybody quietly works around. I'd rather take the time to build something the team actually relies on than announce success before it's real.


