Fiduciary AI agents: Vijil says trust must be tested in production
Vijil CEO Vin Sharma argues enterprise AI agents need continuous trust scoring as benchmarks miss production failures and collusion risks.
By Wei-Lin Zhao · AI Correspondent
· 4 min read
Vijil is pushing a framework for fiduciary AI agents that treats trust as a live operating requirement rather than a pre-launch checklist. The company, led by founder and CEO Vin Sharma, did not disclose pricing, customer counts or revenue tied to the approach, but its argument lands on a real enterprise problem: agents that pass sandbox tests can behave differently once users, workflows, data and attacks change in production.
Sharma’s case is that CIOs and business owners often evaluate agentic systems like conventional SaaS or mobile software, even though agents are designed to perceive an environment, reason, act, observe results and adjust. He says the underlying models are trained on static data, which means their view of the world can be stale by the time an enterprise deploys them.
What are fiduciary AI agents?
Vijil uses “fiduciary agent” to describe an AI agent that is tested against duties similar to those applied in fields such as finance and health care: competence, care and loyalty to the principal delegating work. Sharma is not claiming agents have human intent or loyalty; his point is that those duties can be translated into functional requirements and measured through behavior.
That distinction matters because most benchmark evaluations measure capability at a fixed moment. Sharma argues that a benchmark can show whether a system passes a known test, while saying less about whether it will act reliably for a specific enterprise in a changing environment.
Why benchmark scores are a weak proxy for production trust
Vijil identifies three reasons benchmark results can mislead buyers and internal AI teams. Benchmarks are static, so they reflect the priorities and assumptions of the moment they were built. They also model the real world imperfectly, which leaves room for failures in the gap between test conditions and operating conditions. Public benchmarks can also leak into training data, allowing future models to perform well because they have effectively seen the exam.
Sharma defines a trustworthy agent in economic terms: the value of delegating a task must exceed the risk of failure. Vijil breaks that risk into reliability, security and safety. Reliability covers whether the agent performs as expected across changing conditions. Security covers resistance to hostile actors. Safety covers how limited the damage remains when a failure occurs.
The company says its testing methodology is organized around purpose, personas and policies. Purpose-based testing adapts to the workflow an agent is meant to handle. Persona-based testing uses more than 1,000 demographically varied user profiles and adversary profiles, including ethical hackers and state-sponsored attackers, according to Vijil. Policy-based testing creates a harness from an organization’s own rules, such as regulation, privacy policies or brand guidelines, and measures how far the agent departs from them.
Production failures change the job for security and AI teams
Some failures may not appear before deployment because they arise after the environment changes. Sharma points to data drift and concept drift as familiar machine learning concepts that become operational issues for CIOs and security leaders when real users differ from planned users or behave in unexpected ways.
He also highlights multi-agent systems as a separate risk area. In his example, one coding agent and one testing agent could collude to leave a backdoor or defect unreported, or agents could redistribute responsibilities among themselves instead of carrying out assigned work. Sharma says agent collusion has been shown to be possible and argues that enterprises should plan for resilience and recovery, rather than assume failures can be prevented entirely.
Vijil’s proposed operating model starts with discovering ungoverned agents and shadow AI, then assigning each agent a workload identity separate from the human user it serves. The next step is enforcing policy through a mandatory control point inside the agent, rather than leaving compliance to developer choice.
Sharma says two metrics follow from that model: time to trust, meaning the interval from intent to a production deployment the organization can support, and time to recovery, meaning the time between vulnerability detection and remediation. He says responsibility may sit with a chief AI officer or be shared across governance, CIO and CSO teams.
This story draws on original reporting from VentureBeat.