The Vendor You Evaluated Last Year Is Not the Product You Have Now Licensed

Somewhere in your organisation, there is a slide deck from twelve months ago. It has a scorecard on it. Green, amber, a few reds that got negotiated down to ambers. A signature followed shortly after. The vendor cleared your process, the pilot cleared your committee, and the system went live.

That evaluation was rigorous. And it is now describing a product that no longer exists.
Picture two meetings where this usually surfaces, because it surfaces in two quite different ways.

In the first, for example, someone in Medical Affairs mentions, almost in passing, that the AI drafting tool you evaluated and licensed a year ago has started phrasing things differently – more confident, less hedged, citing sources it used to flag as uncertain. Someone else says the vendor added a feature last quarter that half the team has quietly started relying on, one that was never in the scope your validation package described. You knew this was an AI vendor. You just didn’t know the product you approved and the product now running are two different things.


In the second meeting, nobody uses the word “AI” at all, because nobody thinks to. Someone mentions that the CRM has “gotten smarter” as it’s now suggesting next-best-actions to the field team. Someone else says the document management system quietly started auto-summarising regulatory correspondence last month, and people have been using the summaries in briefing packs. Nobody evaluated either of these capabilities, because from a contracts standpoint, nothing was purchased. It was a version upgrade to software you vetted years ago, as software, for entirely different reasons, long before AI was part of the conversation.

Both meetings describe the same underlying failure. Your organisation evaluated a point in time and has been quietly relying on that evaluation ever since. But the two failures are not identical, and if you only build governance for one of them, you have solved half the problem while believing you solved all of it.

Two categories of vendor, one blind spot

Category one is the AI vendor you evaluated as an AI vendor. You did the work. You assessed the model, the data foundation, the evidence, the oversight model. You signed a licence for a specific capability, tested against specific claims that were true on a specific date. This is the vendor everyone assumes the “re-evaluation” conversation is about – and it is, but it’s only half the conversation.

Category two is the vendor you never evaluated as an AI vendor at all, because when you signed with them, they weren’t one. Your CRM, your ELN, your document management platform, your safety database, your clinical trial management system – the unglamorous, already-approved, already-embedded backbone of how your organisation runs. Most of these vendors have spent the last eighteen months bolting AI features onto their existing products, because every software company in the world is doing the same thing right now, for the same competitive reasons. The feature arrives in a routine release note, gets flipped on by default in many cases, and nobody in your organisation treats it as a new capability requiring evaluation — because to procurement, to IT, to the business unit using it, it doesn’t look like a new vendor. It looks like the same platform it always was, just updated.

That second category is the more dangerous blind spot, for a simple reason: your governance process for category one, however imperfect, at least exists. Someone, somewhere, decided this is an AI vendor and ran it through some version of an assessment. Category two never passes through that gate at all. There is no purchase order that triggers a review, because nothing was purchased. There is no vendor pitch to sit through, no pilot to scrutinise, no line item for legal to flag. The AI capability simply appears inside a system your organisation already trusts, inheriting that trust by default, with none of the scrutiny that trust was originally built on.

Why the quiet arrival is the harder problem

Think about what your original evaluation of that CRM, that document platform, that safety system actually tested. It tested data security, uptime, integration, support, contractual terms – everything appropriate for a deterministic tool doing a defined, predictable job. It did not test data foundation quality for AI reasoning, evidence lineage for AI-generated content, or a human oversight model for an automated suggestion – because none of that existed yet, and none of it was in scope.

Now that same platform is generating next-best-action suggestions, drafting summaries, or surfacing “insights” – often using training data, methodology, and grounding you know nothing about, running inside a piece of software your organisation already implicitly trusts. Your people extend that trust automatically, because the login screen looks the same as it did last year. Nobody had the conversation about whether this new capability, arriving through a routine update rather than a procurement decision, deserves the same scrutiny you’d apply to any other AI system touching pharmacovigilance, regulatory correspondence, or medical communications. In most organisations, that conversation simply never happens, because nothing about the update looked like an event that should trigger it.

This connects to something worth naming plainly: turning on an AI capability inside a platform you already own is a legitimate and often smart strategy – it’s frequently a better choice than buying or building something new, and it’s one of the most under-used correct answers in AI decision-making. The problem isn’t that embedded AI exists. The problem is the difference between choosing it deliberately, with eyes open, versus it choosing itself the moment a vendor flips a feature flag on and nobody in your organisation was in the room when that happened.

The category-one vendor moves too, just more visibly

The AI-native vendor you did evaluate is easier to catch, but only because you know to be looking, not because the drift is actually smaller.

Many AI vendors are not running a model they built and control end to end – they’re building on top of a foundation model from a frontier lab, sometimes several, and routing between them. When that underlying provider ships a new model version or deprecates an endpoint, the vendor’s product changes with it, sometimes overnight, whether or not the vendor’s own roadmap intended a change at all. What you validated the accuracy of in Q1 may be running on materially different reasoning underneath by Q3 – same interface, same brand, different engine.

The scope of automation quietly expands, too. A capability that started as “drafts a first pass for human review” can, release by release, absorb more of the judgement it was originally meant to support, not through a dramatic announcement, but through small usability improvements that each individually looked reasonable. Nobody signs off on “less human oversight than we approved.” It just accretes. And the commercial and security terms shift in the renewal paperwork rather than the product changelog so a sub-processor added, a pricing tier restructured now that you’re past pilot volume, a feature you rely on moved behind a new tier.

None of this is hidden in any sinister sense. It’s published in release notes, footnotes, and terms-of-service updates that are, by design, easy for a busy pharma team not to read line by line. The vendor isn’t obligated to walk you through what changed and what it means for the specific validated use case you built on top of their product nine months ago. That obligation, if it exists at all, sits with you — and it sits with you twice over, once for the vendor you knew was an AI vendor, and once for the vendor you never realised had become one.

Why “we’ll catch it if something goes wrong” isn’t a strategy for either

The instinctive fallback, for both categories, is monitoring for failure, if the outputs degrade or cause a problem, someone will notice, and that’s when re-evaluation happens. This is worse than it sounds, for two reasons specific to AI.

First, AI failure is often not loud. A model that has quietly become slightly less grounded, slightly more confident about things it shouldn’t be confident about, doesn’t announce itself with an error message. It produces plausible output that is wrong in ways that are genuinely difficult to catch without deliberately checking so a property that’s exactly as true of an AI feature buried inside your CRM as it is of a standalone AI product.

Second, by the time a failure is visible enough to trigger a review, it has usually already happened inside a workflow, possibly inside a GxP process, possibly inside something that left the building. Waiting for the failure to be the trigger means the re-evaluation is now an incident investigation, not a governance activity, and for the quietly-embedded category, the investigation starts from an even worse position, because there’s no evaluation file to go back to. There was never one written.

What a standing activity actually looks like for both

The fix isn’t a heavier initial evaluation, and it isn’t re-running a full RFP every quarter, that’s neither realistic nor proportionate. The fix is deciding, deliberately, that re-evaluation is a standing activity with a cadence and an owner, covering both categories of vendor, not just the one that arrived with a procurement trail attached.

A small number of things need to exist that, in most organisations right now, simply don’t.

An owner for each system with AI capability, including the ones that were never bought as AI, whose job includes noticing what changed, not just whether the invoice got paid. Right now, nobody holds this for the embedded category in particular, because nobody has been assigned a system that, on paper, hasn’t changed.

A standing inventory that gets re-checked, not just built once, so a live list of where AI capability exists across your technology estate, including inside platforms bought for entirely different reasons, refreshed on a cadence rather than compiled as a one-time exercise and then left to go stale the same way the vendors do.

A defined trigger set that pulls a re-evaluation forward regardless of the renewal date be it a material model change, a new sub-processor, a feature flag flipped on inside a platform you already run, a shift in what a system is now being used for versus what it was ever assessed for. Some of these are visible in release notes if someone is actually reading them; some require asking directly, on a cadence, rather than waiting to be told.

A lighter-weight review instrument for these check-ins than the one used at initial signature so re-testing the specific claims and evidence that changed, against the same standard you’d apply to a new AI vendor, so that “still fits our governance model” has a current answer rather than an inherited one, whether the capability arrived through a purchase order or a changelog.

Even a well-run cadence has a blind spot worth naming honestly: the gap between checks. If a routine release changes something material on a Tuesday and the next scheduled review isn’t until Friday, the system keeps acting on Monday’s authorisation for three days in between. A standing activity closes the “we’d never find out” problem. On its own, it doesn’t close the “we found out a few days late” problem, and for a narrow slice of your portfolio, a few days is exactly what matters. For that narrow slice, the cadence needs a second layer underneath it: not a periodic look back, but a check at the point of use, on the specific action about to happen, confirming the conditions it was authorised under still hold before it’s allowed to reach a patient, a regulator, or a decision. Most of what you run can live safely on a cadence alone. The handful of uses where a few days of undetected drift is genuinely unacceptable need that second gate as well.

The intensity of all of this should not be uniform across your portfolio. A CRM suggestion feature used for internal prioritisation does not need the same standing scrutiny as an embedded capability touching pharmacovigilance signal triage or feeding a regulatory submission. Building one cadence and applying it everywhere is almost as wasteful as having no cadence at all; the organisations that do this well have already worked out where the asymmetric risk sits in their portfolio, across both categories of vendor, and spend their attention there first.

The part that’s uncomfortable 

If you did a genuinely good evaluation a year ago, the instinct is to treat it as done. And for the vendor you knew was an AI vendor, that instinct is at least defensible, as the evidence was real, for the product as it existed then. For the vendor that quietly became an AI vendor sometime after you signed, there is no instinct to correct, because there was never a decision point to begin with. The capability simply arrived, wearing the credibility of a system your organisation already trusted.

That’s not a criticism of how your team runs procurement. It’s a recognition that AI has changed what “software you already own” means, and most governance processes weren’t built to notice a category of vendor that changes underneath you without ever presenting itself as one.

The organisations getting this right aren’t the ones running heavier initial evaluations on the vendors they know to scrutinise. They’re the ones who’ve stopped assuming that “we already vetted this platform” means anything at all about the AI capability it shipped last month, and who treat both categories of drift as the same standing obligation, owned, triggered by defined events, sitting alongside every vendor relationship for as long as the system stays in use.

 

Where this goes next

How to build a live inventory that actually catches embedded AI arriving through routine updates, what the trigger list should contain for each category, who should own it, and how to fold both kinds of drift into your existing validation and change-control machinery without turning it into a second full-time job for someone on your team, that’s exactly the kind of thing that goes stale if you learn it once and file it away. It’s also exactly what this topic looks like when it’s built out into a working framework rather than a set of principles.

That’s now its own course called After Signature: Standing Re-Evaluation for AI Vendors and Embedded AI which is publishing this month in the Foundations Track, alongside Evaluating, Buying & Building AI in the Pharma AI Enablement Institute. It’s one of a growing number of courses in the Institute, each built around a single critical issue pharma AI leaders are actually facing right now, kept current with monthly updates as the vendor landscape itself keeps moving, and a live monthly Q&A where you can bring your own team’s latest changelog and ask what it actually means for a system you already thought you understood.

Because the uncomfortable truth in this piece applies to your training on this topic too: learn it once, and it starts going stale the same way both kinds of vendor do.

Somewhere in your building, that twelve-month-old scorecard is still sitting in a folder, still treated as settled. That’s worth five minutes of discomfort before it’s worth anything else.

The live inventory template, the trigger list for both vendor categories, who should actually own it, and how it folds into validation and change-control without becoming someone’s second job, that’s what’s inside the course itself.

Book a 15-minute call to find out more about the AI Enablement Institute.

Contact Us

Write you name and email and enquiry and we will get right back to you as soon as we can.