AI Tools for Banks

Buyer's guide

How to evaluate an AI vendor as a community bank or credit union

By the AI Tools for Banks editorial team · Last verified

On this page

Short answer

Start by asking for a named institution your size, on your core, running the product in production today. Then separate shipped features from roadmap dates, find out who owns the integration for the next five years, price the internal governance cost, and get a written not-to-exceed figure before you spend staff time on a pilot.

Most AI evaluations at a community institution fail in the same way. Six weeks go into comparing feature grids, and the thing that decides the outcome, whether anyone your size has ever run this successfully, never gets asked directly. This is the sequence that surfaces that answer first, when it is still cheap to walk away.

Ask the size question before the feature questions

The single most useful sentence in this process is: name a bank or credit union under $2 billion in assets, on our core, running this in production today, and let us call them. A surprising number of well-known vendors cannot complete that sentence.

An inability to answer is information rather than a disqualification. It tells you that you will be an early named institution of your size, which is a legitimate position to take if the product solves a real problem and you price the risk accordingly. What it should change is the contract: shorter initial term, delivery-linked payment, and an exit that does not require a board resolution.

  • Institution name and asset size, not a logo wall
  • Their core, and whether the integration was marketplace-delivered or custom
  • Go-live date, and how long implementation actually took against the estimate
  • What they measured after twelve months, in their own words rather than the vendor's case study

Separate what ships from what was announced

Several products in this market have named AI features with future general availability dates, and the marketing rarely distinguishes them from what is running today. Ask directly which features are generally available, which are in limited release, and which have a date attached.

Then evaluate on the shipped set only. Roadmap can inform the decision between two otherwise equal vendors; it should never be the reason for the decision, because you will be running the current product for at least the first year and possibly the whole initial term.

Trace where the integration cost lands, and who carries it

The difference between a product listed in your core provider's marketplace and one that needs a custom API integration is usually larger than any difference in licence fee. Marketplace and OEM arrangements carry the integration through core upgrades as part of the deal. A custom integration is yours to maintain, and the bill arrives the first time your core changes a field.

Ask what data the product needs, at what frequency, who builds the feed, who fixes it when it breaks, and what happened at a reference institution the last time the core was upgraded. The last question produces the most honest answers.

Price the internal cost, not just the licence

Anything that scores, decides or drafts becomes something your risk function documents and reviews, on a schedule, forever. That is staff time you are committing alongside the subscription.

The vendors that reduce this cost are the ones producing documentation as a by-product of how the product works: fair-lending artifacts generated during model construction, cited answers that carry their source, traceability from a calculated figure back to its document. A product that hands you a number and a user manual leaves the write-up to your compliance officer, which is the most expensive form of cheap software.

  • Model documentation and annual validation effort
  • Integration maintenance across core and LOS upgrades
  • Content curation, which is what stretches conversational AI timelines
  • The human review queue the product creates, and who staffs it

Get a number before you spend staff time

Almost nobody in this market publishes pricing. That makes the evaluation itself expensive, because the only way to learn whether a product is affordable is to run most of a sales cycle.

Ask for a written not-to-exceed figure early, based on your asset size and volume. A vendor that will not give one before a full discovery process is telling you something about how the relationship will run, and that is worth knowing in week two rather than week ten.

A short due-diligence checklist

Run this before the second demo, not after the third.

  • Named reference at your asset size, on your core, that you can actually call
  • Written list of generally available features against roadmap features with dates
  • Security posture: SOC 2 Type II or equivalent, and where the data physically sits
  • Who touches your data, including whether human reviewers are vendor staff or subcontracted
  • Whether your documents or data train the vendor's models, in the contract rather than the FAQ
  • What documentation the product produces for your examiner without extra work
  • A not-to-exceed price and an initial term you can survive if it does not work

Frequently asked questions

Should a community bank be the first institution its size on a product?

Sometimes, and it is a legitimate choice when the problem is real and the vendor is honest about the gap. Price it accordingly: shorter term, payment tied to delivery, and a defined exit. What you should not do is buy on the assumption that peer references exist when nobody has produced one.

How much should we budget for AI software?

There is no useful benchmark, because almost nothing in this market is priced publicly. Two products publish anything usable, and both are horizontal rather than banking-specific. Build the evaluation calendar around getting quotes early rather than late.

Who should own the evaluation internally?

The person who owns the process being fixed, with risk and information security engaged from the first demo rather than at contract stage. Late security review is the most common reason these projects slip a quarter.

What is the most common evaluation mistake?

Comparing products that do not compete. Origination platforms, decisioning layers and document analysis tools all get described as AI lending software, and an institution that puts all three on one feature grid will spend a quarter learning that it was three separate decisions.