Why we didn't wait: A CEO's field notes from two years of applied AI

AI value is compounding, not linear. BTS CEO Jessica Skon shares how experimentation fuels flywheels, and how breakthrough “AI diamonds” emerge and scale.
April 29, 2026
5
min read
Subscribe to the BTS newsletter
Follow us on Linkedin
Follow BTS on Linkedin
Share

Three decisions that changed everything.

Two years ago, we made three deliberate decisions about how BTS would move with Applied AI.

We would become our own Customer Zero.

While others were building strategies, defining governance, and waiting for clarity, we made a different call: we decided not to wait. Not because the stakes were low, but because they were high. And because in a space evolving this quickly, clarity wouldn’t come from planning. It would come from movement.

So instead of starting with a roadmap, we started with three principles:

  1. No top-down mandate. The people closest to the work figure it out.
  2. IT must evolve from gatekeeper to enabler - leading AI trials and fast experimentation.
  3. Don’t wait for certainty.

We set the organization in motion, and once we did, things started to move quickly.

What if we started this company today?

Waiting for certainty is itself a choice, and it’s costing companies more than they realize.

We started where we knew the work best: our simulations. No perfect plan, just teams moving, trying, and iterating.

Simulations are core to who we are at BTS. Companies that simulate don’t just make better decisions; they execute faster and build more engaged cultures.

The team asked a simple question:

"What if we were to start our company today?”

That question started the flywheel.

They asked IT for a few licenses and started building - vibe-coding, writing agents, and testing tools - moving at a pace that would make any VC-backed start-up smile.

The messy middle.

At first, the team was underwhelmed.

The early reports were blunt:

“Not good with math.”
“Poor graph capabilities.”

The team wasn't discouraged. They kept tinkering - jumping between tools, staying on top of new releases, experimenting constantly.

This was a small team, across 24 countries, building off each other’s ideas. Laughing at crazy creations. Breaking things. Iterating in a sandbox alongside real clientwork.

Each cycle produced something:

  • A sharper scenario
  • A faster build
  • A more powerful simulation

The flywheel was turning, and it was generating something real.

When the diamond appeared.

Then something shifted.

The team moved into client trials across five countries. They figured out ISO compliance and built the architecture to handle the complexity, the “spaghetti.”

And what emerged wasn’t incremental:

  • What used to take weeks started happening in days.
  • Limited creativity started to feel like unlimited innovation.
  • Clients became self-serving.
  • Agentic simulations were built directly into client systems for real-time updates and preparation.

This was our first AI diamond - a high-impact outcome created by many cycles of experimentation compounding into real value.

It only appeared because we kept the flywheel turning, each cycle increasing the odds that something would break through.

95% adoption in eight weeks.

Then it was time to take the AI diamond global.

BTS is decentralized and highly entrepreneurial. We operate across 24 countries and 38 offices, where local teams have real autonomy.

And historically? That’s meant a low appetite for adopting something built somewhere else and pushed from the center.

So we expected resistance.

Instead, something surprising happened.

In the first eight weeks, we saw 95% adoption across our global footprint.

It felt completely different from our own digital initiatives, ERP implementations, top-down rollouts of the past.

This moved on its own. Why? 

We realized it didn’t start with a framework or a model, it started with a feeling.

The feeling of being at the leading edge of one’s craft and profession.

  • Joy
  • Excitement
  • Pride

As we watched this play out across teams it stopped feeling like isolated wins.

There was a pattern to it. A repeatable, organic, innovation motion.

And the flywheel didn’t stop with simulations.

It spread across finance, sales enablement, legal, operations, and client delivery. Some cycles led to small improvements, and others revealed new diamonds.

Not becausewe planned for them, but because we built the conditions for people to find them.

The question I'd ask any CEO right now: Is your flywheel turning, or are you still waiting for the perfect plan?

In part 2, I’ll share the key success factors behind the breakthrough, and what we’re now seeing across more than 120 global clients.

Applied AI FAQs

Learn how to design conversations that actually move decisions forward.
Download the report

Related content

Blog
August 19, 2026
5
min read
Everybody's planning an AI reset off-site. Four mistakes will sink most of them.
Planning an AI reset off-site? The agenda decides everything. Four common design mistakes, and how to build two days that change what your company is capable of.

There’s a specific kind of strategy meeting getting scheduled right now, in nice hotels with bad coffee: the AI reset off-site.

And for good reason. In a 2026 WRITER survey, 48% of leaders described their AI rollout as, in their own words, a "massive disappointment." That's nearly half the room.

What that number really measures is the distance between what these tools can do and what people are doing with them. In our experience, that distance is almost entirely human.

Which is why the off-site is the right instinct. Making the most of that time is the harder part.

What separates an AI reset that actually changes the game from an expensive two-day conversation? In our experience, it comes down to avoiding four common design mistakes.

Mistake 1

Blaming the bots

The gap between AI investment and real adoption is almost always about people, not technology. And when adoption stalls, we usually find it's one of four things.

  1. They don't think it will help them
  2. Nobody around them is using it
  3. They don't feel capable
  4. Or they don't have real access to the tools they were promised

Four different problems, and four completely different fixes.

That's why diagnosis comes first. If you don't know which barrier you're dealing with, every intervention becomes an educated guess. And you cannot tell which one you have by staring at a dashboard. A belief gap and a skill gap look identical in a status report and need opposite interventions. Show up guessing, and you'll spend real money teaching people to use a tool they simply don't trust yet. Congratulations - you've just catered the wrong conversation.

Mistake 2

Letting leaders off the hook

One of the biggest predictors of whether change sticks is also one of the most overlooked: leadership.

If your executives show up as observers, nodding along and quietly answering email under the table, your people clock it in about four minutes.

That doesn't mean your CEO has to emcee the thing. It means they use the tools in front of everyone, participate in the conversation, and make it clear this isn't someone else's initiative.

Recently we’ve been working with a Fortune 200 global professional services firm who’s top 120 leaders were at very different points with AI. Some were redesigning entire processes. Others were using it to summarize emails, or not at all. Rather than focus on the technology, the four-hour session focused on what leaders could do with AI, applying it to a live strategic challenge and ending with a personal commitment to lead differently. The response was strong enough that the organization is now cascading the experience globally.

The lesson is simple: when leaders experience AI as a strategic capability, they're better equipped to model the behavior that makes adoption stick. Nothing you build during those two days survives without that entire chain of leadership doing its part.

Mistake 3

Chasing the wrong outcome

Without a behavioral baseline, you have no way to prove anything actually moved. No baseline, no ROI. You're just hoping the energy in the room was good, which is a wonderful feeling and a terrible metric to bring to your CFO.

But the baseline isn't just about proving the off-site worked. It's about understanding where you're starting in the first place. And you'll want that clarity, because the quiet resistance is real. In that same 2026 research, nearly a third of employees admitted to actively working around their company's AI strategy. If you don't win their belief in the room, some of them will keep politely ignoring the whole thing from their desks. You can't measure your way out of that. You have to earn your way out of it.

Which brings us to the biggest reframe of all.

Mistake 4

Leaving follow-through to chance

We've been working with a Fortune 100 medical device company on their AI strategy for three years. It started with their leadership team, a three-hour session built around what those leaders would do differently, and it landed. What became clear afterward was that the same experience needed to happen everywhere else. So, it expanded: 90-minute activations for 15,000 people, and this year intact teams redesigning their own workflows.

Three years in, that first session is the smallest part of the story.

Your event is where momentum gets created. What happens at 30, 60, and 90 days is where results get made.

If you're planning one of these and want to change what happens on Monday, not just how everyone feels on Friday, that the work we do.
We'd be glad to help you design it.
Blog
August 14, 2026
5
min read
Every candidate looks like a great hire now. AI made sure of it.
Polish is no longer a hiring signal. See how organizations use role-relevant simulations and predictive validity data to hire for high-stakes roles.

Candidates now arrive at interviews pre-coached by AI, with their resumes optimized to pass every checkpoint. Polish has stopped being a signal. The traditional hiring process was built to read exactly the cues that AI is now best at producing, and the signals hiring managers once relied on have weakened as a result. And for roles where the wrong hire carries real business consequences, losing the ability to tell who will actually perform is not a minor inconvenience. It is a material risk, and it exposes the business to unnecessary turnover, reduced performance, and heavier investment for talent growth and development.

So how do you observe the behaviors that matter most, before someone is in the role?

Not by asking better questions, but rather by putting candidates in situations designed to elicit that behavior.

The limits of predicting from paper

Credentials tell you what someone has done. Structured interviews tell you what someone says they would do. Neither lets you observe what they actually do in the moments that count.

This distinction matters most in client-facing, relationship-driven roles, where the performance gap between a strong hire and a weak one plays out in real business outcomes (revenue, retention, client growth). Organizations that hire at scale in these roles carry that gap across hundreds of decisions at a time.

The better approach is to watch candidates do the work before you hire them. Put them in simulated, role-relevant scenarios, and pair the simulation with a second, different kind of measure so no single method carries the whole decision. That combination is what lets you evaluate real performance before anyone is in the role. Organization-specific simulations provide a clear read on who is ready and capable of performing on day one. In a world of AI-supported candidate signals, the use of simulations makes the process harder to prep for. It is harder to fake. And, when designed well, it is substantially more predictive than other hiring methods.  

What counts as evidence

Claims about predictive power are easy to make. Evidence for them is rarer than you would expect.

A predictive validity study, the kind that links pre-hire assessment scores to how someone actually performs once hired, is some of the hardest evidence to produce and the rarest to see. Many assessments are validated against proxies: another test, or a theoretical model of the role, rather than real results on the job. Connecting scores to concrete business outcomes and doing the statistical work to show the link holds, takes years of shared data and a level of commitment from both the assessment provider and the client that most partnerships never reach. That is precisely why it is worth asking for. A provider who can show how assessment scores track to training completion, retention, and first-year output is offering something categorically different from one who can only show a correlation with another test.

Why simulation holds up where other methods do not

When a candidate sits across from a trained assessor (someone playing the client or prospect on the other side of the conversation) and has to work through a real situation, they cannot rely on a rehearsed answer. The scenario is specific. The stakes feel real. What you see is close to what you would get on the job.

That is the value of simulation-based assessment: it does not test what candidates know about the role.

It shows how they use what they know when a real person is on the other side of the conversation, before the stakes are real.

For roles that carry significant business responsibility, this distinction is the whole game. The cost of the wrong hire in a high-stakes client-facing role is not just a missed quota for a quarter - It plays out in relationships that do not develop, clients who leave, and productivity losses that compound over time. Getting those hiring decisions right, at scale, with consistency, requires methods that are built for predictive accuracy, not just candidate experience or hiring speed.

What this means for how organizations think about hiring

Most organizations are still optimizing the wrong things in their hiring process. They invest heavily in employer branding, application flow, and interview structure, all of which matter, but less in the core question: does our hiring process actually predict who will succeed in this role?

AI has sharpened the stakes here. If every candidate can present as polished and prepared, screening based on presentation becomes less useful. What holds up is direct observation of the behaviors that the job requires.

A few principles worth building from:

  • Measure what the job requires, not what is easy to measure. Cognitive tests and personality questionnaires have their place, but they do not look much like the job. The closer the assessment is to the actual work, the better it predicts performance in it.
  • Ask what your assessment predicts. Training completion? Retention? First-year output? Most organizations cannot answer that question today, largely because providers have rarely been asked to prove it. It is a fair thing to ask for.
  • Take the human element seriously. In a simulation, a candidate is having a real conversation, responding in real time, navigating a situation that requires judgment. Even with the help of AI, that is hard to game. And it remains one of the strongest predictors of on-the-job performance available.

The data exists to make hiring decisions more accurate, fairer, and more directly tied to business outcomes. For organizations operating in high-stakes roles at scale, there is too much on the line to rely on methods that cannot hold up to that standard.

You may be interested in BTS’ thought leadership in the five talent shifts AI is forcing now.  

Blog
July 31, 2026
5
min read
El GPS no maneja el auto. La IA cambió el mapa, no el viaje…(ES)
La IA ya no es una ventaja competitiva en ventas. Descubre por qué el verdadero diferencial está en el criterio comercial, el conocimiento del negocio y la capacidad de construir relaciones de confianza.

La IA ya forma parte del día a día de las ventas. Hoy cualquier asesor puede llegar a una reunión con datos, tendencias e insights generados en segundos. Sin embargo, disponer de más información no garantiza conversaciones de mayor valor.

A través de una experiencia real con un consultor comercial, este artículo explica por qué la inteligencia artificial funciona como un GPS: ayuda a interpretar el entorno, pero no conduce la conversación ni entiende las prioridades del cliente.

En este artículo descubrirás:

  • Por qué el acceso a la información ya no supone una ventaja competitiva.
  • La importancia del business acumen para interpretar los datos con criterio.
  • Cómo hablar el lenguaje del cliente genera credibilidad y diferenciación.
  • Por qué las relaciones B2B evolucionan hacia relaciones P2P basadas en la confianza.
  • Qué capacidades consultivas seguirán siendo exclusivamente humanas incluso en la era de la IA.

La tecnología seguirá evolucionando, pero la ventaja competitiva estará en quienes sean capaces de combinar inteligencia artificial con conversaciones centradas en el cliente, pensamiento estratégico y relaciones de largo plazo.

Related content

Blog
August 19, 2026
5
min read
Everybody's planning an AI reset off-site. Four mistakes will sink most of them.
Planning an AI reset off-site? The agenda decides everything. Four common design mistakes, and how to build two days that change what your company is capable of.

There’s a specific kind of strategy meeting getting scheduled right now, in nice hotels with bad coffee: the AI reset off-site.

And for good reason. In a 2026 WRITER survey, 48% of leaders described their AI rollout as, in their own words, a "massive disappointment." That's nearly half the room.

What that number really measures is the distance between what these tools can do and what people are doing with them. In our experience, that distance is almost entirely human.

Which is why the off-site is the right instinct. Making the most of that time is the harder part.

What separates an AI reset that actually changes the game from an expensive two-day conversation? In our experience, it comes down to avoiding four common design mistakes.

Mistake 1

Blaming the bots

The gap between AI investment and real adoption is almost always about people, not technology. And when adoption stalls, we usually find it's one of four things.

  1. They don't think it will help them
  2. Nobody around them is using it
  3. They don't feel capable
  4. Or they don't have real access to the tools they were promised

Four different problems, and four completely different fixes.

That's why diagnosis comes first. If you don't know which barrier you're dealing with, every intervention becomes an educated guess. And you cannot tell which one you have by staring at a dashboard. A belief gap and a skill gap look identical in a status report and need opposite interventions. Show up guessing, and you'll spend real money teaching people to use a tool they simply don't trust yet. Congratulations - you've just catered the wrong conversation.

Mistake 2

Letting leaders off the hook

One of the biggest predictors of whether change sticks is also one of the most overlooked: leadership.

If your executives show up as observers, nodding along and quietly answering email under the table, your people clock it in about four minutes.

That doesn't mean your CEO has to emcee the thing. It means they use the tools in front of everyone, participate in the conversation, and make it clear this isn't someone else's initiative.

Recently we’ve been working with a Fortune 200 global professional services firm who’s top 120 leaders were at very different points with AI. Some were redesigning entire processes. Others were using it to summarize emails, or not at all. Rather than focus on the technology, the four-hour session focused on what leaders could do with AI, applying it to a live strategic challenge and ending with a personal commitment to lead differently. The response was strong enough that the organization is now cascading the experience globally.

The lesson is simple: when leaders experience AI as a strategic capability, they're better equipped to model the behavior that makes adoption stick. Nothing you build during those two days survives without that entire chain of leadership doing its part.

Mistake 3

Chasing the wrong outcome

Without a behavioral baseline, you have no way to prove anything actually moved. No baseline, no ROI. You're just hoping the energy in the room was good, which is a wonderful feeling and a terrible metric to bring to your CFO.

But the baseline isn't just about proving the off-site worked. It's about understanding where you're starting in the first place. And you'll want that clarity, because the quiet resistance is real. In that same 2026 research, nearly a third of employees admitted to actively working around their company's AI strategy. If you don't win their belief in the room, some of them will keep politely ignoring the whole thing from their desks. You can't measure your way out of that. You have to earn your way out of it.

Which brings us to the biggest reframe of all.

Mistake 4

Leaving follow-through to chance

We've been working with a Fortune 100 medical device company on their AI strategy for three years. It started with their leadership team, a three-hour session built around what those leaders would do differently, and it landed. What became clear afterward was that the same experience needed to happen everywhere else. So, it expanded: 90-minute activations for 15,000 people, and this year intact teams redesigning their own workflows.

Three years in, that first session is the smallest part of the story.

Your event is where momentum gets created. What happens at 30, 60, and 90 days is where results get made.

If you're planning one of these and want to change what happens on Monday, not just how everyone feels on Friday, that the work we do.
We'd be glad to help you design it.
Blog
August 14, 2026
5
min read
Every candidate looks like a great hire now. AI made sure of it.
Polish is no longer a hiring signal. See how organizations use role-relevant simulations and predictive validity data to hire for high-stakes roles.

Candidates now arrive at interviews pre-coached by AI, with their resumes optimized to pass every checkpoint. Polish has stopped being a signal. The traditional hiring process was built to read exactly the cues that AI is now best at producing, and the signals hiring managers once relied on have weakened as a result. And for roles where the wrong hire carries real business consequences, losing the ability to tell who will actually perform is not a minor inconvenience. It is a material risk, and it exposes the business to unnecessary turnover, reduced performance, and heavier investment for talent growth and development.

So how do you observe the behaviors that matter most, before someone is in the role?

Not by asking better questions, but rather by putting candidates in situations designed to elicit that behavior.

The limits of predicting from paper

Credentials tell you what someone has done. Structured interviews tell you what someone says they would do. Neither lets you observe what they actually do in the moments that count.

This distinction matters most in client-facing, relationship-driven roles, where the performance gap between a strong hire and a weak one plays out in real business outcomes (revenue, retention, client growth). Organizations that hire at scale in these roles carry that gap across hundreds of decisions at a time.

The better approach is to watch candidates do the work before you hire them. Put them in simulated, role-relevant scenarios, and pair the simulation with a second, different kind of measure so no single method carries the whole decision. That combination is what lets you evaluate real performance before anyone is in the role. Organization-specific simulations provide a clear read on who is ready and capable of performing on day one. In a world of AI-supported candidate signals, the use of simulations makes the process harder to prep for. It is harder to fake. And, when designed well, it is substantially more predictive than other hiring methods.  

What counts as evidence

Claims about predictive power are easy to make. Evidence for them is rarer than you would expect.

A predictive validity study, the kind that links pre-hire assessment scores to how someone actually performs once hired, is some of the hardest evidence to produce and the rarest to see. Many assessments are validated against proxies: another test, or a theoretical model of the role, rather than real results on the job. Connecting scores to concrete business outcomes and doing the statistical work to show the link holds, takes years of shared data and a level of commitment from both the assessment provider and the client that most partnerships never reach. That is precisely why it is worth asking for. A provider who can show how assessment scores track to training completion, retention, and first-year output is offering something categorically different from one who can only show a correlation with another test.

Why simulation holds up where other methods do not

When a candidate sits across from a trained assessor (someone playing the client or prospect on the other side of the conversation) and has to work through a real situation, they cannot rely on a rehearsed answer. The scenario is specific. The stakes feel real. What you see is close to what you would get on the job.

That is the value of simulation-based assessment: it does not test what candidates know about the role.

It shows how they use what they know when a real person is on the other side of the conversation, before the stakes are real.

For roles that carry significant business responsibility, this distinction is the whole game. The cost of the wrong hire in a high-stakes client-facing role is not just a missed quota for a quarter - It plays out in relationships that do not develop, clients who leave, and productivity losses that compound over time. Getting those hiring decisions right, at scale, with consistency, requires methods that are built for predictive accuracy, not just candidate experience or hiring speed.

What this means for how organizations think about hiring

Most organizations are still optimizing the wrong things in their hiring process. They invest heavily in employer branding, application flow, and interview structure, all of which matter, but less in the core question: does our hiring process actually predict who will succeed in this role?

AI has sharpened the stakes here. If every candidate can present as polished and prepared, screening based on presentation becomes less useful. What holds up is direct observation of the behaviors that the job requires.

A few principles worth building from:

  • Measure what the job requires, not what is easy to measure. Cognitive tests and personality questionnaires have their place, but they do not look much like the job. The closer the assessment is to the actual work, the better it predicts performance in it.
  • Ask what your assessment predicts. Training completion? Retention? First-year output? Most organizations cannot answer that question today, largely because providers have rarely been asked to prove it. It is a fair thing to ask for.
  • Take the human element seriously. In a simulation, a candidate is having a real conversation, responding in real time, navigating a situation that requires judgment. Even with the help of AI, that is hard to game. And it remains one of the strongest predictors of on-the-job performance available.

The data exists to make hiring decisions more accurate, fairer, and more directly tied to business outcomes. For organizations operating in high-stakes roles at scale, there is too much on the line to rely on methods that cannot hold up to that standard.

You may be interested in BTS’ thought leadership in the five talent shifts AI is forcing now.  

Blog
July 31, 2026
5
min read
El GPS no maneja el auto. La IA cambió el mapa, no el viaje…(ES)
La IA ya no es una ventaja competitiva en ventas. Descubre por qué el verdadero diferencial está en el criterio comercial, el conocimiento del negocio y la capacidad de construir relaciones de confianza.

La IA ya forma parte del día a día de las ventas. Hoy cualquier asesor puede llegar a una reunión con datos, tendencias e insights generados en segundos. Sin embargo, disponer de más información no garantiza conversaciones de mayor valor.

A través de una experiencia real con un consultor comercial, este artículo explica por qué la inteligencia artificial funciona como un GPS: ayuda a interpretar el entorno, pero no conduce la conversación ni entiende las prioridades del cliente.

En este artículo descubrirás:

  • Por qué el acceso a la información ya no supone una ventaja competitiva.
  • La importancia del business acumen para interpretar los datos con criterio.
  • Cómo hablar el lenguaje del cliente genera credibilidad y diferenciación.
  • Por qué las relaciones B2B evolucionan hacia relaciones P2P basadas en la confianza.
  • Qué capacidades consultivas seguirán siendo exclusivamente humanas incluso en la era de la IA.

La tecnología seguirá evolucionando, pero la ventaja competitiva estará en quienes sean capaces de combinar inteligencia artificial con conversaciones centradas en el cliente, pensamiento estratégico y relaciones de largo plazo.