If you have built a branching scenario, you know how fast the paths multiply.
A learner responds.
A customer reacts.
You need another branch.
Soon you have a large flowchart for one conversation.
That is why I am curious about Jev, a decision model from TypeSafe AI.
It may have a role in learning games where a character needs to respond to what just happened.
First, what is Jev?
Most AI tools we use produce words.
Jev returns a structured decision instead. TypeSafe says it does not generate text at all.
You give it the current situation and a question.
TypeSafe offers three kinds of question:
- Choice. You supply the possible answers and it picks one.
- Score. It rates the situation on a rubric you write, ordered from low to high.
- Noul. It estimates whether a statement is true, as a number from 0 to 1.
What does that look like in a game?
Imagine a game character who can ask for information, offer help, or end the conversation.
The developer writes those actions.
The game sends them to Jev as the options in a Choice question.
Jev picks one from that menu.
The game still has to check the answer and carry out the action.
The dialogue comes from written text or a separate writing system. Jev does not write it.
That menu is what I mean by bounded options.
The developer decides which actions the game allows right now. Jev does not create that menu.
And a choice can follow the game's rules and still teach the wrong lesson.
A believable customer reaction might accidentally reward a learner for skipping an escalation step.
Does a high confidence score mean Jev is right?
No.
For Choice and Score, TypeSafe computes confidence from how concentrated the answer's probabilities are.
All of it on one option gives 1.0. The more evenly it spreads, the lower the number.
TypeSafe's own documentation says a high-confidence answer can still be wrong.
So confidence tells you how strongly the model favors an option.
It does not tell you the option fits your policy or your learning goal.
Noul has no separate confidence field. Its 0 to 1 number is the whole answer.
I would test any cutoff against real examples from the actual task.
What has been built in games?
Coffee Under Fire
Coffee Under Fire calls itself an experimental game: a coffee-delivery arena shooter that runs in the browser.
Its documentation lays out a clear split of work:
- Game code builds what one character is allowed to perceive.
- Game code lists the legal actions, with destinations, targets, and durations.
- Jev picks one.
- The app checks the answer, then rechecks that the action is still legal before running it.
The project also records accepted decisions, so a run can be replayed without calling Jev.
This is a documented integration.
It is not an independent test of how well the game plays or teaches.
JevScape
In JevScape, Jev plays RuneScape through about 50 available actions.
Fishing, walking, eating, and closing a dialog are among them.
Before each decision, the program sends the character's status, recent events, and the last server message.
The developer published a recorded fishing session and live-run logs.
The developer also says it plainly: the action catalog carries the game knowledge.
This is an agent playing a game.
It is not a study of people learning from one.
jev-arcade
jev-arcade is a public project that connects Jev to Snake, 2048, and Tetris.
Its author reports a September 21 test using Jev 1.13.0.
Three seeded episodes per game, 60 turns each, with the fallback switched off.
Here are the reported scores, Jev versus a programmed rule-based player:
- Snake: 70 versus 80.
- 2048: 484 versus 497.
- Tetris: 167 versus 2,333.
Do the division.
In Snake, Jev reached about 88 percent of the rule-based score.
In 2048, about 97 percent.
In Tetris, about 7 percent.
That is three episodes per game.
It is a small, specific test, not a verdict on Jev.
The author links the Tetris gap to planning several pieces ahead. The test itself does not isolate the cause.
The project's code also hands the move to the rule-based player when Jev's confidence falls below a threshold the developer sets.
In a separate run at 0.5, every Tetris move went to the rule-based player.
That score belongs to the rule-based player, not to Jev.
What do these projects prove?
They show real ways to wire Jev into a game.
I have not reproduced their runs.
They do not show that a Jev-powered learning game teaches better.
They do not show that its choices will match an instructional designer's judgment.
What would a Jev learning game look like?
Here is the one I would test.
A short practice game for a new customer support employee.
The goal: handle a frustrated customer while following the team's real escalation procedure.
The first scene gives the learner a customer message, an order history, and three actions:
- Ask one clarifying question.
- Offer an approved remedy.
- Escalate.
Each action changes what the customer knows and how the conversation goes.
I would build it around controlled responsiveness.
The customer can react differently. The facts and the teaching rules stay under the designer's control.
With a subject-matter expert, I would write the order history, allowed remedies, escalation rules, scoring, completion criteria, and feedback.
Jev would have no authority to change any of that.
After the learner acts, coded rules would filter a menu of approved customer reactions.
Jev would get that menu and the scenario state, then pick a reaction.
The game would check it, show the written dialogue, and apply the planned consequence.
An unusable answer would trigger a written fallback.
Say the learner asks a good clarifying question.
The customer might reveal a missing detail.
Or stay frustrated.
Either reaction gives useful practice without inventing a policy exception.
I would log the state, the available reactions, the pick, the probabilities, the confidence, and any fallback, so a reviewer can trace what happened.
How would you know if it helps?
This is a proposal to test. I would take four steps.
- Build a review set. Collect 50 to 100 scenario states: missing facts, poor escalation choices, prohibited remedies, repeated frustration, and ambiguous cases. Have at least two subject-matter experts label reactions acceptable, preferred, or unacceptable, with reasons. Hold some cases back for the final check.
- Inspect the decisions. Count invalid outputs, unacceptable reactions, and how often the fallback fires. Check whether low confidence lines up with the genuinely ambiguous cases. Measure whether answers arrive fast enough for the conversation.
- Compare simpler versions. Run the same cases through a fixed branch, a rules-only branch, and Jev choosing approved reactions. Compare quality, speed, and the effort to write and maintain each one.
- Test learning separately. See whether learners choose well in new cases, explain their reasoning, and keep the skill after a delay.
A scenario can feel more responsive without teaching anything better.
Matching an expert's first choice matters less than never picking an unacceptable reaction.
The model has to earn its place over a simpler design.
Start with one scene.
Write the menu of approved reactions before you connect any model.
Sources
Checked September 27, 2026.
- TypeSafe documentation, Introduction, for what Jev is, its three question types, and that it does not generate text: docs.typesafe.ai/introduction
- TypeSafe documentation, Confidence, for how confidence is computed and that it is not a guarantee of correctness: docs.typesafe.ai/confidence
- TypeSafe documentation, Score, for the ordered rubric: docs.typesafe.ai/primitives/score
- TypeSafe documentation, Noul, for the 0 to 1 answer with no separate confidence: docs.typesafe.ai/primitives/noul
- Coffee Under Fire, NPC control documentation, for the decision pipeline and replay: github.com/joaoh82/coffee-under-fire
- JevScape README, for the action catalog, the decision inputs, and the recorded session: github.com/Skyvern-AI/jevscape
- jev-arcade README, for the September 21 benchmark and the threshold 0.5 run: github.com/CankatSarac/jev-arcade
- jev-arcade player code, for the low-confidence fallback: jev_player.py
Want This Capability on Your L&D Team?
I can build your game from your source material, run a Build Day with your team, or you can start free with my tools.
See how we'd work together Try the tools free