Extended reality builds the place. Virtual, augmented and mixed reality put digital content around a person, from a fully simulated room to an object anchored on a real desk. Without conversational AI, what that content says back is set in advance.
Put the two together and a character can listen, work out what was actually said, and answer differently as the conversation goes, instead of running a pre-recorded script.
Adding AI changes what an XR environment can do.
The steps repeat in real time as the conversation continues.
PTR builds the language-model step in that diagram: AI companions and conversational agents that answer from approved knowledge within defined boundaries. The same twin can be reached through chat, voice or an on-screen avatar. See PTR's digital twins for how one spans all three.
The pairing is useful because it combines a spatial setting with a character that responds to what the learner says and does.
Together they let someone rehearse a conversation, a procedure or a decision under pressure, as many times as it takes, with no patient, colleague or piece of equipment at risk.
Combining the two adds problems neither has on its own: response delay, model boundaries, and what the hardware can carry.
Dropping a text window into a VR scene does not, on its own, create the effect people mean by "AI and XR."
Participants practising interpersonal skills rated a VR-embodied conversational AI agent more effective than the same AI reached through text alone in Cornell/Stanford, CSCW 2025 research.
A separate comparison at ACM SIGGRAPH MIG, 2025 found an AI NPC tutor reached moderate to high co-presence but rated well below a human tutor on overall social presence.
| Scripted character | AI-driven character | |
|---|---|---|
| Dialogue options | A fixed set of pre-written choices | Open-ended natural speech or text |
| Adapts to unexpected phrasing | No, falls back to a default line | Usually, within its guardrails |
| Needs re-authoring for a new topic | Yes, every branch is hand-written | Often just an update to the source material |
A learner in a headset, a character who listens, and a reply nobody scripted.
AI and XR combine when a language model decides what a virtual character says and does, and XR gives that character a spatial, embodied place to say and do it.
Either piece alone falls short: XR without AI is a fixed script; AI without embodiment is a chat window.
Still have a question about AI and XR?




Research on conversational agents in VR, social presence, interpersonal skills training and response delays.

An early implementation of a conversational character living inside a VR scene.

The co-presence comparison behind the "more than a chatbot" band above.

The embodiment study behind the "more effective than text alone" note above.

The filler-action study behind the response-latency note above.
Extended reality builds the place. Virtual, augmented and mixed reality put digital content around a person, from a fully simulated room to an object anchored on a real desk. Without conversational AI, what that content says back is set in advance.
Put the two together and a character can listen, work out what was actually said, and answer differently as the conversation goes, instead of running a pre-recorded script.
Adding AI changes what an XR environment can do.
The steps repeat in real time as the conversation continues.
PTR builds the language-model step in that diagram: AI companions and conversational agents that answer from approved knowledge within defined boundaries. The same twin can be reached through chat, voice or an on-screen avatar. See PTR's digital twins for how one spans all three.
The pairing is useful because it combines a spatial setting with a character that responds to what the learner says and does.
Together they let someone rehearse a conversation, a procedure or a decision under pressure, as many times as it takes, with no patient, colleague or piece of equipment at risk.
Combining the two adds problems neither has on its own: response delay, model boundaries, and what the hardware can carry.
Dropping a text window into a VR scene does not, on its own, create the effect people mean by "AI and XR."
Participants practising interpersonal skills rated a VR-embodied conversational AI agent more effective than the same AI reached through text alone in Cornell/Stanford, CSCW 2025 research.
A separate comparison at ACM SIGGRAPH MIG, 2025 found an AI NPC tutor reached moderate to high co-presence but rated well below a human tutor on overall social presence.
| Scripted character | AI-driven character | |
|---|---|---|
| Dialogue options | A fixed set of pre-written choices | Open-ended natural speech or text |
| Adapts to unexpected phrasing | No, falls back to a default line | Usually, within its guardrails |
| Needs re-authoring for a new topic | Yes, every branch is hand-written | Often just an update to the source material |
A learner in a headset, a character who listens, and a reply nobody scripted.
AI and XR combine when a language model decides what a virtual character says and does, and XR gives that character a spatial, embodied place to say and do it.
Either piece alone falls short: XR without AI is a fixed script; AI without embodiment is a chat window.
Still have a question about AI and XR?
Research on conversational agents in VR, social presence, interpersonal skills training and response delays.

An early implementation of a conversational character living inside a VR scene.

The co-presence comparison behind the "more than a chatbot" band above.

The embodiment study behind the "more effective than text alone" note above.

The filler-action study behind the response-latency note above.