Skip to content

SPIKE-16: photorealistic virtual-trainer avatar

Status: research spike complete (2026-07-22). Author: foundry-researcher (Opus). No Azure resources were created, read, or modified for this spike; it is a documentation and first-party source review only. Read-only, exploratory, early-stage: per the tasking, "the technology is not yet mature or affordable enough to recommend adoption" was an acceptable outcome, and that is close to where the evidence lands.

Scope: resolve the open questions behind the decision log D-16 (the refined Charter vision), in which the owner wants to personally produce a photorealistic virtual-trainer avatar as a real production artifact type this repo's methodology should eventually support, alongside the image (SPIKE-01) and voice (SPIKE-02) work already built. The owner did not know the correct technical term for this product category going in, so the first job here is to fix the vocabulary, then survey first-party Azure and credible third-party options. Backlog brief: this spike's brief.

Grounding rule: every load-bearing claim is tied to a first-party source (Microsoft Learn, Microsoft Foundry docs, or a named vendor's own page) with an inline URL. Anything the sources do not state is marked UNKNOWN with the test or doc that would resolve it. No figure here is invented; community-forum figures are labelled as such and not treated as authoritative.


Question

Five questions from the tasking:

  1. What is the correct technical term for this product category, confirmed (not assumed) against real vendor and Microsoft documentation.
  2. What first-party Azure/Microsoft options exist, if any, before assuming a third-party vendor is required.
  3. What third-party vendor options exist, surveyed on photorealism, whether they need a filmed actor or can generate from a photo/description, lip-sync accuracy, commercial licensing, and cost.
  4. How this relates to the existing voice work, cross-referencing SPIKE-02 and SPIKE-07 lip-sync (viseme/phoneme timing) findings, since a vision-only answer without a compatible audio-timing source is incomplete.
  5. Azure AI Foundry integration fit: can a recommended option run inside or alongside the existing Foundry account, or does it need a separate hosting path.

Findings

1. The correct technical term

  • Microsoft's own product name is "Text to speech avatar." Microsoft defines it precisely: it "converts text into a digital video of a photorealistic human (either a standard avatar or a custom text to speech avatar) speaking with a natural-sounding voice." (what-is-text-to-speech-avatar) This is the search term that finds the first-party path; "digital human," "photorealistic avatar," and "neural avatar" do not, on their own, land on the Microsoft product.
  • "Talking-head synthesis" is the accurate technical descriptor for the photo-driven variant, and Microsoft uses it. Microsoft's "photo avatar" is documented as generating "a talking-head video from a single image," and the roster of Microsoft-provided photo avatars is literally titled "Talking heads." (voice-live-how-to, batch-synthesis-avatar-properties) The photo avatar is driven by the VASA-1 base model. (voice-live-how-to)
  • "Neural avatar" is Microsoft's term for the trained model object. The custom-avatar consent docs describe recording an actor "to create neural avatar models." (custom-avatar-create)
  • "Digital human" and "virtual presenter/spokesperson" are marketing-layer umbrella terms, not the precise product category. "Digital human" in the wider industry (for example NVIDIA ACE, Soul Machines, UneeQ) typically implies a real-time interactive agent, a superset of the talking-head video this spike is about; "virtual presenter/spokesperson" is a use-case label the third-party vendors below apply to the same underlying talking-head technology. Neither is wrong, but neither is the term to search Microsoft docs with.

Resolution for future sessions: search first-party material as "text to speech avatar" (and "photo avatar" / "talking head" for the single-image variant); search the third-party market as "AI avatar generator" or "talking-head video / avatar video". Treat "digital human" as the broadest umbrella and "virtual trainer/presenter" as the use-case framing, not the technology name. The owner's instinct ("photorealistic virtual-trainer avatar") maps cleanly onto Microsoft's "photorealistic human ... avatar" definition.

2. First-party Azure/Microsoft options

There is a genuine first-party path: Azure AI Speech "Text to speech avatar," available inside Microsoft Foundry. It is not a separate product from the Speech/AIServices resource this repo already stands up for MAI-Voice-2.

  • Two avatar sources, two synthesis modes.
  • A custom, one-of-a-kind trainer avatar is possible two ways:
    • Custom video avatar: requires at least 10 minutes of video recording of the actor talent as training data plus a recorded consent statement; the resulting avatar "looks the same as the avatar talent in the training data" and Microsoft does not support changing its clothes, hairstyle, or appearance after the fact. (what-is-custom-text-to-speech-avatar)
    • Custom photo avatar (preview): requires only a single photo (and a consent video if the photo is of a real person); it is head-only, VASA-1 driven, and currently a manual offline process with limited access. (what-is-custom-text-to-speech-avatar, custom-photo-avatar-create)
  • Resource and tier: it runs on a Microsoft Foundry (or Speech) resource, Standard S0 only; "Select the Standard S0 pricing tier to access avatars." (real-time-synthesis-avatar) This is the same resource kind and tier SPIKE-02 already provisions for MAI-Voice-2, so no new resource class is introduced. (SPIKE-02)
  • No-code entry point exists: the Foundry portal has a Text to Speech Avatar playground where you pick a standard avatar, background, and voice, type text, and hit Generate. (get-started-text-to-speech-avatar) This is the cheapest possible audition for the owner before any commitment.
  • Governance is built in: avatar output is automatically watermarked and adopts the C2PA Content Credentials standard so audiences can see the video is AI-generated; watermark detection is available to approved users via Microsoft. (transparency-note) Custom avatar is a Limited Access feature, available by registration only and only for certain approved use cases. (limited-access)

3. Third-party vendor options

If the first-party photorealism bar or roster does not fit, the credible market leaders are HeyGen, Synthesia, and D-ID (with Tavus and Colossyan as adjacent names). All figures below are vendor-published or from vendor-comparison articles, reviewed 2026-07-22, and are not first-party Microsoft sources; treat pricing as indicative and re-confirm on the vendor's own terms page before any spend.

  • HeyGen offers photorealistic "Avatar IV" quality and can build an "Instant Avatar" from a single selfie in about five minutes (photo-driven, no filmed shoot required). Plans run free / 29 / 99 / 149 USD per month, but the photorealistic tier is metered at roughly 20 credits per minute, so effective cost is materially higher than the sticker plan. For talking-head realism with facial micro-expressions, vendor-comparison writeups rate Avatar IV at or near the top. (heygen.com best AI avatar generators, eesel HeyGen pricing 2026)
  • Synthesia targets a "Studio Avatar" built from a professional filmed recording of an actor (not a single photo), priced around 1,000 USD per year for the studio avatar, and rates highest for full-body cinematic multi-angle performance in the same comparisons. (d-id.com best AI avatar generators, colossyan HeyGen vs Synthesia)
  • D-ID animates any headshot photo (photo-driven talking head), with a cheapest paid tier around 5.99 USD per month. (d-id.com best AI avatar generators)
  • Photo-vs-actor axis: HeyGen (Instant Avatar) and D-ID generate from a single photo; Synthesia's studio avatar wants a filmed actor. Microsoft's own split mirrors this exactly (photo avatar from one image vs custom video avatar from ten-plus minutes of footage), so the "do I need to film a person?" decision is the same fork on either the first-party or third-party path.
  • Lip-sync: all of these are audio-driven talking-head systems, meaning lip-sync is produced inside the vendor's renderer from the audio track, not something the caller wires up. This is the same architecture as the Azure avatar (see Finding 4).
  • Commercial licensing / children's content: UNKNOWN in detail. The vendor-comparison articles do not spell out commercial-use restrictions or minors/education constraints; each vendor's own terms of service must be read before using an avatar in content aimed at children. This is called out again in the UNKNOWN table.

4. Relation to the existing voice work (cross-reference SPIKE-02 and SPIKE-07)

This is the pivotal finding, and it is good news for the first-party path.

  • On the Azure avatar, lip-sync is handled entirely inside Microsoft's renderer; you do not supply viseme or word-boundary timing. The documented custom-avatar component sequence is: text goes into a text analyzer that emits a phoneme sequence; the audio synthesizer produces the speech audio; then "the text to speech avatar model predicts the image of lip sync with the speech audio, so that the synthetic video is generated." (what-is-custom-text-to-speech-avatar) In real-time, "the avatar is synchronized with the audio output" automatically. (voice-live-how-to)
  • Therefore the single biggest open risk from the voice spikes does not block the first-party avatar's lip-sync. SPIKE-02's headline unknown (#1) is whether MAI-Voice-2 emits usable WordBoundary events, because the reader app's read-along highlight needs client-visible word timing. (SPIKE-02) The first-party Azure avatar does not need that: it derives lip motion from its own internal phoneme sequence and the audio (server-side), so the word-boundary question is orthogonal to whether the avatar's mouth moves correctly. The caller supplies text or SSML plus a voice, not timing data.
  • But for a custom or third-party avatar renderer, SPIKE-07's viseme findings become load-bearing, and they sharpen the MAI-voice risk. SPIKE-07 (now on disk) confirms that Azure native Speech can emit client-side lip-sync data (mstts:viseme SSML plus the SDK VisemeReceived event: redlips_front viseme IDs, and FacialExpression 3D blend shapes as a 60 FPS matrix of 55 ARKit-style facial positions), which is exactly what a self-hosted or third-party avatar rig would consume. (SPIKE-07, how-to-speech-synthesis-viseme) Two caveats from SPIKE-07 matter here: (a) native viseme IDs are en-US neural voices only and blend shapes are en-US and zh-CN neural only, so the en-GB narrator gets no native visemes; and (b) MAI-Voice-2 emits no viseme (or word-boundary) events at all on its REST path. So if a trainer avatar is built on a custom/third-party renderer and keeps the MAI narration voice, there is no native lip-sync signal, and SPIKE-07's decoupled path (audio-to-viseme via Rhubarb Lip Sync, or a phoneme-to-viseme forced aligner) is required. This only bites off the first-party Azure avatar, whose renderer does its own phoneme-to-lip mapping internally regardless of voice.
  • The real voice-side unknown for the avatar is which voice can drive it. The avatar's voice may be a standard voice, a professional/custom neural voice, a personal voice, or (for a custom video avatar) a "voice sync for avatar." (what-is-text-to-speech-avatar, voice-sync-for-avatar) Whether MAI-Voice-2 (the repo's chosen narration voice, itself in preview) is accepted as an avatar voice is not documented and is UNKNOWN. If the owner wants the trainer to sound like the established narration cast (Harper, Lisa, Ethan per SPIKE-02), that compatibility has to be tested, and if MAI voices are not accepted, the fallback is a standard neural voice or a custom/professional voice, which changes the audio identity.
  • On the third-party path, if you bring your own MAI-Voice-2 audio track, the vendor's audio-driven lip-sync consumes the waveform directly, so again word boundaries are not needed; the tradeoff is that the audio then leaves the Azure boundary and the vendor's TTS may be preferred instead.

5. Azure AI Foundry integration fit

  • First-party avatar fits inside the existing Foundry footprint. It is the same Foundry/AIServices S0 resource family as MAI-Voice-2 and MAI-Image-2.5, uses the Speech SDK / REST batch endpoints, and needs no separate SaaS. (real-time-synthesis-avatar, batch-synthesis-avatar)
  • Region caveat, and it matters for this repo. Per the Speech regions table (ttsavatar tab), eastus (the repo's primary region from SPIKE-02/SPIKE-03) supports real-time avatar, batch avatar, custom avatar usage, custom photo avatar creation, and voice sync, but the "Custom video avatar training" column is blank for eastus. Custom video avatar training is only offered in westus2, southeastasia, westeurope, and swedencentral; a model trained there can then be copied to an eastus resource for use. (regions) So a filmed-actor trainer avatar would require a train-elsewhere-then-copy step, whereas a photo avatar or a standard avatar has no such split in eastus.
  • Third-party is entirely outside Foundry. HeyGen/Synthesia/D-ID are external SaaS APIs with their own auth, billing, data handling, and content governance; adopting one means a second vendor relationship and taking generated media (and possibly the narration audio) outside the Azure boundary and the C2PA/watermark guarantees Microsoft provides automatically.
  • New modality for the model registry. An avatar is a talking-head video artifact. The registry schema's kind enum is ["image", "voice", "video", "reasoning", "text", "speech-to-text", "embedding"] as of the schema v2 work recorded in ADR-0018. It held four values when this spike was written. (models/registry.schema.json) If this proceeds past the spike, either reuse video or (cleaner, given the distinct governance and pipeline) add an avatar value to that enum, and add a registry entry whose sourceRef points at the follow-up ADR. Do not do this now; it is a decision for the ADR, noted here so it is not missed.

Content-safety and children's-content rules (called out specifically, per this repo's worked example)

This repo's first proven build produces children's content, so the responsible-AI posture is load-bearing, not boilerplate:

  • Educational/interactive learning is an explicitly approved use case for custom neural voice and, by the same responsible-AI framework, the avatar features: "To create a fictional brand or character voice for reading or speaking educational materials, online learning, interactive lesson plans." (transparency-note use cases) A virtual trainer/teacher persona is squarely in-scope.
  • Exploiting or manipulating children is an explicitly prohibited use. (disclosure-voice-talent)
  • Parental/guardian disclosure is required for content aimed at minors: "If your use case is intended for minors or children, you'll need to ensure that your disclosure is clear and transparent so that parents or legal guardians can understand the role of synthetic media." (transparency-note evaluating)
  • Custom avatar is Limited Access by registration, and if trained from a real person (including the owner) requires a recorded consent statement that Microsoft biometrically verifies against the training footage. (limited-access, custom-avatar-create)

The net: the first-party path carries the disclosure, consent, watermark, and C2PA machinery that a children's brand should want, whereas a third-party vendor's equivalent protections must be verified per vendor.


What is still UNKNOWN

#UnknownWhy it is not resolved hereWhat resolves it
1Exact avatar billing rate. Docs state avatar is billed per second of video output, and for real-time per active second whether speaking or silent, but not the number. (text-to-speech pricing note)The Azure Speech pricing page is client-rendered and did not return text to the fetch tool (same failure SPIKE-02 hit). A community Microsoft Q&A cited roughly 1.44 USD per 6-minute block of video, but that is a forum answer, not authoritative. (community Q&A)Open azure.microsoft.com/pricing/details/cognitive-services/speech-services in a browser (avatar pricing shows only for regions where the feature is available), or read the resource billing meters after a short test render.
2Can MAI-Voice-2 be used as the avatar voice? (Leaning "no", per SPIKE-07.)Avatar-voice docs enumerate standard, professional/custom, personal, and voice-sync voices; they do not name MAI-Voice-2 (itself a preview prebuilt voice). SPIKE-07 reads the same docs as pairing the managed avatar with "standard or custom voice models," not the MAI preview model, treating MAI as excluded from the managed renderer. (SPIKE-07) The enumeration is suggestive of exclusion but not an explicit "MAI not supported" statement, so it stays UNKNOWN pending a live test.In the Foundry avatar playground, select a standard avatar and try to pair it with en-US-Harper:MAI-Voice-2 (or another MAI voice); success or an unsupported-voice error settles it. If unsupported (the likely outcome), the trainer's voice must be a standard/custom voice, which changes the audio identity from the narration cast; a custom rig on MAI audio then needs the SPIKE-07 decoupled lip-sync path.
3Custom avatar training compute cost and endpoint hosting cost.Only community/2024 figures found (approximately 52 USD per training compute hour, approximately 5 USD per hour endpoint hosting, 20 to 48 hours training). (community Q&A)Confirm on the Speech pricing page and via Microsoft sales; note the hosting meter runs continuously while a custom avatar deployment exists, so it must be suspended when idle.
4Does the standard-avatar roster include a persona that fits a "trainer" for the brand demographic?A community Q&A reported the standard avatars "all exhibit Asian traits" and asked for more diversity; Microsoft's standard-avatar and talking-heads lists are documented but the roster's range is not summarized in a way this spike could verify. (community Q&A, standard-avatars)Review the standard-avatars and talking-heads lists directly in the playground; if none fits, the answer is a custom photo or custom video avatar, which raises cost and gating.
5Photorealism quality bar of the Azure avatar vs HeyGen Avatar IV / Synthesia today.Vendor-comparison articles rate HeyGen/Synthesia top for photorealism; a Microsoft Q&A shows users asking Microsoft to improve avatar naturalness. (community Q&A) Quality is subjective and version-dependent.Hands-on: render the same script through the Foundry avatar playground and a HeyGen/D-ID trial, and have the owner judge.
6Limited Access approval timeline and eligibility for the owner's specific virtual-trainer use case.Custom avatar (and custom photo avatar) require registration via aka.ms/customneural; approval timing and per-use-case eligibility are not published. (limited-access)Submit the intake form; this can start in parallel since standard/photo-avatar auditions do not depend on it.
7Third-party commercial licensing for children's/education content.The vendor-comparison articles do not detail minors/education restrictions or commercial-use limits.Read HeyGen / Synthesia / D-ID terms of service directly before any content aimed at children.

None of these are provisioning go/no-go items; they are the cost, quality, voice-compatibility, and gating confirmations an ADR would need. The single cheapest way to close the largest cluster (2, 4, 5) is a no-cost Foundry avatar playground audition.


Recommendation

  • Vocabulary is fixed: the category is "text to speech avatar" (Microsoft), a "talking-head" video of a "photorealistic human"; the photo-from-one-image variant is a "photo avatar" driven by VASA-1. Future sessions should search first-party as "text to speech avatar / photo avatar" and the third-party market as "AI avatar generator." "Digital human" and "virtual trainer/presenter" are umbrella and use-case labels, not the product name. This is the single durable output the tasking asked for so the term is not re-derived.
  • Do not adopt yet. Recommend a low-cost evaluation, not a deployment. This is genuinely early-stage: the exact per-second price (UNKNOWN #1), whether the trainer can keep the narration cast's MAI voice (UNKNOWN #2), custom-avatar training cost and Limited Access approval (UNKNOWN #3, #6), and the photorealism-vs-third-party gap (UNKNOWN #5) are all open, and any one of them could sink affordability for a single owner-produced artifact. "Not yet mature or affordable enough to commit" is the honest position, matching the tasking's expectation.
  • If evaluated, evaluate the first-party Azure text to speech avatar first, for three concrete reasons: it lives in the same Foundry S0 resource as the existing image and voice work (no new vendor, no data leaving the Azure boundary); it handles lip-sync internally from its own phoneme sequence, so the word-boundary/viseme uncertainty tracked in SPIKE-02 and SPIKE-07 does not block it; and it ships automatic watermarking, C2PA content credentials, consent verification, and Limited Access gating that a children's-content brand should actively want. The first step is a zero-commitment audition in the Foundry avatar playground of a standard avatar plus a narration voice.
  • Keep HeyGen (Avatar IV) as the third-party comparator if the owner judges the Azure photorealism insufficient; it can build a photorealistic avatar from a single selfie in minutes and rates top for talking-head realism, at the cost of a separate SaaS relationship, metered per-minute pricing, media and audio leaving Azure, and licensing that must be vetted for children's content. Synthesia is the option if a filmed-actor, full-body studio avatar is wanted; D-ID is the cheapest photo-animation entry.
  • Decide the filmed-actor question early, because it forks both paths identically: a photo avatar (one image, VASA-1) is cheaper and faster but head-only; a custom video avatar (ten-plus minutes of footage, Limited Access, and on Azure a train-in-westus2/southeastasia/westeurope/swedencentral-then-copy-to-eastus step) is higher fidelity and half or full body. For a first owner-produced artifact, a photo/standard avatar is the pragmatic start.
  • A follow-up ADR is required before any deployment (spike then ADR then design then deploy gate), and it must trace to this spike. That ADR should also record the model-registry consequence: an avatar is a new modality, so add an avatar value to the kind enum in models/registry.schema.json (or deliberately reuse video) and create a registry entry referencing the ADR. Do not touch the schema in this spike.

Net: there is a clean first-party path that fits this repo's Foundry footprint and its children's-content governance, and the lip-sync worry from the voice spikes turns out not to apply to avatars. But the cost, voice-compatibility, and photorealism unknowns are real enough that the right next move is a free playground audition and a Limited Access intake, not a build.


Sources

First-party (Microsoft Learn and Microsoft Foundry responsible-AI docs) unless noted, reviewed 2026-07-22:

Third-party vendor and comparison pages (vendor-published, NOT first-party Microsoft; pricing indicative, reviewed 2026-07-22):

Companion spikes in this repo:

  • SPIKE-02 (MAI-Voice-2: voice cast, word-boundary unknown, S0 AIServices resource): SPIKE-02-voice-model.md
  • SPIKE-07 (speech models / lip-sync: native mstts:viseme and FacialExpression blend shapes, en-US/zh-CN locale limits, MAI-Voice-2 emits no visemes, decoupled Rhubarb/forced-aligner path): SPIKE-07-speech-models.md