Skip to content

AI integration across runtime boundaries

PersonAI

AI conversations built across Unity, cloud services, identity, and streaming voice

PersonAI connects language models, retrieval, identity, and voice services in Unity. Its evolution from historical-persona demos to configurable provider integrations demonstrates cloud integration work across runtime boundaries; the current desktop iteration still has presentation and live-verification work remaining.

PersonAIProject overview
  • InteractionText & voiceUnity application
  • OrchestrationProvider bindingsIdentity · cancellation · errors
  • ServicesModels & retrievalLanguage · sources · speech
Conversation, configurable providers, and cloud AI services connected in Unity.
My role
Lead XR Developer
Period
Sep 2023 – Dec 2025
Client
Shenandoah University
Platforms
macOS / Meta Quest
Status
delivered

Tools & technologies

UnityC#Google Vertex AIGoogle Discovery EngineElevenLabs TTS SDKGoogle Cloud Speech-to-TextGoogle OAuth 2.0 PKCEGoogle Service Account JWT

FULL_STACK (17)

UnityC#Google Vertex AIGoogle Discovery EngineElevenLabs TTS SDKGoogle Cloud Speech-to-TextGoogle OAuth 2.0 PKCEGoogle Service Account JWTGoogle Cloud StorageNewtonsoft.JSONUI ToolkitUniversal Render PipelineInput SystemOpenAI Unity SDKSystems IntegrationAI IntegrationsRetrieval-Augmented Generation (RAG)

Impact & Results

  • Connected AI, retrieval, speech, and identity services within a single Unity application.
  • Made provider selection configurable per persona, reducing coupling between conversation flow and individual vendors.
  • Created explicit failure and cancellation paths across service boundaries.
  • Developed a reusable integration architecture for the next desktop conversation experience; current avatar presentation and live verification are still pending.

Overview

PersonAI is a Unity application for conversations with configurable AI personas, originally developed around historical figures at the Shenandoah Center for Immersive Learning. I built the connections between language models, retrieval, voice services, identity, and the Unity runtime.

The project has evolved from SDK-based conversations to custom cloud integrations. Earlier versions included dual-path retrieval and service-account authentication with custom RSA key parsing for Unity compatibility. The current codebase separates chat, retrieval, speech recognition, and speech synthesis behind provider interfaces, with configurable bindings per persona.

The current direction is a desktop 3D conversation experience. XR is a future path for this iteration; avatar presentation and live end-to-end verification remain separate from the implemented service integrations.

Role Summary

  • Led application engineering, from selecting and integrating cloud services to Unity runtime behavior and authoring workflows.
  • Built authentication and networking components across constrained Unity and cloud service boundaries.
  • Evolved the architecture from vendor SDK integration toward configurable provider contracts and shared orchestration.

Non-Technical Summary

PersonAI lets people interact with configured personas through text and voice. It connects cloud AI services to a Unity experience, with source retrieval intended to make answers easier to inspect.

Retrieved sources and citations help users evaluate an answer; they do not guarantee historical accuracy. When external retrieval fails, the orchestration layer can continue the conversation while flagging that failure for the interface.

Highlights

  • Built a Unity AI conversation application integrating language models, retrieval, streaming speech, and cloud authentication.
  • Implemented Google OAuth PKCE and, in an earlier architecture, service-account JWT authentication with custom DER/RSA parsing to handle Unity runtime limitations.
  • Developed provider-based orchestration for chat, retrieval, speech recognition, and speech synthesis.
  • Added cancellation, retry/backoff, typed errors, and retrieval-failure signaling to make integration failures understandable.

Quick Highlights

  • Configurable providers for chat, retrieval, speech recognition, and speech synthesis.
  • Google Vertex AI streaming responses and an Anthropic chat adapter.
  • Vertex AI Search retrieval with citations and explicit retrieval-failure signaling.
  • Google OAuth PKCE, Google speech recognition, and ElevenLabs streaming audio.

Technical Breakdown

A provider registry resolves each persona's chat, retrieval, speech recognition, and speech synthesis bindings. The chat orchestrator handles external retrieval when the chosen chat provider does not ground responses natively.

The Google adapter supports Vertex AI Search grounding and streamed Gemini responses through server-sent events. An Anthropic adapter provides a buffered response behind the same chat contract. Google speech recognition and ElevenLabs streamed PCM audio complete the voice integration.

Google OAuth uses PKCE with loopback and application deep-link return paths. Shared HTTP transport centralizes cancellation, retry/backoff, and typed failures. If external retrieval fails, a successful chat result carries a retrievalFailed flag rather than silently presenting the turn as sourced.

Earlier versions solved different runtime constraints, including a custom DER/PKCS#8 RSA parser for service-account JWT signing where Unity lacked the expected .NET API. Those details remain part of the version history.

Systems Used

  • Unity and C# for application orchestration and runtime integration.
  • Google Vertex AI, Vertex AI Search, and an Anthropic chat adapter.
  • Google OAuth PKCE and shared HTTP transport.
  • Google Cloud Speech-to-Text and ElevenLabs streaming speech synthesis.
  • Provider interfaces, persona configuration, cancellation, retries, and typed error handling.

Deep Dive

The integration challenge is coordinating services with different response models. Google can stream text and perform native retrieval; the Anthropic adapter returns a buffered result and can receive context from a separate retrieval stage. Provider contracts keep those differences out of the main conversation flow.

Retrieval and generation have distinct failure semantics. The orchestrator can preserve a generated answer when external retrieval fails, but marks the result so the interface can explain that sources were unavailable. This is an operational behavior, not an accuracy guarantee.

Voice and identity introduce additional runtime boundaries: streamed audio must reach Unity playback, network work must respond to cancellation, and OAuth must return control to the application. The project's progression shows how those concerns moved from individual SDK integrations into explicit application services.


KEYBOARD_SHORTCUTS

GO_TO

  • g h TERMINAL
  • g p PROJECT_INDEX
  • g e EXPERIMENTS
  • g s SKILL_MATRIX
  • g a ABOUT_ME
  • g r EXPERIENCE

GLOBAL

  • ⌘K COMMAND_PALETTE
  • / COMMAND_PALETTE
  • ` TERMINAL_MODE
  • ? SHORTCUT_PANEL
  • ESC CLOSE_OVERLAY

Chords expire after 1 second. Keys are ignored while typing.