Technology Sep 01, 2026 · 9 min read

Build Real-Time Client-Side TTS in Angular Using Firebase AI Logic and Gemini

In my previous article on Dev.to, I showed how to build expressive text-to-speech using Gemini and Firebase Cloud Functions. While that setup worked beautifully, it required building and deploying custom backend endpoints. Firebase AI Logic enables running Gemini TTS client-side with production secu...

DE
DEV Community
by Connie Leung
Build Real-Time Client-Side TTS in Angular Using Firebase AI Logic and Gemini

In my previous article on Dev.to, I showed how to build expressive text-to-speech using Gemini and Firebase Cloud Functions. While that setup worked beautifully, it required building and deploying custom backend endpoints. Firebase AI Logic enables running Gemini TTS client-side with production security, eliminating custom backend code and automating deployment via Git integration.

1. Technical Stack

Project technical stack:

  • Angular 22: The latest version as of August 2026.
  • Node.js LTS: The LTS version as of May 2026.
  • Firebase Remote Config: Manages dynamic parameters, such as TTS model names, and conditionally configures limited-use tokens.
  • Firebase AI Logic: Interfaces with Gemini generative models.
  • Firebase App Check: Prevents abuse and verifies token authenticity.
  • Firebase App Hosting: Streamlined deployment to the Firebase App Hosting environment. With a single code push, the latest changes deploy to the production environment.

The public Google Gemini Developer API is restricted in my region (Hong Kong). However, the Agent Platform Gemini API (Google Cloud) offers enterprise access that works reliably here, so I chose the Agent Platform Gemini API for this Firebase AI Logic demo.

1.1 Install global dependency

npm install -g firebase-tools

Install or update firebase-tools globally using npm.

firebase logout
firebase login

Log out and re-authenticate with Firebase.

1.2 Install dependencies

npm i --save-exact firebase
npm i --save-exact --save-dev firebase-tools serve

Install the dependencies to call the Firebase AI Logic API, use the Firebase CLI to generate files, and serve the production build.

firebase init

Execute firebase init and follow the prompts to set up the Firebase AI Logic, Emulators, App Hosting, and Remote Config.

If you have an existing project or multiple projects, you can specify the project ID on the command line.

firebase init --project <PROJECT_ID>

After completing the setup steps, the Firebase tools generate the configuration files such as .firebaserc and firebase.json. You can view the .firebaserc and firebase.json in the GitHub repo.

1.3. Generate Firebase Configuration File

We asked antigravity-cli (a terminal-first AI coding agent released by Google) and the Gemini Flash model to create two Node.js scripts to do the following:

# generated firebase configuration
firebase.config.json

Add firebase.config.json to .gitignore to prevent accidental commits.

During build time, Angular bundles both JSON files into the dist directory to provide initial configuration values.

2. Architecture

Architecture of the Firebase AI Logic TTS L16 to WAV Conversion

Users submit text to Firebase AI Logic to synthesize speech. Firebase generates the complete L16 audio payload and returns it to the client. Because the HTML audio element does not support the L16 format, the application converts the audio to a WAV Blob before binding the Blob URL to the element source.

Architecture of the Firebase AI Logic TTS L16 Playback

The second flow streams the audio and sends the L16 chunks to the Angular application. The audio player creates an AudioBufferSourceNode to play the chunk data and cleans up resources to prevent memory leaks.

3. Source Code

While the full codebase is available in the ng-firebase-tts repository, the application relies on Firebase Remote Config to manage configuration, App Check to prevent abuse, and App Hosting to deploy the Angular application.

The following sections illustrate how to initialize the Firebase app and App Check, and how to activate Remote Config values.

4. Firebase Integration

The public/firebase.config.json file contains the public Firebase API key, the sensitive reCAPTCHA Enterprise key, and the App Check debug token to bypass device attestation in local development. These values are critical in Firebase App and App Check initialization.

4.1. Firebase Infrastructure

@Service()
export class ConfigService {
  #app: FirebaseApp | undefined = undefined;
  #remoteConfig: RemoteConfig | undefined = undefined;

  /*... getter methods are omitted... */

  get appConfig(): AppRemoteConfig {
    return this.#appConfig;
  }

  get aiBackend(): AI {
    return this.#aiBackend;
  }

  async initialize(): Promise<void> {
    this.#app = initializeApp(firebaseConfig.app);

    (globalThis as any).FIREBASE_APPCHECK_DEBUG_TOKEN = firebaseConfig.appCheckDebugToken || true;

    initializeAppCheck(this.#app, {
      provider: new ReCaptchaEnterpriseProvider(firebaseConfig.recaptchaEnterpriseKey),
      isTokenAutoRefreshEnabled: true,
    });

    this.#remoteConfig = getRemoteConfig(this.#app);
    this.#remoteConfig.defaultConfig = remoteConfigDefaults;
    await fetchAndActivate(this.#remoteConfig);

    this.#appConfig = {
      vertexAILocation: getValue(this.#remoteConfig, 'vertexAILocation').asString(),
      useLimitedUseAppCheckTokens: getValue(
        this.#remoteConfig,
        'useLimitedUseAppCheckTokens',
      ).asBoolean(),
      geminiTTSModelName: getValue(this.#remoteConfig, 'geminiTTSModelName').asString(),
    };

    this.#aiBackend = getAI(this.#app, {
      backend: new AgentPlatformBackend(this.#appConfig.vertexAILocation),
      useLimitedUseAppCheckTokens: this.#appConfig.useLimitedUseAppCheckTokens,
    });
  }
}

The initialize method initializes Firebase App, configures App Check, sets up Firebase AI, and assigns Remote Config values to appConfig.

The appConfig object holds the location, the Gemini TTS model name, and the limited-use App Check token flag.

4.2. Remote Config and Construct Agent Platform Backend

Configure the TTS model name and the limited-use App Check token parameters in Firebase Remote Config. Both parameters have conditional values. When the Firebase web application is firebase-ai-logic-tts, the TTS model is gemini-3.1-flash-tts-preview, and the limited-use App Check token flag is set to true.

Gemini TTS Model Name Parameter

Use Limited-use App Check Token Parameter

Add condition to Use Limited-use App Check Token Parameter

export const AI_BACKEND = new InjectionToken<AI>('AI_BACKEND');

export function provideFirebase() {
  return makeEnvironmentProviders([
    {
      provide: AI_BACKEND,
      useFactory: () => inject(ConfigService).aiBackend,
    },
  ]);
}

The AI_BACKEND injection token provides a factory function to return the Firebase AI from ConfigService.

export const appConfig: ApplicationConfig = {
  providers: [
    ... other providers ...
    provideAppInitializer(async () => await inject(ConfigService).initialize()),
    provideFirebase(),
  ],
};

provideAppInitializer and provideFirebase initialize Firebase and configure Firebase AI during application startup.

5. TextToSpeechService Definition in Angular

The private createModel method calls getGenerativeModel to create a generative model configured with speechConfig. The subsequent workflows use this model to convert L16 audio to WAV or stream L16 playback directly.

@Service()
export class TextToSpeechService {
  readonly #configService = inject(ConfigService);
  readonly #modelName = this.#configService.appConfig.geminiTTSModelName;

  private createModel(voiceName: string) {
    return getGenerativeModel(this.#aiBackend, {
      model: this.#modelName,
      generationConfig: {
        responseModalities: [ResponseModality.AUDIO],
        speechConfig: {
          voiceConfig: {
            prebuiltVoiceConfig: { voiceName },
          },
          languageCode: 'en-US',
        },
      },
    });
  }
}

5.1 Flow 1: L16 Audio to WAV conversion

async synthesize(text: string, voiceName: string): Promise<Blob> {
  const model = this.createModel(voiceName);
  const result = await model.generateContent([text]);
  const chunk = this.extractValidChunkData(result.response);
  const { data, mimeType } = chunk;
  return convertToWav(decodeBase64(data), mimeType);
}

The extractValidChunkData method extracts binary audio data and the MIME type from the response payload.

The synthesize method retrieves the entire audio payload in L16 format. However, the HTML audio element does not support L16, so the application must convert the data to WAV format before assigning the Blob URL to the element source.

See the conversion to WAV code for the full implementation.

This implementation remains straightforward because it avoids managing response streams and incremental audio chunks. However, users suffer from latency when the text is long and it produces a long audio stream. To eliminate playback latency, the next section explores streaming chunks directly via the Web Audio API AudioContext.

5.2 Flow 2: L16 Audio Instant Playback

async *synthesizeStream(text: string, voiceName: string): AsyncGenerator<RawAudioBinary | undefined> {
    const model = this.createModel(voiceName);
    let firstMimeType = '';
    let sampleRate = DEFAULT_SAMPLE_RATE;

    const responseStream = await model.generateContentStream([text]);
    for await (const chunk of responseStream.stream) {
      const chunkData = this.extractValidChunkData(chunk);
      if (chunkData) {
        const { data, mimeType } = chunkData;
        const decodedData = decodeBase64(data);
        if (!firstMimeType && mimeType) {
          firstMimeType = mimeType;
          sampleRate = parseMimeType(firstMimeType).sampleRate;
        }

        yield { decodedData, sampleRate };
      }
    }

    yield undefined;
}

The synthesizeStream method returns an asynchronous generator that yields the raw binary data and the sample rate. The MIME type is audio/l16; rate=24000; channels=1, so parseMimeType extracts the sample rate for the audio player's AudioContext.

While AudioPlayerService is beyond the scope of this article, the source code demonstrates how to play back audio without using an HTML audio element.

Next, we will build a reactive user interface in Angular that renders an HTML audio element to play audio.

6. Build Reactive UI in Angular

The TextToSpeechComponent delegates audio management to the view service to maintain separation of concerns.

6.1 Text to Speech Component

The TextToSpeechComponent displays three buttons to generate speech from text across three different scenarios:

  • Scenario 1: Synthesizes the entire audio payload and binds the resulting Blob URL to an HTML audio element.
  • Scenario 2: Streams audio chunks incrementally and plays the speech immediately via the Web Audio API. This mode hides the HTML audio element because audio plays directly through the Web Audio API.
@Component({
  selector: 'app-text-to-speech',
  templateUrl: './text-to-speech.component.html',
  styleUrl: './text-to-speech.component.css',
  imports: [SpinnerIconComponent, NgTemplateOutlet],
  providers: [TextToSpeechViewService],
})
export class TextToSpeechComponent {
  private readonly speechService = inject(TextToSpeechViewService);

  interestingFact = input<string | undefined>(undefined);
  audioPrompt = input.required<string>();
  voice = input.required<string>();

  async generateSpeech(mode: GenerateSpeechMode) {
    const fact = this.interestingFact();
    await this.speechService.generateSpeech(mode, {
      prompt: this.audioPrompt(),
      voice: this.voice(),
      fact: this.interestingFact(),
    });
  }
}

6.2 TextToSpeechView Definition in Angular

The TextToSpeechViewService encapsulates TextToSpeechService and AudioPlayerService to coordinate speech synthesis and audio playback.

@Injectable()
export class TextToSpeechViewService {
  private readonly speechService = inject(TextToSpeechService);
  private readonly audioPlayerService = inject(AudioPlayerService);

  #audioUrl = signal<string | undefined>(undefined);
  audioUrl = this.#audioUrl.asReadonly();

  private processStreamChunk(isInitialized: boolean, playbackRate: number, chunk: RawAudioBinary) {
    if (!isInitialized) {
      this.audioPlayerService.initialize(chunk.sampleRate, playbackRate);
      isInitialized = true;
    }
    this.audioPlayerService.processChunk(chunk.decodedData);
    return isInitialized;
  }

  private async handleSync(promptArgs: FactConfig) {
    const blob = await this.speechService.synthesize(promptArgs.prompt, promptArgs.voice);
    this.setAudioUrl(blob);
  }

  private async handleStream(promptArgs: FactConfig) {
    let isInitialized = false;    
    const { prompt, voice } = promptArgs;

    for await (const chunk of this.speechService.synthesizeStream(prompt, voice)) {
      isInitialized = this.processStreamChunk(isInitialized, 1, chunk);
    }
  }

  private setAudioUrl(finalBlob: Blob | undefined) {
    if (finalBlob) {
      const createdUrl = URL.createObjectURL(finalBlob);
      this.#audioUrl.set(createdUrl);
      return createdUrl;
    }
    return undefined;
  }

  async generateSpeech(mode: GenerateSpeechMode, promptArgs: FactConfig) {
    revokeBlobURL(this.#audioUrl());
    this.#audioUrl.set(undefined);

    switch (mode) {
      case 'sync':
        await this.handleSync(promptArgs);
        break;
      case 'web_audio_api':
        await this.handleStream(promptArgs);
        break;
    }
  }
}

handleSync invokes the Firebase SDK to synthesize audio using the Gemini TTS model, generates a Blob URL, updates #audioUrl, and renders the HTML audio element.

handleStream calls the Firebase SDK to use the Gemini TTS model to stream the audio instead. The audio context receives the chunk data, plays it immediately, and avoids rendering the HTML audio element.

Conclusion

The integration of text-to-speech with Firebase AI Logic empowers Angular applications for real-time audio generation.

The Angular application handles text-to-speech entirely on the client. Pushing changes to Git triggers an automatic deployment to Firebase App Hosting.

Try cloning the GitHub repository, uploading an image to generate an obscure fact, and using the Gemini 3.1 Flash TTS preview model to speak it with the specified scene, emotion, and pace.

Resources

DE
Source

This article was originally published by DEV Community and written by Connie Leung.

Read original article on DEV Community
Back to Discover

Reading List