Helm Chart Changes in Version 8.0.0

Version 8.0.0 of Capacity Private Cloud introduces new services, GPU acceleration, a new way of selecting ASR models, and a new optional ingress mode. Most of these are controlled through your Helm values file. This article explains each new setting, shows example values, and lists the changes you need to make when upgrading from 7.x.

Applies to: Capacity Private Cloud 8.0.0 Speech products (charts lumenvox, lumenvox-common, lumenvox-speech and lumenvox-external-services 8.0.0).

For the full list of changes in this release, see the Release Notes 8.0.0.


Before you begin

  • Kubernetes 1.35 or later is required. Version 7.x required 1.33.

  • Update your local Helm chart repository before installing or upgrading:

    helm repo update
    
  • The default image tag is now :8.0. If your values file does not set global.image.tag, your services will move to 8.0 images, and :8.0 always pulls the latest 8.0.x patch.

  • Review your values file before upgrading. Some 7.x settings have been renamed or moved and will stop the install with a message explaining what to change. See Upgrading from 7.x below.


Summary of new settings

FeatureHelm keyDefault
GPU inference (per service)global.gpu.<service>.enabledOff
GPU per languageglobal.asrLanguages[].services.<service>.gpuUses the global default
ASR serviceglobal.asr.enabledOn
Real-time transcription (new)global.transcribeRealtime.enabledOn
Batch transcription (new)global.transcribeBatch.enabledOn
Neural TTSglobal.neuralTts.enabledOn
Legacy TTS for a languageglobal.ttsLanguages[].legacyEnabledOff (Neural TTS is used)
Additional ASR models for a languageglobal.asrLanguages[].extraModelsNone
Default ASR model versionglobal.asrDefaultVersion"8.0.0"
Istio Gateway API ingress (new)global.lumenvox.ingress.className: "istio""nginx"
In-cluster databases for test/dev (new)global.enabled.externalServicesOff

Throughout this article, <service> means asr, transcribeRealtime or transcribeBatch.


1. GPU acceleration for ASR and transcription

ASR, real-time transcription and batch transcription can each run with GPU inference. GPU processing reduces latency and allows additional throughput. It is off by default.

You set a chart-wide default for each service, and can then override it for individual languages.

Turn on GPU for all three services, for every language:

global:
  gpu:
    asr:
      enabled: true
      count: 1              # nvidia.com/gpu devices per pod
    transcribeRealtime:
      enabled: true
    transcribeBatch:
      enabled: true

Or turn on GPU per language (in this example, GPU for English only, with batch transcription using two GPUs):

global:
  asrLanguages:
    - name: "en"
      services:
        asr:
          gpu: true
        transcribeRealtime:
          gpu: true
        transcribeBatch:
          gpu: { enabled: true, count: 2, visibleDevices: "all" }
    - name: "es"            # no gpu block: uses the global.gpu defaults

Available GPU options

OptionDescription
enabledTurns GPU inference on or off.
countNumber of nvidia.com/gpu devices per pod. Default 1.
visibleDevicesSets NVIDIA_VISIBLE_DEVICES (for example "all").
runtimeClassNameContainer runtime class to use (for example "nvidia").

At the language level, a plain true or false overrides only enabled. An object (as in the transcribeBatch example above) can override any of the four options.

Requirements and notes

  • GPU nodes with the NVIDIA device plugin installed.
  • GPU mode is currently available only with high-definition acoustic models. See section 5 for how to add high-definition models to a language.
  • GPU pods automatically tolerate the nvidia.com/gpu taint and the AKS spot-node taint, so they schedule onto tainted GPU node pools without extra settings.
  • Set visibleDevices: "all" on clusters where the device plugin sets NVIDIA_VISIBLE_DEVICES=void.
  • Set runtimeClassName where the node's default container runtime is not the NVIDIA runtime.

2. Turning speech services on and off

Each speech service can now be switched on or off independently. All four are on by default.

  • ASR, real-time transcription and batch transcription are deployed once for each entry in asrLanguages.
  • Neural TTS is deployed once for each entry in ttsLanguages.

Example: ASR and Neural TTS only, with no transcription services

global:
  asr:
    enabled: true
  transcribeRealtime:
    enabled: false
  transcribeBatch:
    enabled: false
  neuralTts:
    enabled: true

Notes

  • These settings sit directly under global, alongside enableNlu and enableNeuron.
  • global.enabled is still used only to turn whole charts on or off (lumenvoxSpeech, lumenvoxCommon, lumenvoxVb, externalServices).
  • global.minimalInstall: true still turns off ASR, transcription and TTS altogether, as in 7.x.

3. Neural TTS and legacy TTS

Neural TTS remains the default engine for every TTS language, as in 7.x. New in 8.0 is a single switch, global.neuralTts.enabled, that turns Neural TTS on or off for the whole deployment. Choosing the engine per language works as before, using legacyEnabled.

Voices must be listed explicitly in your values file so they are installed and available to the system.

Neural TTS (default) for a language

global:
  neuralTts:
    enabled: true           # default
  neuralttsDefaultVersion: "8"
  ttsLanguages:
    - name: "en_us"
      voices:
        - name: "aurora"
        - name: "caspian"

Legacy TTS for a language

global:
  ttsLanguages:
    - name: "en_us"
      legacyEnabled: true
      voices:
        - name: "chris"

Important: Turning Neural TTS off does not switch your languages to legacy TTS. With neuralTts.enabled: false, a language only gets legacy TTS if it has legacyEnabled: true. Any other language will have no TTS at all.

Notes

  • Voice versions default to neuralttsDefaultVersion ("8") for neural voices and ttsDefaultVersion for legacy voices. You can set a specific version for an individual voice with version.
  • The resource service now only advertises the voices for the engine that is actually deployed for each language.
  • If you are upgrading Neural TTS from version 6.0.0, clear both the TTS cache folder and the Neural TTS models folder before starting the upgrade.

4. Real-time and batch transcription

Version 8.0.0 introduces two new services:

  • transcribe-realtime replaces the previous streaming transcription components, with a re-architected, lower-latency pipeline.
  • transcribe-batch is a new dedicated service for high-throughput offline transcription.

Both services are deployed for every ASR language by default. You can:

  • Switch them off using the toggles in section 2.
  • Turn on GPU acceleration for them in the same way as ASR (see section 1).
  • Choose which models each one loads, per language (see section 5).

Optional real-time transcription tuning

lumenvox-speech:
  transcribeRealtime:
    enableFlowSense: false  # set to true to enable FlowSense
    europaBufferMs: "800"   # audio buffer length, in milliseconds

Upgrading from 7.x: because both services are on by default, an upgraded installation gains these deployments (and their resource usage) for every ASR language unless you switch them off. Size your cluster accordingly.

Monitoring: transcription metric names now use the transcribe_realtime_* or transcribe_batch_* prefix. Update any monitoring dashboards and alerts.


5. ASR models per language

All ASR model settings for a language now live in that language's asrLanguages entry. ASR, real-time transcription and batch transcription each load only the models listed for them, rather than every model installed. This lets you run different models, or different model versions, on each service.

Example: English with a high-definition encoder and decoder, a fine-tuned model, and real-time transcription using the high-definition encoder only

global:
  asrDefaultVersion: "8.0.0"
  asrLanguages:
    - name: "en"
      version: "8.0.0"
      fineTuned: true                  # renamed from enableFineTuned
      services:
        transcribeRealtime:
          base: false                  # don't load the base encoder
      extraModels:
        - name: "asr_encoder_hidef_model_en"
          version: "8.0.0"
          # services omitted: loaded by asr, transcribeRealtime and transcribeBatch
        - name: "asr_decoder_hidef_model_en_us"
          version: "8.0.0"
          services: ["asr"]            # only asr loads this one

Settings

SettingDescription
extraModelsAdds models (such as high-definition variants) to a language. Use services on each model to limit which of asr, transcribeRealtime and transcribeBatch load it. If services is omitted, all three load it.
services.<service>.base: falseStops that service loading the language's base encoder. The base decoder is always loaded.
fineTuned: trueDownloads and loads the language's fine-tuned model. This replaces enableFineTuned.
versionModel version for a language or an individual model. If not set, asrDefaultVersion is used.
asrDefaultVersionDefault version for any language or model without its own version. Now "8.0.0".
global.customAsrModelsNow used only for shared download packages that aren't tied to a language.

Safety check: the install fails with an error, rather than deploying a pod that repeatedly crashes, if any service ends up with no models for a language.

Upgrading from 7.x: this section contains three required changes. See items 2 to 4 in Upgrading from 7.x.


6. Istio Gateway API ingress

Version 8.0.0 adds a new ingress mode that uses the Kubernetes Gateway API with Istio. NGINX Ingress remains the default in the 8.0.0 charts and behaves exactly as in 7.x, including with custom Ingress class names such as nginx-public. No changes are needed to your NGINX settings if you continue using NGINX.

Switch to Istio Gateway API mode

global:
  lumenvox:
    ingress:
      className: "istio"               # any other value = Ingress with that class
    gateway:
      namespace: "istio-ingress"
      createIstioNamespace: true       # false if another release owns the namespace
      disableTls: false                # true when TLS is terminated at the load balancer
      tlsSecretName: speech-tls-secret
      infrastructureAnnotations: {}    # cloud load balancer annotations
      serviceExtraPorts: []            # e.g. port 443 for AWS NLB-terminated TLS

Requirements

  • Istio installed (this provides the istio GatewayClass).
  • Gateway API CRDs installed.
  • Kubernetes 1.35 or later.
  • Your TLS secret in the release namespace. The chart copies it into the gateway namespace automatically.

Configuration notes

TopicDetails
Load balancerPut your cloud load-balancer annotations in gateway.infrastructureAnnotations. Istio copies them onto the Service it creates. On Azure (provider: "azure"), the load-balancer health probe is configured automatically.
IP allowlistIn Istio mode, global.lumenvox.ingress.internalAllowlist (a list of CIDRs) restricts access to the admin portal, deployment portal, file-store, management API and reporting API. In NGINX mode, continue to restrict these routes as in 7.x, using a whitelist-source-range entry in httpAnnotations.
TLSIstio mode uses gateway.disableTls. NGINX mode continues to use ingress.disableTls, as in 7.x.

7. In-cluster databases for test and development

The new lumenvox-external-services subchart runs MongoDB, PostgreSQL, RabbitMQ and Redis inside your Kubernetes cluster. It replaces the Docker Compose setup previously used for test and development environments.

Not for production. The subchart's defaults are sized for a proof of concept. For production deployments, we continue to recommend managed or self-hosted database services. See the External Dependency Support Matrix for supported versions.

Step 1: Create the four secrets in the release namespace

kubectl create secret generic mongodb-existing-secret -n lumenvox --from-literal=mongodb-root-password=<password>
kubectl create secret generic postgres-existing-secret -n lumenvox --from-literal=postgresql-password=<password>
kubectl create secret generic rabbitmq-existing-secret -n lumenvox --from-literal=rabbitmq-password=<password>
kubectl create secret generic redis-existing-secret -n lumenvox --from-literal=redis-password=<password>

Passwords must be alphanumeric. The chart never creates these secrets for you; if one is missing, the install stops and shows the exact kubectl command needed to create it.

Step 2: Enable the subchart

global:
  enabled:
    externalServices: true

Notes

  • Redis runs as a cluster by default, which requires the OT Redis Operator to be installed first. The install checks for the operator and shows the install command if it's missing. A single-instance mode is also available; see the subchart README.
  • Storage class:set the storage class for your cluster:
    • lumenvox-external-services.storage.className for MongoDB, PostgreSQL and RabbitMQ.
    • lumenvox-external-services.redis.clusterMode.operator.storage.className for the Redis cluster (for example managed-csi on AKS).
    • For clusters with no storage class at all, the subchart can install the local-path provisioner.
  • Istio service mesh: if you use Istio as a service mesh, these database pods stay out of the mesh unless you set global.serviceMesh.istio.injectDatabaseSidecars: true.

Upgrading from 7.x

Review these items before upgrading. Each one either stops the install with a message explaining what to change, or changes behaviour without warning.

  1. Upgrade Kubernetes to 1.35 or later. Version 7.x required 1.33.

  2. Rename asrLanguages[].enableFineTuned to fineTuned. The old key stops the install with a message giving the new name.

  3. Move per-language models from customAsrModels to extraModels. In 7.x, a model listed in customAsrModels (such as a high-definition encoder) was downloaded and loaded automatically. In 8.0, only models listed under a language are loaded, so per-language entries left in customAsrModels stop the install until they are moved under their language's extraModels.

    Before (7.x):

    global:
      customAsrModels:
        - name: "asr_encoder_hidef_model_en"
      asrLanguages:
        - name: "en"
          enableFineTuned: true
    

    After (8.0):

    global:
      asrLanguages:
        - name: "en"
          fineTuned: true
          extraModels:
            - name: "asr_encoder_hidef_model_en"
    
  4. Models without a version now download at 8.0.0. asrDefaultVersion changed from "7.0.0" to "8.0.0". To keep an older model, set version explicitly on the language or model.

  5. The default image tag is now :8.0. Installations that don't set global.image.tag will move to 8.0 images.

  6. Real-time and batch transcription are on by default. An upgraded installation gains these deployments for every ASR language unless you switch them off (see section 2).

Rollbacks: since version 7.x, rolling back to a pre-7.0 version with the standard helm rollback command is not supported. See How to Roll Back from Version 7.0 to 6.x.

If you have questions about your upgrade path, particularly if you are upgrading from a version earlier than 7.0, contact Capacity Support before proceeding.


Related articles


Was this article helpful?