Skip to main content

VPS - Voice Processing Solution

The Voice Processing Solution (VPS) from e.sigma is a powerful voice processing system consisting of 4 modules. These modules provide a speech recognition component (STT – Speech-to-Text), a translation system, a speech generation component (TTS - Text-to-Speech) and an intent classification. All modules operate offline and require no internet access, ensuring secure use without data exposure.

The combination of the various modules opens up comprehensive possibilities for the development of professional applications – for example a dialog system, the control of graphical user interfaces or natural language interfaces for training and monitoring in phraseology training.

Each module is encapsulated as a Docker container, enabling simple integration and scalability. With a respective RestAPI, all required functions are implemented and can be called transparently depending on requirements.

VPS offers the following advantages for successful deployment

  1. Real-time capable state-of-the-art, speaker-independent speech recognition
  2. Fast and simple creation of customer-specific speakers (TTS)
  3. Intention classification for robust command interface
  4. Translation system for over 20 language combinations
  5. Web-based translation application

Technical Functions and Operational Advantages

  • Simple integration through Docker containers
  • RestAPI support for all functions
  • Customer-specific training of intent classifier (xml2icsf)
    • XML description of intents/slots (xml)
    • Model training and deployment for mapping natural language to intents (icsf)
  • Support of more than 20 languages (STT, TTS, translation, classification)
  • Automated pronunciation and comprehension evaluation
  • Simple creation of non-native speakers
VPS

Products