Overview

R-CMD-check

The goal of text2speech is to harmonize various text-to-speech engines, including Amazon Polly, Coqui TTS, Google Cloud Text-to-Speech API, Microsoft Cognitive Services Text to Speech REST API, and Speechify Text-to-Speech API.

With the exception of Coqui TTS and Speechify, all these engines are accessible as R packages:

You might notice Coqui TTS doesn’t have its own R package. This is because, at this time, text2speech directly incorporates the functionality of Coqui TTS.

Speechify is called directly through the Speechify Text-to-Speech API. Create an API key at https://platform.speechify.ai/api-keys and store it in the SPEECHIFY_API_KEY environment variable (for example in your .Renviron), or pass it with tts_auth("speechify", key_or_json_file = "YOUR_KEY").

Installation

You can install this package from CRAN or the development version from GitHub with:

# Install from CRAN
install.packages("text2speech")

# or the development version from GitHub
# install.packages("devtools")
devtools::install_github("jhudsl/text2speech")

Authentication

Check for authentication. If not already authenticated, users must individually configure it for each service.

library(text2speech)

# Amazon Polly
tts_auth("amazon")
#> Warning in pollyHTTP(action = "voices", verb = "GET", query = query, ...):
#> Forbidden (HTTP 403).
#> Warning in structure(out[["Voices"]], NextToken = out[["NextToken"]]): Calling 'structure(NULL, *)' is deprecated, as NULL cannot have attributes.
#>   Consider 'structure(list(), *)' instead.
#> [1] TRUE
# Coqui TTS
tts_auth("coqui")
#> [1] TRUE
# Google Cloud Text-to-Speech API 
tts_auth("google")
#> [1] TRUE
# Microsoft Cognitive Services Text to Speech REST API
tts_auth("microsoft")
#> [1] TRUE
# Speechify Text-to-Speech API
tts_auth("speechify")
#> [1] TRUE

Voices

List different voice options for each service.

# Amazon Polly
voices_amazon <- tts_amazon_voices()
head(voices_amazon)
#>   voice         language language_code gender service
#> 1 Zeina           Arabic           arb Female  amazon
#> 2 Zhiyu Chinese Mandarin        cmn-CN Female  amazon
#> 3  Naja           Danish         da-DK Female  amazon
#> 4  Mads           Danish         da-DK   Male  amazon
#> 5 Sofie           Danish         da-DK Female  amazon
#> 6 Ruben            Dutch         nl-NL   Male  amazon

# Coqui TTS
voices_coqui <- tts_coqui_voices()
#> Warning in system("tts --list_models", intern = TRUE): running command 'tts
#> --list_models' had status 1
#> ℹ Test out different voices on the CoquiTTS Demo (<https://huggingface.co/spaces/coqui/CoquiTTS>)
head(voices_coqui)
#> # A tibble: 0 × 5
#> # ℹ 5 variables: type <chr>, language <chr>, dataset <chr>, model_name <chr>,
#> #   service <chr>

# Google Cloud Text-to-Speech API 
voices_google <- tts_google_voices()
head(voices_google)
#>                   voice language language_code gender service
#> 1      af-ZA-Standard-A     <NA>         af-ZA FEMALE  google
#> 2      am-ET-Standard-A     <NA>         am-ET FEMALE  google
#> 3      am-ET-Standard-B     <NA>         am-ET   MALE  google
#> 4       am-ET-Wavenet-A     <NA>         am-ET FEMALE  google
#> 5       am-ET-Wavenet-B     <NA>         am-ET   MALE  google
#> 6 ar-XA-Chirp3-HD-Aoede   Arabic         ar-XA FEMALE  google

# Microsoft Cognitive Services Text to Speech REST API
voices_microsoft <- tts_microsoft_voices()
head(voices_microsoft)
#>                                                                voice
#> 1   Microsoft Server Speech Text to Speech Voice (af-ZA, AdriNeural)
#> 2 Microsoft Server Speech Text to Speech Voice (af-ZA, WillemNeural)
#> 3 Microsoft Server Speech Text to Speech Voice (am-ET, MekdesNeural)
#> 4  Microsoft Server Speech Text to Speech Voice (am-ET, AmehaNeural)
#> 5 Microsoft Server Speech Text to Speech Voice (ar-AE, FatimaNeural)
#> 6 Microsoft Server Speech Text to Speech Voice (ar-AE, HamdanNeural)
#>                        language language_code gender   service
#> 1      Afrikaans (South Africa)         af-ZA Female microsoft
#> 2      Afrikaans (South Africa)         af-ZA   Male microsoft
#> 3            Amharic (Ethiopia)         am-ET Female microsoft
#> 4            Amharic (Ethiopia)         am-ET   Male microsoft
#> 5 Arabic (United Arab Emirates)         ar-AE Female microsoft
#> 6 Arabic (United Arab Emirates)         ar-AE   Male microsoft
# Speechify Text-to-Speech API
voices_speechify <- tts_speechify_voices()
head(voices_speechify)
#>     voice language language_code gender   service
#> 1    aadi    Hindi         hi-IN   male speechify
#> 2  aaliya     Urdu         ur-IN female speechify
#> 3   aamir     Urdu         ur-IN   male speechify
#> 4   abhay    Hindi         hi-IN   male speechify
#> 5 abhijit  Bengali         bn-IN   male speechify
#> 6 abirami    Tamil         ta-IN female speechify

Convert text to speech

Synthesize speech with tts(text = "TEXT", service = "ENGINE")

# Amazon Polly
tts("Hello world!", service = "amazon")

# Coqui TTS
tts("Hello world!", service = "coqui")

# Google Cloud Text-to-Speech API 
tts("Hello world!", service = "google")

# Microsoft Cognitive Services Text to Speech REST API
tts("Hello world!", service = "microsoft")

# Speechify Text-to-Speech API
tts("Hello world!", service = "speechify")

The resulting output will consist of a standardized tibble featuring the following columns:

mirror server hosted at Truenetwork, Russian Federation.