Botnoi Voice

Botnoi Voice

#374 most-used

Convert text to natural-sounding speech at scale

MarketingCommunicationProductivityDocumentsAICourses & LMS

Botnoi Voice is an AI-powered text-to-speech and subtitle generation platform from Thailand, supporting 23 languages with a broad library of voices (V1 and V2 speaker sets). Connect it to Actionist and your agents can generate narration for videos, e-learning modules, and customer communications, produce subtitles from any audio file, and automatically convert written scripts into MP3 or WAV audio without anyone opening a recording studio.

Average time saved
14 hours
per person · per month
≈ 2 workdays back

Eliminates manual work. Agents eliminate the manual cycle of recording, editing, uploading, and captioning audio content — from blog narration and training module voice-overs to IVR prompts and webinar subtitles.

Schedule

What your Botnoi Voice agent runs on autopilot

A week of scheduled jobs your Actionist agent will execute on your behalf.

28Scheduled jobs
7Agents at work
24/7Always on
Agents
Wed–Fri
Wed
Thu
Fri
7a
8a
9a
10a
11a
12p
1p
2p
3p
4p
5p
6p
Multi-app workflows

Botnoi Voice × every other app you use

End-to-end automations that span multiple apps — each one a real business outcome.

6Workflows
8Apps spanned
~32 hrsSaved / week
6Personas served
For marketing
Featured4 apps

Blog post to podcast episode, fully automated

When a new blog post is marked published in Notion, the agent fetches the article body, sends it to Botnoi Voice to generate an MP3 narration in the content team's chosen language, uploads the audio file to the podcast hosting platform, and posts the episode URL to the #content Slack channel — turning every written post into a listenable episode without any manual recording.

~5 hrs

Time saved for your team — every week, on autopilot

The flow
Trigger·When a blog post is published in Notion
Result
Generate MP3 narration via Generate Voice from Text (V2)Upload audio file to podcast folderPost episode link to #content channel
The win
Saved per run
~1 hrs
Runs / week
~5×
Every blog post becomes a podcast episode with zero recording time
Driven byMarketing Agent
ROI

Savings

What your team gets back — two angles: what you stop doing manually, and what that's worth.

Without Actionist

What you do manually today

With Actionist

What your agent runs for you

  • Sales
    90 min / week
    Manual audio recording for pitches

    Sales reps open a recording app, script and record each personalised pitch, edit out mistakes, and share the file manually — 30+ minutes per prospect.

    Sales Agent
    0 min
    Agent generates personalised audio pitches automatically

    When a deal moves to Proposal stage the agent reads the CRM record, generates a personalised audio pitch, and writes the hosted URL to the deal — ready in under a minute.

  • Marketing
    120 min / week
    Manual studio narration for each content piece

    The content team books a recording session, records narration, edits the audio, exports, and uploads — a half-day process per article that limits how often audio is produced.

    Marketing Agent
    0 min
    Agent converts every blog post to audio automatically

    When a blog post is published the agent generates a V2 narration, uploads the MP3, and posts the episode link — turning the weekly publishing batch into a dual text-and-audio release.

  • Customer Support
    0 min / week
    No audio versions of help articles

    Support articles are text-only. Customers who prefer audio or have accessibility needs have no alternative, and the support team has no scalable way to produce audio for 200+ articles.

    Customer Support Agent
    0 min
    Agent generates multilingual audio for every FAQ update

    Each time a FAQ is updated the agent generates audio in all active customer languages and writes the URLs to the knowledge base — every article has an audio version with no extra effort.

  • Human Resources
    180 min / week
    Manual narration bookings for training modules

    The L&D team books a voice-over artist for each new module, waits for delivery, edits captions manually, and uploads — a process that takes days per module and creates a content backlog.

    Human Resources Agent
    0 min
    Agent narrates and captions training modules on approval

    When a module is approved the HR agent generates V2 narration and SRT captions automatically, uploads both to the LMS, and marks the module narrated — no studio booking required.

  • Finance
    30 min / week
    Written-only financial briefings

    Executives receive dense written reports and must find time to read them. Audio summaries require a finance team member to personally record and share, which rarely happens.

    Finance Agent
    0 min
    Agent narrates weekly financial briefings automatically

    Every Monday the agent generates a professionally narrated financial audio briefing and posts it to Slack — leadership receives an audio brief they can listen to on the commute.

  • Operations
    60 min / week
    Manual webinar captioning via transcription services

    Operations submits recordings to a transcription vendor, waits 24 to 48 hours for delivery, reviews the transcript, reformats it as SRT, and uploads manually — a slow, costly process for every recording.

    Operations Agent
    0 min
    Agent generates SRT captions from audio URL within about a minute

    When a recording is uploaded the agent submits the audio URL to GenSub, receives SRT data within about a minute, and attaches it to the video record — no transcription vendor, no waiting.

  • Legal
    45 min / week
    Dense text-only compliance and policy communications

    Staff receive long compliance notices and policy updates as text documents. Comprehension is inconsistent and legal has no way to confirm staff have engaged with the content.

    Legal Agent
    0 min
    Agent generates audio policy briefs in all working languages

    Legal publishes a policy update and the agent generates audio in every regional language at a measured pace — staff can listen to compliance notices in their own language during the working day.

+ 100s of other Botnoi Voice automations
Average time saved
53 hrs / person / month
Calculator

Calculate what your team saves

Team size
8 people
Hourly rate
$35 / hr
Hours saved / week
28
Hours saved / year
1,400
Annual ROI
$49,000

Based on Botnoi Voice's typical team usage — the visible tasks plus a few other automations the agent runs: ~3.5 hrs / person / week of admin work automated.

Connect

How to plug Botnoi Voice into Actionist

Pick the connection method that suits your environment.

Connect using a Botnoi Voice API Key. Generate a key from your Botnoi account dashboard and paste it into Actionist. The key is sent as a Botnoi-Token header on every request.

1
Log in to Botnoi Voice

Go to botnoigroup.com/botnoivoice and sign in to your account.

2
Generate your API key

Navigate to your account settings and generate an API key. Copy it and store it securely — treat it like a password.

3
Paste into Actionist

In the Actionist Apps tab, find Botnoi Voice, click Connect, paste your API key, and click Test connection. Actionist runs a minimal test call to verify the key is valid.

Credentials you'll need
API Key*
Find your API key at botnoigroup.com/botnoivoice under your account settings
Actions

15 actions your agent can call

Read and write operations available to your Actionist agent.

FAQs

Questions about Botnoi Voice + Actionist

How does Actionist connect to Botnoi Voice?
Go to the Apps tab, find Botnoi Voice, and click Connect. You will be prompted to paste your Botnoi Voice API key, which you can generate from your account dashboard at botnoigroup.com/botnoivoice. Once you paste the key and click Test connection, Actionist runs a minimal test audio generation call to confirm the key is valid. The API key is stored securely and sent as a Botnoi-Token header on every request.
Which languages does Botnoi Voice support for text-to-speech?
Botnoi Voice supports 23 languages including Thai, English, Japanese, Korean, Chinese, Arabic, German, Spanish, French, Vietnamese, Indonesian, Malay, Burmese, Lao, Cambodian, Filipino, Portuguese (Brazil), Russian, Dutch, Hindi, Singaporean, Turkish, and Italian. You pass a language code with each generation request. The V1 and V2 speaker libraries both cover multiple languages, though the exact speaker options vary by version.
What is the difference between V1 and V2 speakers?
V1 speakers use the original Botnoi Voice speaker library and are reliable for a wide range of use cases. V2 speakers are the newer generation with improved prosody, more natural-sounding delivery, and additional voice characters. V2 audio is generally preferred for customer-facing content, e-learning modules, and marketing material. Both versions are available via the Actionist integration — you choose the speaker version and the specific speaker ID when configuring the generation action.
Can Botnoi Voice generate subtitles as well as speech?
Yes. In addition to text-to-speech, Botnoi Voice includes a GenSub endpoint that generates timed subtitle segments from an existing audio file URL. You provide a publicly accessible audio URL, set the maximum segment duration and silence threshold to control how captions are split, and receive timed subtitle data with optional SRT format output. This covers the reverse direction: audio-in, subtitles-out, as opposed to text-in, audio-out for TTS.
Does Botnoi Voice store the generated audio files?
If you enable the save_file option in the Generate Voice action, the Botnoi Voice API retains the audio file on its servers and returns a hosted audio URL in the response. This URL can be embedded directly in emails, web pages, or linked from a knowledge base without requiring you to handle the file upload yourself. If save_file is disabled, the API returns only the audio data stream and does not host the file.
How do API credits work with Botnoi Voice?
Botnoi Voice uses a credit system where each text-to-speech generation and subtitle generation request consumes a certain number of points. The API returns your remaining point balance and monthly credit reset value in every response, so your Actionist agent can read the balance after each call. You can configure the Operations Agent to run a credit balance check before bulk generation batches and alert the team if credits are running low before a job starts.
Can I control the speed and volume of the generated audio?
Yes. Both the speed and volume parameters are configurable per generation request. Speed accepts a value from 0 to 2 (1.0 is normal pace, 0.5 is half speed, 2.0 is double speed). Volume accepts a value from 0 to 2 (1.0 is standard output level). This lets you produce slow-paced pronunciation guides, fast-paced promo bumpers, or quiet background narration tracks — all from the same API endpoint without post-processing.
What audio formats does Botnoi Voice output?
Botnoi Voice can output audio in MP3 or WAV format. MP3 is suitable for most web, podcast, and messaging use cases. WAV is the uncompressed format preferred for broadcast, telephony (IVR prompts), and professional audio production workflows where post-processing will be applied. You specify the format with the type_media parameter in the Generate Voice action.
Get started

Connect your apps in minutes.

Start with a free instant demo, or talk to our team about a deployment designed for your business.