I'm using a TTS system that lets you design a voice from a caption. You hand it a text description of the voice plus a random seed, and it speaks in exactly that voice.

{

"input": "こんにちは。本日はお集まりいただき、ありがとうございます。",

"irodori": {

"caption": "落ち着いた知的な大人の女性の声。滑らかで聞き取りやすく、上品で信頼感のある話し方。",