qwenspeak-x-5

内容来源:clawhub · 原始地址 · 查看安装指南

原始内容


name: qwenspeak description: Text-to-speech generation via Qwen3-TTS over SSH. Preset voices, voice cloning, voice design. Use when the user wants to generate speech audio, clone voices, or work with TTS. compatibility: Requires ssh and a running qwenspeak instance. QWENSPEAK_HOST and QWENSPEAK_PORT env vars must be set. metadata: author: psyb0t homepage: https://github.com/psyb0t/docker-qwenspeak

qwenspeak

YAML-driven text-to-speech over SSH using Qwen3-TTS models.

For installation and deployment, see references/setup.md.

Security & safety

  • Voice cloning requires consent. Reference audio (ref_audio) is biometric data of a real person's voice — only clone a voice you have explicit consent for. Cloning someone's voice without consent enables impersonation and fraud; never clone from audio scraped or supplied without the speaker's permission, and never use a clone to impersonate a specific named individual without their say-so. See Modes below.
  • Every call leaves your host. scripts/qwenspeak.sh execs ssh to tts@$QWENSPEAK_HOST:$QWENSPEAK_PORT and pipes text, YAML job configs, and any audio you put/get over that connection — none of it stays local. Only point QWENSPEAK_HOST/QWENSPEAK_PORT at an instance you run or explicitly trust; see SSH Wrapper below.
  • Both env vars are required — the wrapper hard-fails (QWENSPEAK_HOST not set / QWENSPEAK_PORT not set) if either is empty, so there is no silent unauthenticated fallback. Auth itself is SSH public-key only (see below).

Security model

qwenspeak is not a general-purpose shell. The instance runs inside a lockbox-hardened container, and this skill only ever talks to an instance you (or your operator) already run and trust:

  • Key-auth only — SSH accepts public-key auth only (no passwords), connecting as a restricted tts@ user. There is no interactive shell and no PTY.
  • Fixed command set — the SSH channel dispatches only the tts command plus lockbox's built-in, scoped file operations (the tables below). Anything else is refused; the remote never spawns a shell, so there is no command-injection surface.
  • Work-dir confined — every file path resolves under the instance work directory; traversal is blocked. The sandbox cannot read or write your host filesystem.
  • Consumer-only — this skill submits TTS jobs and moves files to/from a running instance. It never provisions, escalates, or installs anything on your machine (server setup is a separate, operator-side step — see setup.md).

SSH Wrapper

Use scripts/qwenspeak.sh for all commands. It handles host, port, and host key acceptance via QWENSPEAK_HOST and QWENSPEAK_PORT env vars.

scripts/qwenspeak.sh <command> [args]
scripts/qwenspeak.sh <command> < input_file
scripts/qwenspeak.sh <command> > output_file

External transmission. Every invocation sends its command, YAML job body, and any piped stdin/stdout (text, transcripts, audio bytes) over SSH to whatever QWENSPEAK_HOST/QWENSPEAK_PORT point at — data leaves your host. Point these only at a service you run or explicitly trust.

TTS Generation

Submit YAML, get a job UUID back immediately, poll for progress. Jobs run sequentially — one at a time, the rest queue up.

# Get the YAML template
scripts/qwenspeak.sh "tts print-yaml" > job.yaml

# Submit job
scripts/qwenspeak.sh "tts" < job.yaml
# {"id": "550e8400-...", "status": "queued", "total_steps": 3, "total_generations": 7}

# Check progress
scripts/qwenspeak.sh "tts get-job 550e8400"

# Follow job log
scripts/qwenspeak.sh "tts get-job-log 550e8400 -f"

# Download result
scripts/qwenspeak.sh "get hello.wav" > hello.wav

YAML Structure

Global settings + list of steps. Each step loads a model, runs all its generations, then unloads. Settings cascade: global > step > generation.

steps:
  - mode: custom-voice
    model_size: 1.7b
    speaker: Ryan
    language: English
    generate:
      - text: "Hello world"
        output: hello.wav
      - text: "I cannot believe this!"
        speaker: Vivian
        instruct: "Speak angrily"
        output: angry.wav

  - mode: voice-design
    generate:
      - text: "Welcome to our store."
        instruct: "A warm, friendly young female voice with a cheerful tone"
        output: welcome.wav

  - mode: voice-clone
    model_size: 1.7b
    ref_audio: ref.wav
    ref_text: "Transcript of reference"
    generate:
      - text: "First line in cloned voice"
        output: clone1.wav
      - text: "Second line"
        output: clone2.wav

Modes

custom-voice — Pick from 9 preset speakers. 1.7B supports emotion/style via instruct.

voice-design — Describe the voice in natural language via instruct. 1.7B only.

voice-clone — Clone from reference audio. Set ref_audio and ref_text at step level to reuse across generations. x_vector_only: true skips transcript.

Consent & privacy. ref_audio is a voice sample of a real person — treat it as sensitive personal data. Only clone a voice you have explicit consent to clone; an agent must NEVER clone a voice from audio of someone who hasn't agreed to it, and must never use a clone to impersonate a specific named individual (e.g. for fraud, harassment, or deceptive/synthetic-media purposes). If the user's request or the reference audio's provenance is unclear, ask before proceeding.

Emotion trick for cloned voices

Upload references with different emotions, use separate steps:

scripts/qwenspeak.sh "create-dir refs"
scripts/qwenspeak.sh "put refs/happy.wav" < me_happy.wav
scripts/qwenspeak.sh "put refs/angry.wav" < me_angry.wav
steps:
  - mode: voice-clone
    ref_audio: refs/happy.wav
    ref_text: "transcript of happy ref"
    generate:
      - text: "Great news everyone!"
        output: happy1.wav

  - mode: voice-clone
    ref_audio: refs/angry.wav
    ref_text: "transcript of angry ref"
    generate:
      - text: "This is unacceptable"
        output: angry1.wav

Job Management

scripts/qwenspeak.sh "tts list-jobs"              # list all
scripts/qwenspeak.sh "tts list-jobs --json"        # JSON output
scripts/qwenspeak.sh "tts get-job <id>"            # job details
scripts/qwenspeak.sh "tts get-job-log <id>"        # view log
scripts/qwenspeak.sh "tts get-job-log <id> -f"     # follow log
scripts/qwenspeak.sh "tts cancel-job <id>"         # cancel

Statuses: queuedrunningcompleted | failed | cancelled

Completed jobs auto-cleaned after 1 day, all jobs after 1 week. UUID prefixes work (e.g. first 8 chars).

File Operations

All paths relative to the work directory. Traversal blocked.

Command Description
put <path> Upload file from stdin
get <path> Download file to stdout
list-files [--json] List directory
remove-file <path> Delete a file
create-dir <path> Create directory
remove-dir <path> Remove empty directory
move-file <src> <dst> Move or rename
copy-file <src> <dst> Copy a file
file-exists <path> Check if file exists (true/false)
search-files <glob> Glob search (** recursive)

Speakers

Speaker Gender Language Description
Vivian Female Chinese Bright, slightly edgy young voice
Serena Female Chinese Warm, gentle young voice
Uncle_Fu Male Chinese Seasoned, low mellow timbre
Dylan Male Chinese Youthful Beijing dialect, clear natural timbre
Eric Male Chinese Lively Chengdu/Sichuan dialect, slightly husky
Ryan Male English Dynamic with strong rhythmic drive
Aiden Male English Sunny American, clear midrange
Ono_Anna Female Japanese Playful, light nimble timbre
Sohee Female Korean Warm with rich emotion

YAML Options

All settings cascade: global > step > generation.

Field Default Description
dtype float32 float32, float16, bfloat16 (float16/bfloat16 GPU only)
flash_attn auto FlashAttention-2: auto-detects, auto-switches float32→bfloat16
temperature 0.9 Sampling temperature
top_k 50 Top-k sampling
top_p 1.0 Top-p / nucleus sampling
repetition_penalty 1.05 Repetition penalty
max_new_tokens 2048 Max codec tokens to generate
no_sample false Greedy decoding
streaming false Streaming mode (lower latency)
mode required Step only: custom-voice, voice-design, or voice-clone
model_size 1.7b Step only: 1.7b or 0.6b
text required Text to synthesize
output required Output file path
speaker Vivian custom-voice: speaker name
language Auto Language for synthesis
instruct - custom-voice: emotion/style; voice-design: voice description
ref_audio - voice-clone: reference audio file path
ref_text - voice-clone: transcript of reference audio
x_vector_only false voice-clone: use speaker embedding only