1 · CAPE-Vr form

Consensus Auditory-Perceptual Evaluation of Voice – Revised (CAPE-Vr)

Form content adapted from Kempster, Nagle & Solomon (2025), Journal of Voice, doi:10.1016/j.jvoice.2025.01.022, under CC BY 4.0. This tool is independent and is not affiliated with or endorsed by the CAPE-Vr authors, ASHA, or the publisher. See “About, disclaimers, and attribution” at the bottom.

Recording Conditions

Audio recorded
In person / Virtual
Environment

Stimuli

Vowels: /ɑ/ and /i/. Sustain each for 3–5 seconds; one or more productions in typical speaking voice. /ɑ/ and /i/ (wording hidden)

Sentences:

  1. The blue spot is on the key again.Sentence a
  2. He helped her hurry home.Sentence b
  3. We were away a year ago.Sentence c
  4. I eat eggs every evening.Sentence d
  5. My mama makes lemon muffins.Sentence e
  6. Papa took a piece of the cake.Sentence f

Extemporaneous Speech: “Tell me about a place you have gone or would like to go.” standard prompt (wording hidden)

Rating Conditions

Live voice / Recorded voice
Headphones / Speakers
Auditory anchors

Visual analog scales

Overall Severity
Roughness
Breathiness
Strain

Click or drag on a line to place the mark (or type the score when numerical values are shown). Keyboard: arrow keys move by 1, Page Up/Down by 10. “Clear” returns a scale to not rated. Scores are always saved, whether or not the numbers are shown.

Pitch:
Loudness:
Resonance:
Nasality:

Inconsistencies:

(describe)

Instabilities:

Additional features:

CAPE-Vr instructions: administration and scoring (click to show / hide)

Table 1. CAPE-Vr Protocol: Instructions for Administration and Scoring. From Kempster GB, Nagle KF, Solomon NP. Development and Rationale for the Consensus Auditory-Perceptual Evaluation of Voice—Revised (CAPE-Vr). Journal of Voice (2025). doi:10.1016/j.jvoice.2025.01.022. © 2025 The Authors. Published by Elsevier Inc. on behalf of The Voice Foundation. Reproduced under the Creative Commons Attribution (CC BY 4.0) license (creativecommons.org/licenses/by/4.0). The text is unchanged; it has only been formatted into sections, and the notes marked “Note for this app” have been added.

Demographics

Complete identifying information about the examinee (Name or ID, Gender, Age), examiner (Name), and date of recording as indicated at the top of the form.

Recording Conditions

Seat the examinee comfortably in a quiet environment. Place a microphone at a fixed mouth-to-microphone distance and audio record the voice and speech stimuli. Note on the form whether the stimuli were audio recorded or not; whether the examination was in person or virtual; the room environment; the recording device and/or platform; and the mouth-to-microphone distance.

Tasks and Stimuli

The three primary tasks for the CAPE-Vr may be completed in any order.

Vowels. Instruct the examinee to say the vowels /ɑ/ and /i/ using a typical speaking voice. Each vowel should be sustained for 3-5 seconds. A single trial of each vowel is adequate if the production sounds representative of that individual’s speaking voice. Modeling is discouraged to avoid imitation of the clinician’s pitch and voice quality.

Sentences. Instruct the examinee to read the following sentences aloud in their typical speaking voice. The sentences can be printed out in large, easy-to-read font. If the individual has difficulty reading, the clinician may model the sentences; in this case, check the box provided to the right of the sentences on the CAPE-Vr form.

Note for this app: there is one check box after each sentence, so modeling can be recorded for each sentence separately.

The sentences (and their primary features of interest) follow:

  1. The blue spot is on the key again. (English corner vowels)
  2. He helped her hurry home. (Word-initial /h/)
  3. We were away a year ago. (All voiced phonemes)
  4. I eat eggs every evening. (Vowel-initial words)
  5. My mama makes lemon muffins. (Nasal consonants)
  6. Papa took a piece of the cake. (High-pressure consonants)

Extemporaneous speech. Elicit at least 20 seconds of natural speech with the prompt: “Tell me about a place you have gone or would like to go.” Another prompt that elicits content unrelated to voice use is also acceptable.

Note for this app: if another prompt is used, record it in “Prompt used, if different”.

Reading passage (optional). Indicate the reading passage if one is included as part of the auditory-perceptual evaluation of voice.

Rating Conditions

Ratings are expected to be based on recordings of the voice, but the rater can indicate if they were completed on live voice in real time instead. When rating voice recordings, listen to the stimuli as many times as desired and document the number of repetitions. Also indicate: the use headphones or speakers; the use of standard samples of disordered voices as auditory anchors (ie, reference samples); the identity of the rater; and the date.

Voice Characteristics and Ratings

Attributes Rated along Visual Analog Scales (VASs)

The salient perceptual vocal attributes included in the CAPE-V and CAPE-Vr were identified by the original consensus committee authors as commonly used and easily understood: Overall Severity, Roughness, Breathiness, and Strain. These vocal attributes are generally defined as follows:

  • Overall Severity: global, integrated impression of deviation from normal voice
  • Roughness: perceived irregularity in the voicing source
  • Breathiness: perceived air escape in the voice
  • Strain: perceived vocal effort, tension, or press

Each of these perceptual attributes is accompanied by a 100-mm line forming a VAS. One blank VAS is included on the form if the examiner prefers to rate an additional attribute on a continuous scale. The words “Normal” and “Extreme” appear above the lines on the left and right sides, respectively, to indicate the direction of the perceptual ratings.

The examiner marks the degree of perceived deviance for each attribute with a small vertical line (aka a ‘tick mark’). The examiner may mark each VAS at any location.

Scoring VASs. After marking each of the VASs, measure the distance (in mm) from the left end of the line to the tick mark. Write the value in the blank space to the right of the line. Confirm that the lines are 100 mm long; if not, cross out 100 and insert the actual length of the line. Corrections can be made by dividing the distance measured by the length of the line and multiplying that result by 100.

Note for this app: measuring in millimetres does not apply here. The app scores each line directly: the position of the tick mark is recorded as a whole number from 0 (left end) to 100 (right end) and shown in the box to the right of the line.

Attributes Rated Descriptively

Options are provided for auditory-perceptual judgments of pitch, loudness, resonance, and nasality, inconsistencies, instabilities, and additional features, with room for examiner comments. Examiners may select one or more options per attribute.

  • Pitch: perceived average pitch of the voice, as Normal, Low, or High. Pitch variability may be noted separately under Inconsistencies or Instabilities.
  • Loudness: perceived average loudness or sound level of the voice, as Normal, Quiet, or Loud. Loudness variability may be noted separately under Inconsistencies or Instabilities.
  • Resonance: perceived focus of the sound within the oral cavity, as Normal, Front (forward or “in the mask”), and Back (pharyngeal or “throaty”).
  • Nasality: perceived balance of oral and nasal resonance reflecting patency of the nasal passageways and velopharyngeal function, as Normal, Hyponasal, or Hypernasal.
  • Inconsistencies: Indicate and describe inconsistencies of voice according to task. If there are no notable inconsistencies, circle None. If there are, circle Present and provide a description of the inconsistencies.
  • Instabilities: If relevant, circle one or more categorical descriptors of vocal instabilities from the options provided or add another term to indicate vocal instabilities.
  • Additional features: If relevant, select one or more descriptors from the options provided or list alternate descriptors.
  • Overall impression: State the overall severity of the voice problem and describe the voice in a few words or phrases.

Note for this app: “circle” means click the option; click it again to un-choose it.

2 · Acoustic analysis with Praat optional upload recordings → make the Praat package → run it in Praat → import the results

Recordings

Drag WAV files here or use “Add WAV files…”. Each file gets a CAPE-Vr task, suggested from its file name; please check it. The app only reads each file’s duration, sample rate, channels, and bit depth. All acoustic measures come from Praat. The audio is not stored in the session file: after loading a session, add the same files again to play or export them.

Speaker and analysis settings

Speaker profile

Choosing a profile fills in its pitch and formant values; you can change them. These values only set the ranges Praat searches in (analysis settings); they are not norms or expected values for the speaker.

Advanced analysis settings (optional)
CPPS method
CPPS on speech: voiced parts only
Intensity on speech: voiced parts only
Intensity averaging
HNR method
Jitter measures
Shimmer measures
Jitter/shimmer period range
Formants
Stereo files

Every setting is written into the package and echoed in results.json, and the report’s methods paragraph lists the settings used.

Acoustic results (from Praat)

Make the Praat package

3 · Documentation tools optional clinical summary, EMR draft, CSV, clinic-defined rules

Documentation

Statements from your clinic’s rules (ticked statements are included in the Clinical Summary and EMR draft; click “why?” to see the rule and value)
Clinical Summary
EMR Documentation Draft
Interpretation Rules (clinic-defined; all start empty)

Clinic-defined rules, not part of the CAPE-Vr. Categories and reference statements here come only from rules your clinic sets. The CAPE-Vr removed the mild/moderate/severe labels from the rating form on purpose, because such labels act as anchors and bias ratings (Kempster, Nagle & Solomon, 2025, Modification 1). No published cut-offs are implied, and any example numbers are placeholders only. Nothing here is a diagnosis. Every generated line shows the value and the rule behind it.

4 · Report

Generate and print report

The report is the CAPE-Vr form filled in, followed by the acoustic results and methods (if imported), and the Clinical Summary and EMR draft. It opens as a print view: use “Print / Save as PDF” there.

Rules that were switched off

Documentation and reporting aid. It does not diagnose and does not replace clinical judgment.

About, disclaimers, and attribution