Consensus Auditory-Perceptual Evaluation of Voice – Revised (CAPE-Vr)
Form content adapted from Kempster, Nagle & Solomon (2025), Journal of Voice, doi:10.1016/j.jvoice.2025.01.022, under CC BY 4.0. This tool is independent and is not affiliated with or endorsed by the CAPE-Vr authors, ASHA, or the publisher. See “About, disclaimers, and attribution” at the bottom.
Recording Conditions
Stimuli
Vowels: /ɑ/ and /i/. Sustain each for 3–5 seconds; one or more productions in typical speaking voice. /ɑ/ and /i/ (wording hidden)
Sentences:
- The blue spot is on the key again.Sentence a
- He helped her hurry home.Sentence b
- We were away a year ago.Sentence c
- I eat eggs every evening.Sentence d
- My mama makes lemon muffins.Sentence e
- Papa took a piece of the cake.Sentence f
Extemporaneous Speech: “Tell me about a place you have gone or would like to go.” standard prompt (wording hidden)
Rating Conditions
Visual analog scales
Click or drag on a line to place the mark (or type the score when numerical values are shown). Keyboard: arrow keys move by 1, Page Up/Down by 10. “Clear” returns a scale to not rated. Scores are always saved, whether or not the numbers are shown.
Inconsistencies:
(describe)Instabilities:
Additional features:
CAPE-Vr instructions: administration and scoring (click to show / hide)
Table 1. CAPE-Vr Protocol: Instructions for Administration and Scoring. From Kempster GB, Nagle KF, Solomon NP. Development and Rationale for the Consensus Auditory-Perceptual Evaluation of Voice—Revised (CAPE-Vr). Journal of Voice (2025). doi:10.1016/j.jvoice.2025.01.022. © 2025 The Authors. Published by Elsevier Inc. on behalf of The Voice Foundation. Reproduced under the Creative Commons Attribution (CC BY 4.0) license (creativecommons.org/licenses/by/4.0). The text is unchanged; it has only been formatted into sections, and the notes marked “Note for this app” have been added.
Demographics
Complete identifying information about the examinee (Name or ID, Gender, Age), examiner (Name), and date of recording as indicated at the top of the form.
Recording Conditions
Seat the examinee comfortably in a quiet environment. Place a microphone at a fixed mouth-to-microphone distance and audio record the voice and speech stimuli. Note on the form whether the stimuli were audio recorded or not; whether the examination was in person or virtual; the room environment; the recording device and/or platform; and the mouth-to-microphone distance.
Tasks and Stimuli
The three primary tasks for the CAPE-Vr may be completed in any order.
Vowels. Instruct the examinee to say the vowels /ɑ/ and /i/ using a typical speaking voice. Each vowel should be sustained for 3-5 seconds. A single trial of each vowel is adequate if the production sounds representative of that individual’s speaking voice. Modeling is discouraged to avoid imitation of the clinician’s pitch and voice quality.
Sentences. Instruct the examinee to read the following sentences aloud in their typical speaking voice. The sentences can be printed out in large, easy-to-read font. If the individual has difficulty reading, the clinician may model the sentences; in this case, check the box provided to the right of the sentences on the CAPE-Vr form.
Note for this app: there is one check box after each sentence, so modeling can be recorded for each sentence separately.
The sentences (and their primary features of interest) follow:
- The blue spot is on the key again. (English corner vowels)
- He helped her hurry home. (Word-initial /h/)
- We were away a year ago. (All voiced phonemes)
- I eat eggs every evening. (Vowel-initial words)
- My mama makes lemon muffins. (Nasal consonants)
- Papa took a piece of the cake. (High-pressure consonants)
Extemporaneous speech. Elicit at least 20 seconds of natural speech with the prompt: “Tell me about a place you have gone or would like to go.” Another prompt that elicits content unrelated to voice use is also acceptable.
Note for this app: if another prompt is used, record it in “Prompt used, if different”.
Reading passage (optional). Indicate the reading passage if one is included as part of the auditory-perceptual evaluation of voice.
Rating Conditions
Ratings are expected to be based on recordings of the voice, but the rater can indicate if they were completed on live voice in real time instead. When rating voice recordings, listen to the stimuli as many times as desired and document the number of repetitions. Also indicate: the use headphones or speakers; the use of standard samples of disordered voices as auditory anchors (ie, reference samples); the identity of the rater; and the date.
Voice Characteristics and Ratings
Attributes Rated along Visual Analog Scales (VASs)
The salient perceptual vocal attributes included in the CAPE-V and CAPE-Vr were identified by the original consensus committee authors as commonly used and easily understood: Overall Severity, Roughness, Breathiness, and Strain. These vocal attributes are generally defined as follows:
- Overall Severity: global, integrated impression of deviation from normal voice
- Roughness: perceived irregularity in the voicing source
- Breathiness: perceived air escape in the voice
- Strain: perceived vocal effort, tension, or press
Each of these perceptual attributes is accompanied by a 100-mm line forming a VAS. One blank VAS is included on the form if the examiner prefers to rate an additional attribute on a continuous scale. The words “Normal” and “Extreme” appear above the lines on the left and right sides, respectively, to indicate the direction of the perceptual ratings.
The examiner marks the degree of perceived deviance for each attribute with a small vertical line (aka a ‘tick mark’). The examiner may mark each VAS at any location.
Scoring VASs. After marking each of the VASs, measure the distance (in mm) from the left end of the line to the tick mark. Write the value in the blank space to the right of the line. Confirm that the lines are 100 mm long; if not, cross out 100 and insert the actual length of the line. Corrections can be made by dividing the distance measured by the length of the line and multiplying that result by 100.
Note for this app: measuring in millimetres does not apply here. The app scores each line directly: the position of the tick mark is recorded as a whole number from 0 (left end) to 100 (right end) and shown in the box to the right of the line.
Attributes Rated Descriptively
Options are provided for auditory-perceptual judgments of pitch, loudness, resonance, and nasality, inconsistencies, instabilities, and additional features, with room for examiner comments. Examiners may select one or more options per attribute.
- Pitch: perceived average pitch of the voice, as Normal, Low, or High. Pitch variability may be noted separately under Inconsistencies or Instabilities.
- Loudness: perceived average loudness or sound level of the voice, as Normal, Quiet, or Loud. Loudness variability may be noted separately under Inconsistencies or Instabilities.
- Resonance: perceived focus of the sound within the oral cavity, as Normal, Front (forward or “in the mask”), and Back (pharyngeal or “throaty”).
- Nasality: perceived balance of oral and nasal resonance reflecting patency of the nasal passageways and velopharyngeal function, as Normal, Hyponasal, or Hypernasal.
- Inconsistencies: Indicate and describe inconsistencies of voice according to task. If there are no notable inconsistencies, circle None. If there are, circle Present and provide a description of the inconsistencies.
- Instabilities: If relevant, circle one or more categorical descriptors of vocal instabilities from the options provided or add another term to indicate vocal instabilities.
- Additional features: If relevant, select one or more descriptors from the options provided or list alternate descriptors.
- Overall impression: State the overall severity of the voice problem and describe the voice in a few words or phrases.
Note for this app: “circle” means click the option; click it again to un-choose it.