From the Desk of Gus Mueller

You regular readers of this column might recall when audiologist Larry Humes visited the 20Q pages last October, discussing the hearing-related outcome measures selected by the NASEM committee. To refresh your memory, NASEM is an acronym for the National Academies of Sciences, Engineering, and Medicine. Several organizations, including the Department of Veterans Affairs, the National Institute on Aging, and the National Institute on Deafness and Other Communication Disorders asked NASEM to examine the state of the science in outcomes research for interventions in adult hearing health care. As a result, a committee was formed, and after much deliberation, they selected two different domains, and three representative outcome measures (see Larry’s 20Q for details).
The outcome measures selected were the RHHI (a recent replacement of the HHIE/HHIA), the APHAB (Abbreviated Profile of Hearing Aid Benefit), and the Words In Noise (WIN) test. So, are you ready to implement these? Most of you probably have used the HHIE/A, so you’re okay with the RHHI. The APHAB has been around for 30 years, and many of you have used this outcome measure too, or at least know about it. But what about the WIN test? Not too commonly used, and certainly not as an outcome measure.
While ~80% of audiologists used aided speech testing as part of the hearing aid fitting protocol in the 1970s, that changed quickly when prescriptive procedures and probe-mic verification became the standard in the 1980s. My guess is that most of today’s audiologists have never conducted aided speech testing. You now might be asking, where do I get the test, how is it conducted, when do I do this testing, where do I do this testing, what presentation levels do I use, and that’s just for starters.
And, there are other issues related to implementation of all three of the outcome measures. When is a difference (e.g., unaided vs. aided) really a difference? There are statistical differences which have been published, but our primary concern is what is the difference that indicates a true aided benefit, which is referred to as the minimal clinically important difference (MCID). Not something we often think about. But, never fear, help is on the way!
And, we didn’t pick “just anyone” to explain all this—joining us this month is someone who, over the years, has been instrumental in developing her own inventories. There’s the PAL (Profile of Aided Loudness), the PEW (Patient Expectations Worksheet), and the DIAL (Developmental Index of Audition & Listening). She also authored the germinal article describing everyone’s favorite self-assessment tool, the SOQ (Solodar’s One Question; yes, only one question).
Catherine Palmer, PhD, is Chair and Professor in the Department of Communication Science and Disorders at the University of Pittsburgh, and Professor in the Department of Otolaryngology. She also serves as the Director of Audiology for the UPMC Integrated Health System.
You all know her from her many journal articles and book chapters, and her 200-plus national and international presentations. Her extensive research has been in the areas of auditory learning post hearing aid fitting, the relationship between hearing, cognitive health and health outcomes, and matching technology to individual needs.
Dr. Palmer serves as Editor-in-Chief of Seminars in Hearing, has been on numerous committees and boards of the American Academy of Audiology, and has served as President of this organization. She is the lead presenter of the popular AAA convention Featured Session “Hearing Aids In Review,” which next April will return for the 25th consecutive year!
I suspect that the NASEM committee realizes that not all of you fitting hearing aids are going to jump in and start implementing these three outcome measures. But many of you are at least thinking about it, and some of you already have started—you all will want to keep this informative article from Catherine by your side!Gus Mueller, PhD
Contributing Editor
Browse the complete collection of 20Q with Gus Mueller CEU articles at www.audiologyonline.com/20Q
20Q: Clinical Implementation of the NASEM Outcome Measures
Learning Outcomes
After reading this article, professionals will be able to:
- Identify the measures that have been recommended by NASEM to evaluate outcomes in the domains of understanding speech in complex listening situations and hearing-related psychosocial health.
- Describe how questionnaires can practically be integrated into clinical practice.
- Describe two possible setups for efficiently completing the WIN test to serve as a treatment outcome measure.

1. I recall that Larry Humes made a visit to these pages last October, and introduced us to the NASEM outcome measures, but I hear that you are the expert on the clinical implementation?
I don’t know about the expert part, we’ll see how things go. But yes, Larry provided some great information about the process and recommendations from the committee. I hope everyone reads Larry’s 20Q before jumping into this article, so they’ll have the background on why the two outcome domains (i.e., understanding speech in complex listening situations and hearing-related psychosocial health) and how the specific tests to measure these domains were chosen (Humes, 2025). As Larry indicated, the committee recommended testing understanding speech in complex listening with two tests, the WIN (words in noise, Wilson, 2003) and the APHAB Global (Abbreviated Profile of Hearing Aid Benefit, Chisolm et al., 2005; Cox & Alexander, 1995). The committee went on to recommend using the RHHI (Revised Hearing Handicap Inventory, Dillard et al., 2025; Ventry & Weinstein, 1982) to evaluate hearing-related psychosocial health.
2. Thanks for reminding me of the specific measures, but to be honest, I haven’t really gotten around to implementing them in my clinic, so I’m really looking forward to this conversation.
I’m sure you’re not the only one who has yet to implement them. Clinicians can definitely jump right in and start using these measures, but many may be concerned about perceived barriers related to integrating new or additional measures into their busy clinical routines, so some of us on the original committee thought it might be helpful if we suggested ways to implement these recommendations.
3. Do NASEM committees usually follow up and give talks about the recommendations and include information about implementing the recommendations?
Actually, they don’t. I think in this case, a subset of the committee that included audiologists felt compelled to get the word out, and our various national meetings have been really supportive in giving us time to speak on this topic. As we organized these talks, we all believed that the information might be a lot more useful if we also talked about operationalizing the recommendations. Several of us on the committee are clinicians, myself included, so we have real-life experience in integrating the measures into a busy clinical day.
It is important to reiterate what Larry said about the scope of professional guidelines providing recommendations for implementation, but we thought it was worth getting the conversation (and action) started by talking about what we are doing in our clinics. Just as a reminder, I’m just one person in a larger group who have been working to support clinicians and researchers (audiology group includes Larry Humes, Nick Reed, Tom Powers, Colleen LePrell, and Sherri Smith) and we are just a subset of the larger group who worked on the recommendations (NASEM, 2025).
4. What do we gain by clinicians and researchers all having data on the same exact outcome measures?
Having a minimum set of outcome measures allows a clinic to benchmark its treatment success over time, across patients, and across treatments. On a larger scale, having the same data allows us to pool these results in larger databases to assess various treatments. You can imagine the same utility in research when any given study may only have a small number of participants, but if all the studies are using the same outcomes, then these data can be combined and provide a much more robust interpretation of the outcomes of any specific treatment. Having this type of data can support FDA decisions, reimbursement decisions, and demonstrate the value of the treatments audiologists provide.
Before you ask, I want to highlight why I said “minimum set of outcome measures.” These recommendations are really saying, “at a minimum, include these measures.” That does not mean you can’t include other measures. For instance, in a study of a pharmacological treatment for hearing loss, the researchers may want to include otoacoustic emissions (OAEs) to evaluate return of outer hair cell function as an outcome of their treatment. They should do this, but in addition, they should include the minimum set of outcome measures so their treatment can be compared directly to other treatments. A clinician may want to assess an individual’s perceived localization abilities and they should do that if that is one of the goals of their treatment, but they also should include the minimum set of outcome measures.
5. What do you think the possible barriers are to implementing these recommendations?
Most clinicians are very busy and immediately worry about how they will add additional measures to their routine. Others are not familiar with these particular measures and may be concerned with the time it takes to become familiar. Others may feel that they already have satisfied patients and don’t need specific outcome measures to document treatment success. Hopefully, some of the things we’ll talk about will help support clinicians in accessing these measures, identifying that they are straightforward (not a big learning curve), and provide tips for integrating them into practice.
In terms of whether you need outcome measures, I think it is important to reflect on any medical treatment you have ever received and think about whether there was an outcome measure to assess the success of the treatment. I’m willing to bet that there was. As a profession, we are not in line with expected practice in terms of outcome assessment. Assessing outcomes has face validity for your patient and importantly, supports the interventions that we recommend and becomes even more critical as we see increased interest in reimbursement for our services.
6. Did the NASEM committee think about whether clinicians and researchers could even get their hands on these measures, or whether they would need to pay for access?
Yes, one of the metrics that was used to assess candidate measures for the two domains was feasibility. This included ease of access and ease of use. All three of the measures are available at no cost and with no restrictions (open access). Before I go into any more detail, I want to share a link that takes you to a resource that provides direct access to all three measures as well as other resources to help you get set up in your clinic to use and interpret these measures. Figure 1 illustrates the directory of items that are available. Our thanks to Dr. Lori Zitelli, who created this resource and continues to update it.
Figure 1. Directory of items on the Google Drive resource for implementing the NASEM Outcome recommendations.
7. Okay, so everyone can access these measures, that’s great. How are you suggesting clinicians implement these tests to measure outcomes of hearing aid fittings?
Just as a point of clarification, the domains were identified after seeking input from a variety of community partners, including individuals with hearing loss. The question was what is meaningful to the person seeking hearing care: what is it that is of concern to them, and that they want to see improved. This is an important distinction. For instance, a researcher may care a lot about OAEs returning when applying a novel pharmaceutical treatment, but the person with the hearing loss wants to hear better in noise. So, the outcome measures reflect what matters to the patient. Importantly, measures were chosen that could be applied to any treatment, not just amplification. These measures could be applied pre- and post-auditory training, communication strategies, pharmaceutical treatment, as well as amplification.
8. Good to know, thanks. I’ll try again. How does a clinician implement these outcome measures for any hearing intervention in an adult?
We are talking about two self-report scales (APHAB Global and RHHI) as well as one objective measure (WIN). In terms of outcome assessment, the goal is to administer all three pre- and post-treatment. For the sake of this discussion, we’ll focus on amplification as the treatment, but this could be applied to any treatment that I mentioned earlier. The APHAB Global and the RHHI can be provided in paper and pencil format or on a tablet when the individual is waiting to be called back for their appointment. The downside to this is the timing of when the clinician has the information, if they want to use it as part of their discussion. The RHHI is very simple to score: 18 items that are rated yes, sometimes, or no, and a score of 4, 2, or 0 is assigned to these ratings, respectively. Easy math. The APHAB is more complicated to score. The APHAB Global has 18 items, which include the 3 sections related to hearing in various environments (the global version leaves off the section that assesses reaction to adverse sounds). As soon as this recommendation came out, Dr. Jani Johnson, who oversees the HARL Laboratory at the University of Memphis, where the original APHAB was developed by Robyn Cox, worked with her group to produce an electronic version of the APHAB Global with automatic scoring (Hearing Aid Research Lab (HARL); https://harlmemphis.org/). This is a free download. You’ll want to use this on a tablet, so it is scored for you. The APHAB is also available through the NOAH software (with automatic scoring), but this is still the longer version and does not produce the global score.
The NASEM report ended every chapter with recommendations and had a final list of things that still need to be done. One item on that list was the need for existing e-records to incorporate the APHAB Global and RHHI into their systems. This would allow the scales to be automatically sent to patients prior to their appointments to be completed, so the clinician would have the data in hand. This would be for the pre- and post-test administrations. When this is available, there won’t be any impact of these scales on clinic time (everything will be completed prior to arrival at the appointment).
Currently, the best recommendation is to put these scales on a tablet to be completed in the waiting room. The RHHI takes 2-5 minutes to complete. The APHAB can take between 5-10 minutes to complete, depending on the patient. This method also assumes a level of literacy, so you may find that for some patients, it will work better to complete the scales as an interview. If a clinic works with assistants, this might be a good role for the assistant, and then the clinician would be provided with the data to further plan treatment.
9. What does it take to actually implement these questionnaires in a practice? Can clinicians start using them right away, or is there setup involved first?
You can implement these questionnaires on Monday morning, but in terms of making them as efficient as possible, you will need to team up with your e-record administrators to have these integrated and produced in a way that allows them to be pushed out to your patients before specific appointments (to capture pre- and post-treatment). The good news is that for the large e-records used across the country, once this is possible, it will benefit anyone using that e-record. For audiology practice-specific e-records, you may find that customizing these capabilities can happen pretty quickly.
10. What about the WIN test? That was also recommended, right?
Yes, the WIN is a speech in noise test, which you can find at the link I provided earlier or through the NIH Toolbox. Researchers may be more comfortable using the NIH Toolbox version, which downloads onto an iPad for administration. Although it is technically free through the Toolbox, there is a fee for the version that provides automatic scoring. In the resources at the link we have provided, you’ll have the sound files and a scoring worksheet that significantly reduces the time for scoring.
For the clinician, the biggest question with the WIN is when and where to do the test if you plan to use it as an outcome measure. If you think about the setup when a person is having a diagnostic hearing test, it is probably cumbersome to add a soundfield test of the WIN into this routine (especially if you are not sure the person will be pursuing treatment). You can’t use your diagnostic WIN scores for the pre-test because these will have been ear-specific (under earphones) at a louder level. For busy clinics like ours, it also isn’t attractive to take a patient from a treatment room over to a booth to administer the WIN as an outcome measure (our booths are much too busy to accommodate this).
We discovered (others probably already knew this!) that we could play the WIN through our real-ear system (in our case, the Verifit). This allows us to collect the pre- and post-test data in one sitting; in our case, we are measuring this at the 3-4 week check when we have new hearing aid users come back into the clinic. The audio files and instructions you’ll need to set this up (it isn’t hard) are in the resources at the link we provided. It is important to go through the instructions so you are sure you are getting accurate measures. I want to specifically thank Dr. Michael Kurth at the University of Minnesota for helping us have accurate sound files to be used with the Verifit, as well as discussions about the presentation level. At this time, we recommend presenting this test at 70 dB SPL (the real-ear system calibrates this for you, which is convenient) since this will be loud enough to activate much of the hearing aid's signal processing but not so loud that patients can’t tolerate it. Remember, it is the noise that will stay at this level with the words getting quieter in 5-word segments (from +24 dB SNR to +4 dB SNR).
11. I have never conducted the WIN test, but I thought the recommended presentation level was much higher than 70 dB SPL?
The WIN was not developed as an outcome measure, but rather a diagnostic measure. So if you read about the WIN test administration, depending on the patient’s hearing loss, you’ll see that you are meant to start the test as high as 104 dB SPL under earphones; the goal being to achieve the best score for diagnostic interpretation. The noise stays at this level, and the speech signal decreases from a +24 dB SNR to a +4 dB SNR (this is finally 80 dB SPL for the speech). If you are using hearing aids, you will be listening in the sound field (not ear-specific), and this will be very loud.
The current recommendation would be to present the test at 70 dB SPL in the sound field, both the speech and noise signals are delivered from the same loudspeaker, located at 0-degree azimuth, unaided and then aided bilaterally (pre- and post-treatment). The WIN is literally words in noise, and the test decreases the sound level of the words, leaving the noise at the same level. The test does this automatically so the clinician isn’t adjusting the level after the start of the test. The WIN takes about 2 and a half minutes to administer. The clinician scores the responses on the score sheet and can interpret the result immediately.
12. I already conduct the QuickSIN pretty routinely. Is it okay if I simply substitute that for the WIN?
As I indicated earlier, the APHAB Global, RHHI, and WIN are the minimum set of outcome measures. A clinician certainly could measure anything in addition to this. The reality, however, is that none of us want to add measures to a busy clinical appointment. I know Larry talked about the decision to recommend the WIN over the QuickSIN in his 20Q, and the report goes into lots of detail (NASEM, 2025). There has been very limited research comparing the two tests (Wilson et al., 2007). I am hoping someone will do more work in this area; that would be ideal, then you might be able to use either.
13. Now that we’ve figured out how and when to use these tests, how do we interpret the results?
I will start by saying there is more work to be done in this area, but this should not stop people from starting to implement these measures. The challenge is that other than the APHAB, these measures weren’t designed to be outcome measures, therefore all of the data one might want in terms of interpreting change isn’t available. Let me explain. When we think about whether a score on a specific measure is different from a score obtained on the same measure at a different time, we can think about this in three ways. The terms that are used are:
- Test/retest reliability
- Minimal detectable difference (MDD)
- Minimal clinically important difference (MCID)
Test/retest reliability asks “how stable is this measure over time?” All three of these measures have test/retest reliability data, so we can know when a change is larger than what we would expect just by the fact that there is always going to be some variability in a measure.
14. OK, that’s reliability, but what is Minimal Detectable Difference and Minimal Clinically Important Difference?
Minimal Detectable Difference (MDD) provides us with the smallest difference that can be detected through statistical analysis. Similar to test/retest reliability, MDD doesn’t really get at whether this difference is meaningful to a patient. Minimal clinically important difference (MCID) tells us the smallest change in treatment outcome that a patient perceives as beneficial. Ultimately, MCID is what we want to interpret these three measures. But in the meantime, we need to use what we have; we don’t want the need for additional data to stop us from moving forward. On the other hand, we want to be ready to update our interpretations as additional data becomes available.
15. When will we have all of this information?
I don’t know that I can answer that question, but I can tell you what we know now. We have the test/retest reliability for all three measures. Chisolm et al. (2025) report a +/-17% change on the APHAB Global, indicating a change outside of test/retest. A +/- 6 change on the RHHI can be considered a change (Dillard et al., 2025). And for the WIN, a change of +/- 3 dB is outside test/retest reliability (Wilson, 2003). Larry Humes has recently published data from large data sets of APHAB (Humes et al., 2025b) and RHHI (Humes et al., 2025a) results to provide recommendations about MCID.
There are different ways to approach this question, and Humes et al. (2025a, 2025b) have based their recommendations on data that use the unaided scores as a guide. I am definitely not qualified to discuss the pros and cons of using unaided scores as a guide to providing MCIDs, but I know there are different views about this among test developers. Using unaided scores as a guide means that we’ll interpret meaningful change related to the “starting” point for a given person (the starting point being their unaided score). So, in Larry’s work, the MCID for the APHAB Global varies whether the unaided score was 12-18%, 19-37%, or 38-81% (MCID being 6%, 8%, and 12%, respectively, Humes et al., 2025b). Humes et al. (2025a) report that for the RHHI, the unaided score should be divided by 3 to establish what the required difference will be to state that there has been a MCID.
16. What about MCID for the WIN?
I believe Dr. Humes and colleagues are working on this now, but currently, you should feel comfortable with +/- 3 dB indicating a difference in scores.
17. What if I’d like to measure pre- and post-treatment subjective tinnitus perception? Are the NASEM guidelines basically saying “don’t bother”?
Not at all. If impacting tinnitus perception is important for a particular patient and treatment was targeted at this, then this is a great outcome to measure. If part of the intervention was addressing hearing, then you’ll want to include the minimum set of outcome measures recommended by the NASEM committee as well: APHAB Global, RHHI, and the WIN.
18. It seems like a busy clinician could integrate these measures into their routine with some automation, but then the question is, how do they bill for these additional activities?
CPT code 92638 is defined as Behavioral Verification and would be the appropriate code for this testing. Remember, just because something has a code does not mean it is reimbursed by Medicare or other insurers. Currently, audiologists are defined as diagnosticians, and we are not reimbursed by Medicare (most commercial payers follow Medicare) for treatment. Although there is continued work to try to change this, currently, you won’t typically see reimbursement from insurance for this code.
19. So, audiologists can’t get paid for this work?
I didn’t say that. Just because insurance doesn’t reimburse something doesn’t mean you can’t be paid. Ideally, audiologists will bill this code so it shows that audiologists are using it and doing this work. The patient will need to pay this bill. This works in an unbundled model, or if you bundle charges related to treatment, it would be ideal to list outcome assessment as part of what is included in that bundled price.
20. Any final thoughts?
Decide what you can do tomorrow and do it. If you feel you can start using the 18-item RHHI, even in paper and pencil form, when a patient arrives for their appointment (pre- and post-treatment), do it. It will take less than a minute to score, even in a paper and pencil version.
Take 30 minutes and set up your real ear system to deliver the WIN. We have found that patients really like this activity; it has tremendous face validity. They came in saying hearing in noise is difficult, and we are testing them with and without their hearing aids in noise. This is meaningful to them.
If you don’t already have some, get a few tablets for the clinic and download the APHAB Global. This will allow automatic scoring. This allows you to do the outcome measure, but the responses can also help direct further treatment when you can see where people are still having difficulty. It is also great to have tablets around so you can run a speech-to-text app for a patient who may be struggling to understand you prior to treatment.
References
Chisolm, T. H., Abrams, H. B., McArdle, R., Wilson, R. H., & Doyle, P. J. (2005). The WHO-DAS II: Psychometric properties in the measurement of functional health status in adults with acquired hearing loss. Trends in Amplification, 9(3), 111–126. https://doi.org/10.1177/108471380500900303
Cox, R., & Alexander, G. (1995). The Abbreviated Profile of Hearing Aid Benefit. Ear and Hearing, 16(2), 176–186.
Dillard, L. K., Matthews, L. J., & Dubno, J. R. (2025). Change on the Revised Hearing Handicap Inventory and associated factors: Results from a longitudinal cohort study. International Journal of Audiology, 64(5), 460–470. https://doi.org/10.1080/14992027.2024.2364197
Humes, L. (2025). 20Q: Meaningful outcome measures for hearing interventions in adults. AudiologyOnline, Article 29355. www.audiologyonline.com
Humes, L. E., Dhar, S., & Singh, J. (2025a). Some considerations in the use of the Hearing Handicap Inventory for the Elderly and its derivatives as hearing-aid outcome measures. International Journal of Audiology, 1–20. https://doi.org/10.1080/14992027.2025.2511218
Humes, L. E., Dhar, S., & Singh, J. (2025b). Some considerations for the use of the Abbreviated Profile of Hearing Aid Benefit (APHAB) as a hearing-aid outcome measure. Trends in Hearing, 29. https://doi.org/10.1177/23312165251359755
National Academies of Sciences, Engineering, and Medicine [NASEM]. (2025). Measuring meaningful outcomes for adult hearing health interventions. https://doi.org/10.17226/29104
Ventry, I. M., & Weinstein, B. E. (1982). The hearing handicap inventory for the elderly: A new tool. Ear and Hearing. https://doi.org/10.1097/00003446-198205000-00006
Wilson, R. H. (2003). Development of a speech-in-multitalker-babble paradigm to assess word-recognition performance. Journal of the American Academy of Audiology, 14(9), 453–470. https://doi.org/10.1055/s-0040-1715938
Wilson, R. H. (2007). An evaluation of the BKB-SIN, HINT, QuickSIN, and WIN materials on listeners with normal hearing and listeners with hearing loss. Journal of Speech, Language, and Hearing Research, 50(4), 844–856. https://doi.org/10.1044/1092-4388(2007/059)
Citation
Palmer, C. (2026). 20Q: Clinical Implementation of the NASEM Outcome Measures. AudiologyOnline, Article 29741. Available at www.audiologyonline.com
