Abstract
Voice cloning technology has developed rapidly, and current synthesis techniques can produce humanlike voices. Recent research has shown that cloned voices can now be more intelligible than their human originals in background noise, a signal-extrinsic degradation. Here, we aimed to establish whether this intelligibility benefit for voice clones degraded using noise-vocoding, a signal-intrinsic manipulation that severely degrades fine-grained spectrotemporal and harmonic information in the speech signal. We also investigated whether listeners can perceptually adapt to noise-vocoded voice clones. We compared the intelligibility of ten six-band noise-vocoded voice clones with their human originals. Listeners heard 80 sentences, in two counterbalanced blocks: a block of 40 sentences by voice clones and a block of 40 sentences by human, to compare perceptual adaptation to both speech types. Listeners showed 13.2% higher accuracy for voice clones. Overall, listeners adapted to the noise-vocoded speech and improved by around 16% over the course of 40 sentences in the first block only, regardless of whether voices were human or cloned. The intelligibility benefit for cloned voices thus persisted for noise-vocoded speech, and listeners adapted to noise-vocoded cloned voices as much as for noise-vocoded human voices. Our results have implications for applications of voice clones, in particular in future development of assistive listening technologies.
Introduction
Voice cloning uses Artificial Intelligence (AI) to resynthesise an individual’s voice, aiming to capture their voice source, intonation, pitch, accent, and articulation patterns. This technology can generate lifelike voice clones based on as little as ten seconds of audio (Patel et al., 2024), using deep learning models. Voice cloning technology has developed rapidly, and cloned voices have humanlike quality (Farid, 2022; McGettigan et al., 2024). Potential future applications of this technology include assistive communication for individuals with speech impairments (for example, preserving voices of those with degenerative conditions or following cancer treatment), assistive hearing technologies, enhancing AI agents, public announcement systems, and improving dubbing in television, film, and online media, or creating realistic AI-generated voices for research, education and entertainment.
Recent research has seen a change from a focus on scrutiny of potential misuses of voice cloning technology (Amezaga & Hajek, 2022; Farid, 2022; Hutiri et al., 2024; McGettigan et al., 2025) to also include the consideration of its potential benefits, such as potential greater intelligibility of cloned voices (Adank & Wang, 2026; Calandruccio et al., 2025). Yet, most past intelligibility research focused on the relative intelligibility of general synthetic voices (Clark, 1983; Cooke et al., 2013; Ibelings et al., 2022; Ma & Tang, 2024; Nixon & Anderson, 1986; Pisoni et al., 1985), while only two studies thus far have evaluated the intelligibility of cloned voices (Adank & Wang, 2026; Calandruccio et al., 2025). Older studies showed that synthetic speech was less intelligible than speech produced by humans, especially in degraded listening conditions (Clark, 1983; Cooke et al., 2013; Nixon & Anderson, 1986; Pisoni et al., 1985). Yet, as Text-To-Speech (TTS) synthesis technology has improved, synthetic voices have started to show intelligibility benefits (Ibelings et al., 2022; Ma & Tang, 2024). In their Experiment 2, Ibelings et al. report on male- and female-sounding voices from three different TTS systems (Acapela US, Acapela DNN, Google Wavenet) and compared these to natural voices. They showed that TTS speech yielded slightly better intelligibility scores (1.2dB lower Speech Reception Threshold, SRT), and they concluded that their results demonstrate the viability of TTS for creating efficient, high-quality and intelligible speech materials. Ma & Tang evaluated the intelligibility of speech generated by three commercially synthesised voices (Google TTS, Amazon’s Polly, and Microsoft’s Azure) for normal-hearing and hearing-impaired listeners, in various noise conditions. Compared to the human female baseline voice, listeners’ performance suggested that the synthetic voices were more intelligible even in adverse listening conditions, for the normal-hearing listeners but not for the hearing-impaired listeners.
Adank & Wang used ten cloned voices, based on ten human voices from various accent regions in England and thus used a stimulus set reflecting more idiosyncratic speaker-specific variation. They presented listeners with 80 phonetically balanced sentences (IEEE, 1969) at four signal-to-noise (SNR) levels (+3dB, 0dB, -3dB and -6dB) in an online experiment. They report that cloned voices were more intelligible than their human counterparts; participants showed on average 13.4% higher word recall accuracy for the cloned voices. Individual voices varied in their relative intelligibility benefit for cloned voices; the average benefit ranged between 4.9% and 22.3% per voice. Adank & Wang also conducted an extensive acoustic analysis on the 80 sentences, to evaluate what acoustic characteristics differed between the cloned and human speech. They measured 47 acoustic factors, grouped broadly into spectrotemporal measures including formant measurements, objective intelligibility measures, and voice quality measures including jitter, shimmer, and harmonic-to-noise ratios in several frequency bands between 0 and 3,500Hz. They reduced 38 non-collinear acoustic measures into five rotated Principal Components (PCs) using Principal Component Analysis (PCA), then ran a Linear Discriminant Analysis (LDA) to determine which components best classified sentences as cloned or human. The five PCs reflected vocal stability, harmonic-to-noise variability, spectral tilt and energy, mid-frequency energy, and formant structure. Using these PCs, the LDA classified speech type with 78.1% accuracy. Vocal stability and harmonic-to-noise measures contributed most to the separation, with cloned voices showing smoother source characteristics, i.e., lower jitter and shimmer and higher harmonicity in the 500-3,500Hz range. Adank & Wang also related the acoustic factors to the intelligibility results for the human and cloned voices separately. For the human voices, a regression model was constructed using 38 non-collinear acoustic measures, which accounted for 32% of the variance in intelligibility scores. The most influential features were utterance duration, fundamental frequency (pitch), harmonic-to-noise ratio (voice clarity), amplitude stability, and mid-frequency formant structure. For cloned voices, a model was constructed using 38 acoustic measures, explaining 23% of the variance. Here, the strongest predictors were fundamental frequency, overall harmonic-to-noise ratio, high-frequency spectral balance, and measures of temporal organisation (e.g., utterance duration). Notably, features describing very fine-grained voice instability (e.g., jitter and shimmer) were not selected as important predictors for cloned voices. These results indicate that the intelligibility of voice clones may be governed more by global spectral and temporal properties than by the micro-scale vocal perturbations that are relevant for natural human speech.
Here, we aim to evaluate whether the intelligibility benefit for cloned voices persists for noise-vocoded speech (Shannon et al., 1995a), which degrades harmonic information and fine-grained spectrotemporal detail. A noise-vocoder uses the amplitude envelopes of the speech signal to modulate the corresponding bands of a carrier signal (McGettigan et al., 2014). The typical configuration involves restricting the overall spectral range (e.g., up to 5,000-7,000Hz) and applying a bank of bandpass filters, often with equally spaced centre frequencies on a logarithmic scale to reflect cochlear frequency mapping. Within each band, the signal is half-wave rectified and low-pass filtered (typically with a cutoff of 30-50Hz, although higher cutoff values, e.g., up to around 300Hz, are also used to preserve temporal envelope fluctuations (Dorman et al., 1997; Souza & Rosen, 2009; Whitmal et al., 2007) to extract the temporal envelope, which modulates a carrier noise restricted to the same frequency band. The vocoding procedure thus removes harmonic detail while preserving low-frequency amplitude modulations and temporal information. The fidelity of the degraded speech is critically dependent on the number of analysis-synthesis bands, the envelope cutoff frequency, and the filter slopes. Lower band counts (e.g., 4 or 8 bands) produce a signal sparse in spectral definition while temporal information remains largely present, while higher counts (e.g., 16 or 32 bands) progressively restore spectral definition. Noise-vocoding has been used to simulate the speech processing of cochlear implant users, who receive similarly limited spectral information via a small number of electrodes (Faulkner et al., 2000; Rosen et al., 1999). Noise-vocoding is sometimes characterised as a distortion intrinsic to the speech signal, whereas added speech-shaped noise (as used in Adank & Wang) can be regarded as a signal-extrinsic distortion (Mattys et al., 2012, 2025). In normal-hearing listeners, intelligibility increases logarithmically with the number of bands meaning that performance over bands increases more for lower numbers of bands, and asymptotes near natural speech intelligibility at around 16-24 bands depending on the task and listening conditions (Shannon et al., 1995a, 2004). Thus, noise-vocoded speech can serve as a model for a speech signal that omits fine-grained spectral and harmonic detail. If the intelligibility benefit for cloned voices persists for noise-vocoded speech, then it suggests that this benefit does not predominantly rely on the enhanced spectral and harmonic (voice-source) features in the cloned voices as such, but on features of the cloned voices that remain in the signal after noise-vocoding. Moreover, if cloned voices are more intelligible after noise-vocoding, then this result would imply that the intelligibility benefit extends to intrinsic degradations.
Moreover, the use of this type of degradation allows us to examine whether listeners can perceptually adapt to cloned noise-vocoded speech (Davis et al., 2005; Drouin & Flores, 2024; Hervais-Adelman et al., 2008, 2011; Huyck & Johnsrude, 2012; McGettigan et al., 2014; Wang et al., 2023). Listeners tend to improve their speech perception performance after short-term exposure (i.e., 4-10 sentences) of noise-vocoding. Overall, over the course of 40-80 sentences, listeners can show an improvement of up to 14% in correct word recall (e.g., (Wang et al., 2023). We were aiming for a similar amount of perceptual learning for our speech stimuli, and therefore we decided to replicate the vocoder parameter settings from Wang et al. (2023) in the current study.
By examining whether and how perceptual adaptation to noise-vocoded speech differs for cloned versus human voices, we can evaluate whether the perceptual strategies are specific to harmonically and spectrally rich (un-vocoded) speech or extend to perceptually challenging, spectrally degraded, conditions.
Finally, we aimed to determine whether the order of exposure to human or cloned noise-vocoded speech would affect adaptation, a concept known as a transfer of learning effect (Adank & Janse, 2009; Banai & Lavner, 2019; Bradlow & Bent, 2008; Delhommeau et al., 2002; Haskell, 2001; Loebach et al., 2009; McClaskey et al., 1983; Perkins & Salomon, 1992). Transfer of learning refers to the application of prior skills or knowledge learned in one context to another context. Prior learning can facilitate performance on a new task or domain, an example of positive transfer. However, such learning can also inhibit learning on a new task, an example of negative transfer. Finally, prior exposure could also have no effect, an example of zero transfer, indicating that the learning was highly specific to the task or to the context. For example, Adank & Janse evaluated transfer of learning effects in perceptual adaptation for artificially time-compressed fast speech and natural fast speech. They tested two groups of listeners on a speeded sentence verification task, in which they had to decide as quickly as possible whether a sentence (e.g., “Tomatoes grow on plants”) was semantically true or false. The first group verified the natural fast before the time-compressed sentences, while the second group verified the time-compressed before the natural fast sentences. They found that positive transfer of learning when the (more intelligible) time-compressed sentences preceded the (less intelligible) natural fast sentences, but zero transfer of learning when natural fast sentences preceded the time-compressed sentences. If the intelligibility benefit persists for noise-vocoded cloned speech, then, based on the pattern in Adank & Janse’s results we might expect that transfer of learning effects depends on which of the conditions turns out to be more intelligible. Specifically, we expect that we would find a similar positive transfer of learning effect if the cloned speech precedes the human speech and we would predict the opposite effect if the human vocoded speech is more intelligible, following Adank & Janse’s finding that a positive transfer of learning effect occurs when a more intelligible speech type precedes a less intelligible speech type.
The Current Study
Participants listened to six-band noise-vocoded human and cloned versions of 80 IEEE sentences (40 human, 40 cloned) for ten voices in an online study, in a between-subjects design. In Order 1, participants first listened to 40 sentences produced by the human voices, followed by 40 sentences produced by the cloned voices. In Order 2, the cloned voices preceded the human voices. We selected six-band vocoding following Wang et al. (2023), who showed that listeners improved their keyword accuracy by around 14% on average across 40 sentences for their single task condition in their Experiment 1. We predicted, first, that the intelligibility benefit for cloned voices would be retained, albeit potentially reduced due to severe degradation of harmonic and fine-grained spectral information in the speech signal, based on results in Adank & Wang. Second, we predicted that listeners could perceptually adapt to the human and to the cloned noise-vocoded voices. Third, we expected positive transfer of learning between the cloned voices and the human voices, depending on whether the human or the cloned voices were more intelligible, with a positive transfer or learning effect predicted between the more intelligible and the less intelligible speech type. If there was no difference between the human and cloned voices, we predicted zero transfer of learning.
Methods
Participants
The study was conducted online; participants were recruited on Prolific (prolific.co) and the experiment was hosted on Gorilla (gorilla.sc). We used Bayes’ stopping rule (Rouder, 2014) to determine our sample size, per Adank & Wang. We set our minimum sample size to 80 participants following Adank & Wang. After this minimum sample was collected, we used Bayes Factors
The Bayesian Information Criterion (BIC) measure of the fitted model (H1) was extracted with the BIC() function and was compared against the same measure from the model excluding the interaction term (H0) to obtain the
Participants were between 18-35 years old, native speakers of British English, lived in the United Kingdom while completing the experiment, were raised monolingually, had completed a minimum of five studies on Prolific with an approval rate of >= 90% and were required to run the study on a computer, using Chrome, with wired headphones. 148 participants initiated the experiment, 48 failed the headphone screening (Milne et al., 2021, cf. Procedure), 15 did not provide consent, three were over-recruited due to an error in Prolific (their data was not retained in the experiment, but they were remunerated for their time). Two participants were replaced due to poor performance on the intelligibility task; their average accuracy was more than 2SD below the group mean. The final sample consisted of 80 participants (40 female, 40 male, average age 28.3 years, SD 4.6 years, range 19-35 years old). Participants received payment commensurate to £10 per hour. The study was approved by the Research Ethics Committee of University College London (UCL) and conducted under Project IDs #0599.001/005.
Materials
We used the same materials as used by Adank & Wang, who used materials from a publicly available database: the ARU (Acoustics Research Unit) corpus recorded at the University of Liverpool, United Kingdom (Hopkins et al., 2019). This corpus comprises single-channel recordings of all 720 IEEE sentences (IEEE, 1969) spoken by 12 adult native British English voices (ten from England, two from Wales) recorded in an anechoic chamber using a Type 4190 microphone on a Brüel & Kjær 2669 preamplifier (sampling at 65,536Hz), connected to a B&K Nexus conditioning amplifier, connected to a Type 3160-A B&K LAN-XI multipurpose generator module, linked to a computer running B&K Pulse Time Data Recorder v20. All individuals included gave consent for their recordings to be made publicly available. We included the ten voices from England; six female and four male voices, between 21 and 47 years of age (see Table S1 in Supplemental Materials). They all spoke British English as a first language and completed primary and secondary schooling in the UK. They also had accents that were not strongly regional and were screened to ensure they met this requirement. None of them smoked nor did they have any history of hearing or speech impairments, and they all had good hearing at the time of the recordings as evidenced by a pure tone audiometry test in an audiometric booth according to BS EN ISO 8253-1:2010. All had thresholds of 20dB HL or better (age adjusted) at frequencies ranging from 125Hz to 8kHz. Further specifics on the recording conditions are detailed in (Hopkins et al., 2019).
We selected 240 IEEE sentences from the ARU corpus, 120 test sentences, and 120 training sentences. We used Praat (Boersma & Weenink, 2018) to resample them to 22.05kHz, single channel (mono) and saved them as.wav files. The 120 training sentences were uploaded into ElevenLabs (https://ElevenLabs.io) to create a voice clone for each voice in December 2024. ElevenLabs’ “Voice Design” model was used to create the cloned voices, using the default settings across all voices, except for Voice 11 in the ARU corpus (cf. Supplemental Table S1). For Voice 11 the Signal-to-Noise Ratio was adversely affected in the generated sentences, and we therefore repeated the cloning process with ElevenLabs’ setting “noise suppression” on. The final 2,400 files were subsequently saved as.wav files and sampled to 22.05kHz, 65dB. We selected the first 80 of the 120 sentences generated for each voice for inclusion in the main experiment as in Adank & Wang. The final set of stimuli of the experiment thus consisted of 1,600 sentences (10 human voices, 10 cloned voices, and 80 sentences per voice).
Speech stimuli were subsequently processed using a noise-excited channel vocoder implemented using custom Matlab (Natick, MA, USA) scripts. The vocoder parameters were chosen to match those used in Wang et al. (2023), where they were established during piloting to yield intelligible yet degraded speech with performance away from floor and ceiling, allowing perceptual adaptation to be observed. Each speech signal was first bandpass-filtered into six contiguous frequency channels spanning 50 to 5,000Hz, with channel centre frequencies and bandwidths distributed according to the Greenwood cochlear function (Greenwood, 1990), using sixth-order Butterworth filters. Within each band, the temporal envelope was extracted via half-wave rectification and low-pass filtered using a 4th-order Butterworth filter with a cutoff frequency of 300Hz, applied with zero-phase forward-backward filtering. Each band’s amplitude envelope was extracted using a low-pass filter (cut-off at 300 Hz) followed by rectification. This envelope was used to modulate a white noise, which was then filtered by the same band-pass filter used to extract the envelope, before all the band outputs were summed together. The choice of a 300Hz cutoff avoided overly severe degradation and ensures that performance would not be at floor level. Each band’s amplitude envelope was extracted using a low-pass filter (cut-off at 300 Hz) followed by rectification. This envelope was used to modulate a white noise, which was then filtered by the same band-pass filter used to extract the envelope, before all the band outputs were summed together. This process ensured that the output of each channel remained spectrally confined to its original analysis band. The individual channel outputs were summed, and the composite waveform was amplitude-normalised to the root-mean-square level of the original bandpass-filtered speech. Finally, using Praat, the intensity of each of the 1,600 files was scaled to 65dB and saved as the highest quality.mp3 files (101kbps, 16-bit) available in Praat. Eighty sentences were included in the experiment, 40 per speech type (Human vs. Clone), 4 per voice, with all factors counterbalanced across participants in a counterbalanced Latin-square design.
We also included a two-alternative forced choice (2AFC) task for listeners to complete after the main experiment, as in Adank & Wang, to evaluate whether the noise-vocoding procedure affected participants’ ability to determine whether a sentence was spoken by a human voice. In this 2AFC task, participants listened to a single sentence repeatedly, “Pack the records in a neat thin case”. This sentence contained three point vowels for British English, namely,/ɑː/in pack,/ɔː/records, and/i:/in neat, and spoken by all voices (ten human and ten cloned).
Procedure
The experiment was conducted on an online testing platform, Gorilla.sc, (Anwyl-Irvine et al., 2020). Participants read the information sheet before giving consent. They were then asked to turn on their browser’s audio auto-play. Participants were instructed to use wired headphones, and passed a headphone screening (Milne et al., 2021). Finally, they were presented with a speech sample, which they could replay to adjust their volume to a comfortable level.
Before starting the main experiment, participants practised the task with three sentences not included in the main task, spoken by three of the voices (two human and one cloned), at a highly intelligible level of vocoding (for instance, more than 16 bands), to expose listeners to the sound of vocoding and the task. Next, the main experiment started and consisted of two blocks of 40 sentences with an option to take a two-minute break in between (Figure 1). The task had a between-group design. In the first group (Order 1, human before clone), participants first listened to 40 sentences produced by the human original voices, followed by 40 sentences produced by the cloned voices. In the second group (Order 2, clone before human), this order was reversed, with the cloned sentences preceding the human voices. The cloned and human version of the same sentence were not presented to the same participant. Eighty unique sentences were presented, and sentence order was randomised within each block. Each participant heard the same 80 sentences, spoken by different voices and in the cloned or in the human condition. Participants were told that the experiment would progress automatically after the response period for each sentence had passed. After each sentence had been presented, participants had 25 seconds to type in what they had perceived and were asked to always give a response. Each sentence had five key words. Design of the main task, which used a between-group design with Order counterbalanced across both speech types (human and clone). The break lasted up to 2 minutes, 2AFC: two-alternative forced choice.
The intelligibility task was followed by a 2AFC human-clone identification task (2AFC task). Sentences were presented vocoded in the 2AFC task. Before starting the task, participants were informed that the experiment included human and cloned voices. Second, they heard two versions of the same sentence and asked to decide whether the first or second sentence was spoken by the human voice, in 20 trials. Participants were to respond within five seconds, before the task progressed automatically. After the experiment, participants could provide feedback. On average, participants took 30 minutes to complete the study.
Data Processing and Analysis: Intelligibility Task
The proportion of correctly recognised key words for each sentence was the dependent measure for the intelligibility task. Words with incorrect suffixes (e.g., -s, -ed, -ing) were scored as correct, but words (including compound words) reported in part (e.g., ‘raindrops’ instead of ‘raincoats’) were scored as incorrect (Adank & Wang, 2026; Wang et al., 2023). Trials without a response were coded as 0% correct. Accuracy was scored blinded for condition.
Participants’ sentence recognition accuracy, measured as the number of correct words per sentence (0–5), modelled as a proportion between 0-1, was analysed using GLMMs implemented by the mixed() function in the afex R-package (version 1.4-1; (Singmann et al., 2023). Accuracy was modelled as a binomial outcome with a logit link function. The fixed-effects structure included Type (Human vs. Clone), log(Trial) (running between 1-40 per block), Order (1 or 2), and their three-way interaction, allowing us to examine both adaptation effects over trials and potential transfer effects across both orders.
The initial model specified random intercepts and log(Trial)-by-Type random slopes for Participant, as well as Type-by-Order random slopes for Sentence (i.e., maximal random-effects structure justified by the experimental design; Barr et al., 2013). To select an optimal fitting model for our data, we first removed random effects that caused a convergence failure. Next, we excluded the random effects whose inclusion yielded inaccurate estimates of the raw responses, a sign of over-fitting (Nannen, 2003). Lastly, we applied a backward model-selection procedure using the anova() function in the afex package, which compared the goodness-of-fit (i.e., maximised log-likelihood) of two models given the data while penalising for the complexity of the models. Each time we performed a comparison between a model and a simpler model excluding a certain random effect and removing the effect where it did not significantly contribute to the model fit. We continued such comparisons until we identified the best-fitting model. The best fitting model for sentence recognition accuracy included random intercepts for Participant and random slopes for log(Trial) and Type per Participant.
To analyse the three-way interaction of log(Trial), Type, and Order on speech accuracy, the Analysis of Variance (ANOVA) tables were generated for the GLMM with the nice() function in the afex package. Here, hypothesis testing on the main effects and interaction were conducted via a chi-square model comparison on the log likelihood between a simpler model and the model having one more fixed-effect term in each step. Estimated marginal means at each level of each factor were obtained by the emmeans() function, pairwise categorical contrasts were estimated by the pairs() function, and pairwise contrasts for log(Trial) were estimated by the emtrends() function, all under the emmeans R-package (Version 1.11.1, Lenth, 2023). The p-value threshold was adjusted by Bonferroni’s method for pairwise comparisons.
Data Processing and Analysis: 2AFC Task
The correctness of response (1 or 0) was the dependent measure for the 2AFC task. Here, we fitted a GLMM with the glmer() function from the lme4 R-package (version 1.1-37; Bates et al., 2015). The model included the intercept as a predictor to examine whether responses were significantly different from chance performance (0.5) and included random intercepts for Participant and Voice.
Results
Intelligibility
Descriptive Statistics for Speech Recognition Accuracy by Type and Order
Note. N: number of samples; SD: standard deviation; SE: standard error; CI: 95% confidence interval.
ANOVA Table for the Three-Way Effects of log(Trial), Type, and Order on Speech Recognition Accuracy
Note. df: degree of freedom; Chisq: Chi-squared statistics.
*p<.05, **p<.01, ***p<.001. Significant effects were boldened.

GLMM-estimated percent of correctly recorded keywords per trial of the intelligibility task for cloned (orange) and human voices (blue) presented for Order 1 (top panel) and Order 2 (bottom panel), separated by Block (left panel Block 1, right panel Block 2), for the 40 trials per block. Filled areas represent 95% confidence intervals. Points denote the raw mean accuracy on each trial. Error bars indicate standard error of the mean.
2AFC Responses
Data from one participant was not recorded due to an error in Gorilla.sc, so the final number of data points was 1580. Performance in the 2AFC task did not differ reliably from chance. A GLMM with random intercepts for participants and voices showed that accuracy was not significantly different from chance (β (SE) = -0.24 (0.12), p=.054). This result suggests that listeners could not discriminate human and cloned voices. Overall participants correctly identified the human voice of the pair in 44.4% of trials, showing a weak bias towards classifying voices as cloned instead of as human.
Discussion
Main Findings
We aimed to establish if the intelligibility benefit previously found for cloned speech persists for noise-vocoded speech, if listeners can perceptually adapt to cloned noise-vocoded speech, and determine to which extent the order of exposure to human or cloned noise-vocoded speech affected adaptation. We predicted that the intelligibility benefit for cloned speech would persist for noise-vocoded speech, that listeners would adapt to both speech types equally, and that intelligibility of human voices would be improved if listeners were exposed to cloned speech first. The results confirmed all but the third prediction (i.e., transfer of learning). First, cloned voices retained their intelligibility benefit and listeners showed 13% higher accuracy for voice clones on average compared to human voices (across both orders). Second, listeners adapted to the human and cloned noise-vocoded voices to a similar extent, as both groups (Orders 1 and 2) improved ∼16-17% over the course of the first block of speech. Participants improved equally for the first blocks of both orders, adapting a similar amount for both types of voices. We predicted that Order 2 (Clone → Human) would lead to a positive transfer of learning effect based on Adank & Janse. Adank & Janse found that participants showed better speech perception for natural fast speech when preceded by artificial time-compressed speech (which was easier to recognise), but this prediction was not borne out. Instead, we found a small negative transfer of learning effect from human to cloned voices; cloned speech was 2.2% less intelligible when preceded by human speech. This transfer of learning effect did not support our hypothesis that predicted a positive transfer of learning between the cloned voices and the human voices, if the cloned voices retained their intelligibility benefit after noise-vocoding. This result indicates that it may not be sufficient for a speech type to be easier to result in a positive transfer of learning effect.
Finally, listeners were unable to reliably decide which voice was human, as evidenced by the 2AFC task, as they scored 44% accuracy (chance level was 50%). This result contrasts with Adank & Wang, who found that their listeners could correctly identify the human voice out of a pair in 70% of cases, but it should be noted that they used clear, undegraded, speech in their 2AFC task. It appears that noise-vocoding successfully removed all (harmonic and fine-grained spectral) acoustic information necessary for listeners to decipher which voice was human in this comparison.
Links to Previous Studies
The overall difference in intelligibility between the speech types, with cloned voices showing 14% higher accuracy, mimics results reported in Adank & Wang, who report a comparable intelligibility benefit for cloned voices in speech-shaped noise of 13%. It therefore appears that removal of nearly all harmonic information and fine spectral detail did not affect the relative intelligibility of cloned voices. Thus, it seems unlikely that the intelligibility benefit for cloned voices relies entirely on source-related characteristics such as pitch, pitch variation, and voice smoothness (jitter, shimmer), and overall harmonic-to-noise ratio of the voice, as suggested by the results of Adank & Wang. Our results thus suggest that the advantage may reflect more robust or more predictable global spectrotemporal and amplitude patterns that remain accessible even when speech is reduced to coarse envelope cues. These cues might have also been present in the cloned voices before vocoding, as our results are consistent with recent work showing that modern neural vocoding methods used in TTS systems (and thus in voice cloning) are optimised to reconstruct intelligible speech from compressed representations (Li et al., 2025). Neural vocoding refers to the process of converting an intermediate acoustic representation (typically a mel-spectrogram) into a waveform using a trained neural network, which learns statistical regularities of speech and reconstructs perceptually relevant structure (Wang et al., 2017). As a result, the generated speech signal may contain redundancies or patterns that support intelligibility even after substantial degradation. This aspect of TTS may contribute to the observed robustness of the intelligibility advantage for cloned speech under noise-vocoding, although the precise mechanisms remain to be determined. Alternatively, listeners could have focused their attention on different cues during perception of noise-vocoded cloned speech that remained accessible, such as fine-grained temporal information.
We did not replicate results previously reported for transfer of learning between speech types (Adank & Janse, 2009; Banai & Lavner, 2019; Bradlow & Bent, 2008; Loebach et al., 2009). Perceptual learning of cloned vocoded speech appeared to not facilitate recognition of human vocoded sentences. Instead, perceptual adaptation to human speech slightly impaired recognition of cloned vocoded speech. Our results suggest that it is not sufficient for an easier speech type to precede a more difficult speech type to generate positive transfer of learning between two speech types. Perhaps the lack of this transfer of learning effect is due to the fact that the cloned sentences were the result of text-to-speech synthesis, whereas the human sentences were not. Adank & Janse re-synthesised both their stimulus types using a technique called PSOLA (Pitch Overlap and Add, (Moulines & Charpentier, 1990). Time-compressed speech, which was re-synthesised, versus the natural fast speech, which was not re-synthesised, by also re-synthesising the natural fast speech, a confound was avoided between speech type and synthesis status. Thus, our study differs in that it did not compare two synthetic stimulus types, and that we found that the synthesised speech type (cloned speech) did not benefit from previous exposure to natural speech. This difference suggests that aspects of adaptation to natural speech somehow make it more difficult to recognise cloned speech (in Order 1) than when no speech preceded the synthesised speech (Order 2), resulting in a small negative transfer effect (2.2%). Future research could unpack this finding further, for instance by evaluating transfer of learning effect between synthetic speech types that vary in difficulty systematically, such as four- and six-band noise-vocoded (cloned) speech.
The finding that listeners perceptually adapted to cloned speech at a similar rate and magnitude (i.e., human: 16.1% and cloned: 17% over 40 trials, see Table S3 for test statistics) as to human speech suggests that the auditory perceptual system treats synthetic and human speech in a similar way during the adaptation process. Despite the severe degradation, listeners showed systematic improvements over exposure that mirrored those observed for human voice, to support short-term experience-dependent recalibration. Adaptation in our experiment was largely of a similar magnitude to that reported in our previous work, which used vocoding parameter settings identical to ours (Wang et al., 2023). In Wang et al., participants listened to 40 six-band noise-vocoded sentences in a dual-task design, and they report an increase of 14% for their single-task condition (comparable with our design).
Finally, the intelligibility advantage for the cloned speech was not accompanied by a reliable ability to identify cloned speech with the 2AFC task. Participants identified the human voice at below-chance levels, indicating a slight bias towards classifying voices as cloned. This dissociation suggests that the perceptual mechanisms supporting intelligibility operate independently of explicit voice categorisation, and that listeners may exploit acoustic regularities in cloned speech without conscious awareness of its artificial origin.
Limitations and Future Directions
The design of the current study has several limitations that could be addressed in future studies. First, the study was designed to evaluate how listeners adapted to two types of speech. However, as is the case for any perceptual speech learning experiment, there is no return to baseline once learning has been initiated. Past studies have shown that perceptual learning to certain degradations (such as time-compressed speech) can be retained for up to 10 months after training (Murai & Riquimaroux, 2023). It was therefore not possible to conduct a within-subject design in which we exposed listeners to both types of speech consecutively. Instead, we had to use a more complex between-subject design to allow us to directly compare (by comparing the first 40 trials of each Order) adaptation to cloned and human noise-vocoded speech. Future studies could exploit perceptual adaptation to both speech types more directly by evaluating how listeners adapt to different types of degraded (e.g., noise-vocoding and time-compressed speech) cloned and human speech. For instance, listeners could first be exposed to a block of time-compressed speech (human), followed by a block of noise-vocoded (cloned) speech, and counterbalance the order of degradations and speech types across listeners. Such a design would give further insights into how transfer of learning effects between different degradations interacts with cloned and human speech. Moreover, using time-compressed speech would disrupt the fine phonetic-phonological structure of both speech types, which would allow for investigation of the role of these factors in the increased intelligibility of cloned speech. Finally, future studies could explore an extension to our design. They could do so by also including human-human and clone-clone conditions, where transfer of learning effects could be examined, e.g., between six-band and four-band humans blocks of sentences or six-band and four-band cloned blocks. Such a design would also allow for a more detailed examination of transfer of learning effects and allow researcher to establish whether the negative transfer human-clone reflected an actual “cost” to transfer of learning or can be attributed to a fatigue or saturation effects.
Second, our noise-vocoder was set to use a low-pass envelope cutoff of 300Hz, which preserves slower temporal envelope fluctuations as well faster amplitude modulations. This choice was made to avoid overly severe degradation and ensure that intelligibility remained above floor level with six frequency bands. Using a lower envelope cutoff (e.g. 30-50Hz), as commonly employed in classical cochlear implant simulations (Rosen et al., 2015), would provide a stronger manipulation by restricting the signal primarily to slower temporal envelope cues. Such a manipulation may reduce or eliminate the observed advantage for cloned voices and could therefore help further constrain the acoustic cues underlying this effect. Future studies aimed at isolating the locus of the intelligibility benefit under more strongly degraded listening conditions could systematically manipulate different cutoff frequencies as well as other specific vocoder settings such as the high-frequency cutoff (e.g., 8,000Hz instead of 5,000Hz as used in this study) to establish whether the intelligibility benefit persists for cloned voices and to get a clearer idea of the spectro-temporal locus of the intelligibility benefit for cloned voices. Nevertheless, even in noise-vocoding studies (cf. Shannon et al., 1995) in English listeners (i.e., a non-tonal language!) higher envelope cutoffs do not seem to improve performance much. However, the situation might be different in tasks that are more sensitive to F0 variations. While these variations matter in English, the effect is small. It might be the case, that higher cutoff in tonal languages might improve performance to a greater extent. Moreover, a natural next research direction would be to evaluate the effect of overall degradation severity on the intelligibility benefit for the voice clones, yet here care should be taken to avoid floor and ceiling effects.
Future studies could evaluate the relationship between the cloning intelligibility benefit and adaptation patterns for different voices, as the current study’s design and sample size did not allow us to conduct an exhaustive voice-specific analysis. The cloning process improved overall intelligibility of all voices: for all voice pairs the cloned voice showed higher accuracy scores, between 3.1-24.4% for the current study, cf., Supplemental Table S2), and between 4.9-22.3% for Adank & Wang. It could be the case that an individual voice’s intelligibility benefit post-cloning interacts with the extent to which listeners can adapt to this voice. We could not explore this aspect in our study, as each voice contributed only four sentences to each block, and we did not control when in the block these sentences occurred. Future studies could address this issue by comparing adaptation rates across cloned voices displaying differential intelligibility benefits.
Additionally, we do not know which aspects of the synthesis process allows for the higher intelligibility of the cloned voices. Adank & Wang conducted an extensive acoustic analysis and report that cloned voices differed at various levels, including low-level acoustics, objective intelligibility measured, voice source acoustics, and segmental acoustics. However, they could not directly relate these differences between the cloned and human voices to the intelligibility benefit. Future studies could systematically evaluate how specific differences between cloned and human voices affect intelligibility by manipulating specific aspects (e.g., voice jitter and shimmer, global pitch variability, speaking rate) and measuring any differences in intelligibility before and after manipulation.
A final limitation is our use of a commercial voice cloning system, ElevenLabs (https://ElevenLabs.io), which was used to generate the 10 voice clones we used from Adank & Wang (2026). We chose to use this system because it was also used in Adank & Wang to generate the initial 10 voices we used in the current experiment. We aimed to replicate and extend their findings and were thus limited in our choice of voice cloning software. Note that this decision to use voices originally created using ElevenLabs as the basis for our noise-vocoded stimuli limits the replicability of the research. ElevenLabs does not provide information on its voice-cloning algorithm, model architecture, training procedure, or speaker embedding methods. Moreover, ElevenLabs does not publish sufficient detail on model updates, and it is thus unclear how any update affects the acoustic characteristics of the generated voice clones. Uncertainties about the specifics of commercial models means that we can only treat them as point-in-time snapshots of the technology and thus cannot ensure that voice clones generated using this system are replicable from month to month. We have therefore made the stimuli available via OSF https://osf.io/9z63d/overview, the 80 noise-vocoded sentences per voice plus the original 80 clean sentences originally generated for Adank & Wang. The original (human) recordings for all 10 voices can be found on the ARU database website, https://datacat.liverpool.ac.uk/681/. We recommend that future studies use open source voice cloning software, such as F5TTS (Chen et al., 2025), to ensure replicability and transparency of the cloning process.
Conclusion
Our study showed that the intelligibility benefit for cloned voices persists for noise-vocoded speech, with cloned voices showing 13% higher accuracy compared to their human counterparts. This result demonstrated that the intelligibility benefit originally reported for speech in background noise was not solely due to enhanced harmonic and fine spectral detail in the cloned voices. Listeners also adapted to cloned voices at the same rate as the human voices (16-17%).
Our findings may have implications for future assistive applications, including hearing aids and cochlear implants. More broadly, our results suggest that voice-cloning technologies can be used to enhance intelligibility directly, but also to shape perceptual adaptation in ways benefiting everyday communication. As voice cloning tools mature, they have the potential to form part of a new generation of personalised, evidence-based hearing-support systems. In sum, our results highlight that cloned voices are not only viable experimental tools but can also drive measurable perceptual benefits. As methods for generating these voices continue to advance, they may become a central component of both research, diagnostic, and clinical intervention in hearing and speech science.
Supplemental Material
Supplemental Material - Perceptual Adaptation to Noise-Vocoded Voice Clones
Supplemental Material for Perceptual Adaptation to Noise-Vocoded Voice Clones by Han Wang, Carolyn McGettigan and Patti Adank in Trends in Hearing.
Footnotes
Acknowledgements
This study was supported by UCL. We would like to thank the participants for their time and effort. We thank Stuart Rosen for his expertise regarding vocoder parameter setting.
Ethical Considerations
All participants provided informed consent, and ethics approval was provided by UCL Research Ethics Committee (UREC) to Patti Adank, #0599.005.
Author Contributions
Han Wang: Conceptualization, Methodology, Formal analysis, Data Curation, Writing - Review & Editing, Visualization; Carolyn McGettigan: Methodology, Writing - Review & Editing; Patti Adank: Supervision, Conceptualization, Methodology, Investigation, Validation, Data Curation, Writing - Original Draft, Writing- Reviewing and Editing, Project administration, Funding acquisition.
Funding
The authors disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This study was supported by UCL.
Declaration of Conflicting Interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Data Availability Statement
Analysis R scripts and data can be found on GitHub, at https://github.com/hwanguc/glm_adaptation_cloned_voice. Clean and vocoded versions of all stimuli from the 10 human and 10 cloned voices can be downloaded from
.
Declaration of Generative AI and AI-assisted Technologies in The Writing Process
No AI tools or Generative AI or AI-assisted technologies were used in the writing process.
Supplemental Material
Supplemental material for this article is available online.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
