Neuro-Cognitive Dynamics of
AI Attribution in Music Perception
Can you trust your ears? 🎧
Scan to experience the experiment, view full results, and explore the analysis logic.
Imagine someone plays you a beautiful song. They tell you it was written by AI.
Can you trust your ears? 🎧
Scan to experience the experiment, view full results, and explore the analysis logic.
Abstract
Generative-AI music systems now match human composers on surface acoustic features, yet listeners systematically devalue music attributed to AI. The neural mechanisms of this attribution bias, and how they interact with culturally rooted musical genres, remain underspecified. The present study combines functional near-infrared spectroscopy (fNIRS) with subjective rating measures to isolate the contribution of source labeling from actual acoustic content. A within-subjects 2 × 2 × 3 factorial design (N = 30 behavioral, n = 13 fNIRS retained) crossed Label (AI/Human) × Producer (AI/Human) × Genre (Arabesque/Blues/Electronic). Participants rated 12 musical excerpts on Emotional Investment, Authenticity, Trust, Liking, and Quality (7-point Likert) while prefrontal cortex hemodynamics were recorded. Results revealed a selective AI-attribution penalty: Human-labeled music received higher Authenticity (η²p = .285), Emotional Investment (η²p = .276), and Quality (η²p = .227) ratings, while Liking remained unaffected (p = .270). The label-by-genre interaction was the largest effect for Trust (η²p = .339), with Arabesque showing the steepest AI penalty (Δ = −1.14) and Electronic exhibiting a reversal (Δ = +1.32; AI label increased trust). fNIRS results showed a tentative right-lateralized pattern: AI-labeled trials produced a higher Right PFC HbO2 peak than Human-labeled trials at the participant level as a trend, with a secondary trial-level LME sensitivity model showing a similar direction. Because only 13 participants were retained after quality control, neural findings are interpreted as exploratory. Findings reframe algorithmic aversion as a culturally moderated phenomenon with detectable neural signatures, supporting genre-AI fit expectation accounts.
Keywords: fNIRS, algorithmic aversion, source attribution, prefrontal cortex, cultural essentialism, music perception, placebo effect
Method
Participants: Participants were 30 Turkish university students aged 18–25 years (13 female, 17 male), recruited through convenience sampling. All reported normal hearing and no neurological diagnosis. Twenty-two participants completed fNIRS recording. After quality control, 13 participants were retained for exploratory fNIRS analyses.
Stimuli & Design: 12 excerpts (25s) across Arabesque, Blues, Electronic. 2 Human-composed, 2 AI-generated (Suno v5) per genre. Within-subjects 2×2×3 factorial (Label × Producer × Genre).
Procedure: Each trial began with an unlabeled music excerpt. After 10 seconds, a source label appeared: AI-Generated or Human-Composed. The excerpt continued until 25 seconds. Participants then rated the excerpt on five 7-point Likert scales: Liking, Quality, Trust, Authenticity, and Emotional Investment. A 10-second fixation interval followed each trial.
Measures: Behavioral DVs were Liking, Quality, Trust, Authenticity, and Emotional Investment. fNIRS HbO/HbR signals were inspected; inferential analyses focused on HbO2 Peak as the primary neural metric.
Preprocessing Pipeline: fNIRS preprocessing used raw .oxy HbO/HbR files, participant/file quality control, a 4th-order zero-phase Butterworth band-pass filter (0.01–0.2 Hz), local baseline correction using −10 to 0 s before label reveal, and a 4–15 s post-label HRF response window. ROI-level HbO2 peak values were extracted for Right and Left PFC.
fNIRS preprocessing followed a transparent, literature-informed workflow. Raw .oxy files were inspected for oxygenated and deoxygenated hemoglobin signals (HbO/HbO2 and HbR/Deoxy). Participant and file quality control was applied before inferential analysis. Retained signals were filtered with a 4th-order zero-phase Butterworth band-pass filter from 0.01 to 0.2 Hz to attenuate slow drift and high-frequency physiological/instrumental noise while preserving task-related hemodynamic fluctuations. A local trial-level baseline was computed from the −10 to 0 s interval before label reveal and subtracted from the post-label response. HbO2 Peak was then extracted from a 4–15 s post-label HRF response window for Right and Left PFC ROIs. This window was selected as a literature-informed sensitivity window to capture the delayed hemodynamic response after label onset. Because the retained fNIRS sample was small (n = 13), all neural analyses are reported as exploratory.
Why HbO2 as the primary metric?
Both HbO/HbO2 and HbR/Deoxy were inspected. Inferential analyses focused on HbO2 Peak because HbO is commonly used in task-based fNIRS and typically shows stronger task-related response amplitude, while HbR was treated as a complementary signal.
Why 0.01–0.2 Hz filtering?
This band-pass range is commonly used to reduce very slow drift and higher-frequency physiological/instrumental noise while retaining task-related hemodynamic fluctuations.
Why local baseline −10–0 s?
A local pre-label baseline controls for trial-to-trial drift and captures the participant’s hemodynamic state immediately before the source label appears.
Why 4–15 s HRF window?
fNIRS hemodynamic responses are delayed relative to stimulus onset. The 4–15 s window was used as a literature-informed sensitivity window to capture the expected rise and peak response after label reveal. It should be reported as exploratory/sensitivity analysis, not as a universal standard.
What We Wondered
1. Does the label trick us?
We wondered: if we tell people a song is 'AI-made' or 'Human-made' — but it's actually the same song — would they rate it differently? Spoiler: yes.
2. Does it depend on the music?
Turkish Arabesque feels emotional. Electronic music feels modern. Maybe people forgive AI for making Electronic but not Arabesque? We were right.
3. Can we see it in the brain?
We used a special light-based brain scanner to watch what happens when people see the word 'AI'. The brain works harder, like it's saying 'wait, is this real?'
Neural Results
Right PFC Peak HbO2 (µM): AI [0.247], Human [0.193]
Exploratory Findings
- Yücel et al. (2021) for fNIRS reporting best practices
- Pinti et al. (2019) for transparent preprocessing/reporting of fNIRS pipelines
- Hocke et al. (2018) for common fNIRS filtering practices
- Brigadoi et al. (2014) for motion artifacts and preprocessing concerns
- Fishburn et al. (2014) for HbO sensitivity in cognitive/task-based fNIRS
- Holper et al. (2022) for PFC/fNIRS relevance to trust/source evaluation contexts
- Zhang et al. (2025) for AI-label evaluative bias/neural context if already used in the thesis/poster
Exploratory Behavioral Findings
Beyond the main label effects, the behavioral data suggested several genre- and dimension-specific patterns.
1. Trust was the most label-sensitive judgment
Among the rating dimensions, Trust appeared especially sensitive to source labeling. AI labels reduced trust more strongly than basic liking, suggesting that AI attribution primarily affects perceived agency and reliability rather than simple enjoyment. Trust showed a strong label-related effect, η²p = .339, p < .001.
2. The AI penalty was selective, not global
Participants did not simply dislike AI-labeled music across all dimensions. The AI label mainly affected higher-order judgments such as Trust, Authenticity, Quality, and Emotional Investment, while Liking was comparatively more preserved.
3. Genre changed the meaning of the AI label
The effect of the AI label depended on genre. Arabesque showed a stronger AI-label penalty, Blues appeared more label-invariant or weaker, and Electronic music showed a weaker or partially reversed AI penalty, consistent with genre-label congruency. Trust (Human − AI): Arabesque +1.14, Blues +0.10, Electronic −1.32. This suggests that AI attribution is not a universal bias; it depends on whether the label fits the cultural and acoustic expectations of the genre.
4. Arabesque showed cultural sensitivity to AI labeling
Arabesque appeared especially sensitive to AI labeling. Because it is culturally and emotionally loaded, the AI label may have conflicted with expectations of human expression, lived experience, and cultural authenticity.
5. Electronic music showed possible label congruency
Electronic music showed a weaker or partially reversed AI penalty. This may reflect genre-label congruency: listeners may perceive algorithmic or technological production as more compatible with electronic music than with culturally rooted acoustic genres.
6. Actual producer and displayed label came apart
The design separated what the track actually was from what participants were told it was. This allowed the study to test whether evaluation followed the sound itself or the source attribution attached to it. The displayed source label was manipulated independently of actual producer.
7. Liking and authenticity dissociated
One of the most important behavioral patterns was the dissociation between liking and authenticity. Participants could still like a piece of music while rating it as less authentic, less trustworthy, or less emotionally meaningful when it carried an AI label. This suggests that AI-label bias targets perceived human intention more than immediate pleasure.
References
- Agbangla, N. F. et al. (2022). HRF window standards in fNIRS.
- Bhatt, M. et al. (2020). Scalp Coupling Index Methodology.
- Dietvorst, B. J. et al. (2015). Overcoming algorithmic aversion.
- Pollonini, L. et al. (2014). SCI evaluation in continuous-wave fNIRS.
- Shank, D. B. et al. (2023). Source attribution in AI art.
What We Found
🎵 People don't dislike AI music
Listeners enjoyed AI-labeled music just as much. But they said it felt less 'authentic' and 'soulful' — even when it was the exact same song. The label changes what we believe, not what we feel.
🇹🇷 Arabesque protects itself
When we slapped an 'AI' label on Arabesque, people punished it hardest. It feels like a betrayal of something cultural and human.
🤖 Electronic music does the opposite
In Electronic music, the 'AI' label actually INCREASED trust. People expected AI to make electronic music, so it made sense.
🧠 Preliminary fNIRS pattern
The moment you see the 'AI' word, the right side of your prefrontal cortex showed a higher HbO2 response pattern — like you're switching into critical-thinking mode.
🎭 The label is powerful
Across all genres, the source label was the biggest factor in how people rated the music. The actual quality of the music almost didn't matter.
What Else Did We Notice?
The biggest surprise was that people did not simply dislike AI music. Instead, the AI label changed specific judgments. People were more likely to question whether the music felt trustworthy, authentic, high-quality, or emotionally meaningful. Liking was more protected.
Trust changed most
AI labels mainly affected whether people trusted the music as meaningful or reliable.
Genre mattered
Arabesque was more sensitive to AI labeling, while Electronic music sometimes made the AI label feel more fitting.
Liking is not the whole story
People can enjoy a song while still judging it as less authentic or less emotionally human.
These are behavioral patterns from the study sample. They do not mean every listener reacts the same way.
How We Did It
Discussion
Our data suggests that human perception of art is heavily mediated by mental models held about the creator. The AI label acts as a 'distrust cue', increasing evaluative effort in the right prefrontal cortex. The Trust reversal in Electronic music specifically highlights that algorithmic aversion is not uniform, but heavily dependent on genre-AI fit expectations.
Limitations
- This study was NOT pre-registered.
- fNIRS sample size (n=13) limits statistical power. Neural findings are exploratory.
- One-tailed tests used for directional neural hypotheses.
- Belief in stated source was inferred rather than assessed per trial.
- Sample restricted to Turkish university students.
- 2 Hz fNIRS sampling constrained cardiac-band SCI computation.
fNIRS Analysis Logic & Literature Basis
This section explains the logic behind the fNIRS preprocessing and analysis choices. It is included for transparency and methodological clarity; no analysis scripts are provided on this website. The fNIRS results remain exploratory.
A. What was analyzed?
Raw fNIRS exports were inspected for oxygenated and deoxygenated hemoglobin signals (HbO/HbO2 and HbR/Deoxy). The primary neural summary focused on HbO2 Peak in Right and Left PFC ROIs because HbO is commonly used in task-based fNIRS and often shows stronger task-related response amplitude.
B. How was the signal cleaned?
Signals were filtered using a 4th-order zero-phase Butterworth band-pass filter from 0.01 to 0.2 Hz to reduce slow drift and higher-frequency physiological/instrumental noise while preserving task-related hemodynamic fluctuations.
C. How was baseline handled?
A local trial-level baseline was defined as the −10 to 0 s interval before label reveal. This local pre-label baseline captures the participant’s hemodynamic state immediately before the source attribution appears and helps control trial-to-trial drift.
D. How was the response quantified?
After label reveal, HbO2 Peak was extracted from a literature-informed post-label HRF response window for each trial and ROI. The analysis focused on Right PFC and Left PFC to test whether the AI-label pattern was right-lateralized.
E. How were conditions compared?
For each participant, AI-labeled and Human-labeled trials were averaged separately. AI − Human difference scores were tested at the participant level. The printed poster additionally reports a secondary trial-level LME sensitivity model with participant as a random intercept.
F. Why is this exploratory?
The retained fNIRS sample was small (n = 13) and participant/channel exclusion was substantial. Therefore, fNIRS findings are interpreted as preliminary neural evidence and require replication. Behavioral findings from the full N = 30 sample are the stronger evidence in this thesis.
Why This Matters
Hypotheses & Results
Prediction: AI-labeled music will evoke greater HbO activation in prefrontal cortex than Human-labeled music.
Result: Right PFC Peak HbO2: AI = 0.247 µM, Human = 0.193 µM, Δ = +0.055 µM. Paired t(12) = 1.60, p₁ = .068 (one-tailed). LME: β = 0.058, SE = 0.027, p₁ = .047.
Prediction: Human-labeled music will receive higher Authenticity, Quality, and Emotional Investment ratings.
Results: EI: F(1,29) = 11.07, p = .002, η²p = .276. Auth: F(1,29) = 11.55, p = .002, η²p = .285. Quality: F(1,29) = 8.49, p = .007, η²p = .227.
Prediction: AI-generated music presented under Human label will receive higher ratings.
Results: Emotional Investment: F(1,29) = 4.23, p = .049, η²p = .127.
Prediction: Electronic music will receive lower Auth and EI ratings than Arabesque or Blues.
Results: Auth F(2,58) = 11.81, p < .001. EI F(2,58) = 41.18, p < .001 (Largest single effect).
Prediction: Arabesque will score highest under Human label but show largest AI penalty.
Results: Label × Genre on EI: F(2,58) = 3.42, p = .039. AI penalty: ΔEI = −0.88.
Prediction: Trust gap (Human − AI) will be larger in acoustic genres than Electronic.
Results: Label × Genre on Trust: F(2,58) = 14.86, p < .001. Arabesque +1.14. Electronic −1.32 (REVERSAL: AI label increased trust).