Department of Psychology, Hebei Normal University, Shijiazhuang, China
1 Introduction
In Chinese, the word “蓝牙” (/lan2ya2/, meaning Bluetooth) denotes a wireless communication technology and consists
of two free morphemes, “蓝” (/lan2/, meaning blue) and “牙” (/ya2/, meaning tooth), both of which can function as independent
words and carry distinct semantic representations and concreteness values. When listeners encounter such spoken input, a fundamental question concerns how the cognitive system accesses constituent morphemes while retrieving the holistic meaning of compound words in the study of morphologically complex word processing. As a primary form of complex words, compound words are formed by the combination of two or more free morphemes, a phenomenon widely observed across numerous languages (Juhasz et al., 2003). Linguistic typological statistics indicate that compound words account for the vast majority of the modern Chinese lexicon (Academia Sinica, 1998); therefore, investigating the cognitive mechanisms underlying compound word processing is of considerable theoretical significance.
The role of constituent morphemes in compound word processing has been the focus of a long-standing theoretical debate among three major classes of models. Sublexical models posit that lexical processing obligatorily depends on the early decomposition of morphemes; thus, the whole-word representation can be activated only after constituent morphemes are accessed. For instance, the File Drawer Model proposed by Taft and Forster (1975) and the Hierarchical Model formulated by Taft (2003, 2004) both suggest that morphemes are activated first, thereby guiding subsequent access to the whole-word lemma. Conversely, supralexical models argue that morphologically complex words are stored independently and accessed directly via a whole-word route (Giraudo & Grainger, 2000, 2001; Grainger & Jacobs, 1996). Dual-route models hypothesize that the decomposition route and the direct whole-word route operate in parallel during lexical processing (Caramazza et al., 1988; Pollatsek et al., 2000, 2011; Schreuder & Baayen, 1995).
This debate predominantly centers on visual word recognition. In research on alphabetic scripts, investigators have focused on how morphemic features (e.g., frequency and length) and whole-word properties modulate the time course of lexical processing. Although early studies argued for serial identification (Hyönä & Pollatsek, 1998), subsequent eye-movement studies have provided substantial evidence in favor of parallel processing (Bertram & Hyönä, 2003; Pollatsek et al., 2000). Using a large-scale lexical decision task, Kuperman et al. (2009) demonstrated that whole-word frequency and first-morpheme frequency exert simultaneous effects during the early stages of processing and consequently proposed a multi-route parallel processing model. In Chinese compound word processing, converging evidence has likewise demonstrated the parallel activation of morphemic and whole-word representations (Peng et al., 1994; Taft et al., 1994; Zhang & Peng, 1992). For example, Taft et al. (1994) manipulated constituent morpheme frequencies while controlling for whole-word frequency and found that the frequencies of both the first and second morpheme significantly influenced lexical decision latencies. According to the inter-intra connection model (Peng et al., 1999), orthographic input maps onto lexical networks and simultaneously activates whole-word and morpheme units, suggesting that access to whole-word and morphemic representations unfolds without strict temporal precedence and that the two levels interact continuously throughout processing.
The dynamics of compound word processing appear to change when these model debates are extended to spoken language comprehension. Because auditory input unfolds incrementally and continuously over time, complete acoustic information about a compound word becomes fully available only after the offset of the second constituent. Consequently, constituent-level information may become available before complete whole-word information has accumulated. Researchers have therefore focused on the temporal question of whether the initial constituent is activated before the whole word during real-time spoken word recognition. On the one hand, several studies support early morphological decomposition. For instance, an event-related potential (ERP) study by Holle et al. (2010) demonstrated that a non-word constituent condition elicited a significant N400 effect at the acoustic offset of the first morpheme in spoken compound words. Shen et al. (2017) using the printed-word version of the Visual World Paradigm (PW-VWP), found that visual competitors semantically related to the first constituent of the spoken target elicited significantly more fixations immediately following the acoustic offset of the first syllable. On the other hand, alternative studies have failed to observe such early morphemic activation and instead argue for the temporal precedence of whole-word access. For example, Isel et al. (2003) found that semantic priming effects related to the initial morpheme emerged only after the completion of the spoken compound word. Similarly, in spoken Chinese comprehension, Zhou and Marslen-Wilson (1994) used an auditory lexical decision task and observed only a robust whole-word frequency effect on response latencies to real words, failing to detect independent constituent-frequency or syllabic-frequency effects.
This contradiction is particularly prominent in spoken Chinese, given its unique morpheme-syllable mapping and the prevalence of homophones. When listeners process an initial syllable, the absence of visual orthographic cues for disambiguation allows the acoustic signal to activate a large cohort of candidate representations (Zhou & Marslen-Wilson, 1995). The activation of these homophonic morphemes creates competition, which may weaken the morpheme frequency effects. This may explain why traditional behavioral studies (e.g., Zhou & Marslen-Wilson, 1994) failed to observe robust morphemic facilitation effects. Although Zou et al. (2019) demonstrated that shared morphemic semantics and shared whole-word semantics elicit distinct N400 patterns during later processing stages—supporting separable representational pathways—the temporal dynamics during the initial access stage remain unclear.
During the processing of disyllabic compound words, following the initial activation of the first morpheme, how the time courses of the second morpheme and whole-word semantics unfold in the subsequent acoustic stream remains an open question. As acoustic information from the second syllable unfolds incrementally, the second constituent is presumably decoded prior to complete whole-word access. To date, few studies have examined the temporal dynamics of semantic activation associated with the second constituent. Zhou and Marslen-Wilson (1995) using an auditory priming paradigm, reported a clear position asymmetry: the initial homophonic morpheme produced inhibition, whereas the final morpheme produced facilitation. This contrast suggests that the cognitive system dynamically adjusts its processing strategy as the acoustic stream unfolds. According to Libben’s (2006) maximal opportunity account, constituent morpheme activation and whole-word activation are not strictly sequential stages but instead co-occur within an interactive network.
In the present study, we employed the PW-VWP to examine the complete time course of semantic access during auditory Chinese compound word processing. Upon hearing the spoken word, four semantic conditions were included in the visual display: first-morpheme related words, second-morpheme related words, whole-word related words, and unrelated words. This design aimed to determine whether semantic access unfolds in a strictly serial manner or follows a cascaded time course characterized by early first-morpheme guidance followed by late parallel activation of whole-word and second-morpheme representations. Based on the properties of Chinese words and previous empirical findings (Shen et al., 2017; Zou et al., 2019), the following hypotheses were formulated: a temporal precedence effect of the first morpheme is expected to emerge during the early stage of auditory word input, marked by an early deviation from baseline in fixation probabilities for the first-morpheme-related condition. Conversely, during the late stage, as the second syllable unfolds, whole-word and second-constituent semantics are expected to be activated in parallel to meet increased demands for disambiguation.
Thirty-six undergraduate students (8 males, 28 females) were recruited from a university in Hebei Province. Their ages ranged from 22 to 26 years (M = 22.38, SD = 1.13). All participants were Chinese native speakers, were right-handed, and had normal or corrected-to-normal vision. None of the participants had taken part in similar experiments, and all provided informed consent before the experiment.
During the experiment, while listening to a spoken compound word (e.g., 蓝牙, /lan2ya2/), participants viewed a visual scene consisting of four words displayed at the four corners of the screen (see Figure 1). Based on the semantic relationships between the constituent morphemes or the whole-word meaning of the spoken word and the visual words, four experimental conditions were created: (1) the first-morpheme related condition, where a visual word (e.g., 天空, /tian1kong1/, meaning sky) was semantically related to the first morpheme of the spoken word (e.g., 蓝, meaning blue); (2) the second-morpheme related condition, where a visual word (e.g., 嘴巴, /zui3ba0/, meaning mouth) was semantically related to the second constituent of the spoken word (e.g., 牙, meaning tooth); (3) the whole-word related condition,
where a visual word (e.g., 手机, /shou3ji1/, meaning mobile phone) was semantically related to the meaning of the spoken word;
and (4) the unrelated condition, where a visual word with no relationship to the spoken word. The spatial positions of the four visual words were completely randomized.
Figure 1 An Illustrative Example of the Visual Display
During the material selection process, an initial pool of 60 auditory target words and their corresponding related words was compiled by the experimenter. We ensured that the auditory targets shared no phonological or orthographical features with visual words. Because spoken words differ fundamentally from written words in their presentation format, a preliminary norming study was conducted on the first constituents of the auditory targets to ensure that the majority of listeners would activate the same identical Chinese character upon hearing that syllable. Fifteen participants completed a dictation test targeting the first morphemes of auditory stimuli. We retained items with dictation accuracy above 70% and excluded 25 morphemes that failed to meet this criterion.
Subsequently, a norming task was conducted to evaluate the degree of semantic relatedness between the spoken words and their corresponding first-morpheme, second-morpheme, whole-word related visual words, and unrelated distractors. Based on the 35 remaining word sets (each set comprising one spoken word and four visual words), a semantic relatedness questionnaire was constructed using a 7-point Likert scale (1 = completely unrelated, 7 = highly related). The questionnaire was distributed to 25 students who did not participate in the subsequent formal eye-tracking experiment. The results of the semantic relatedness scores, alongside other psycholinguistic attributes, are summarized in Table 1. A one-way analysis of variance (ANOVA) revealed a highly significant main effect of semantic condition, F (3,136) = 2769, p < 0.001. Post-hoc comparisons indicated that the semantic relatedness scores for the first-morpheme, second-morpheme, and whole-word related conditions were each significantly higher than that of the unrelated condition, ps < 0.001. Crucially, no significant differences in relatedness scores were observed among the three related conditions in pairwise comparisons, ps > 0.05. The stroke numbers and morpheme frequencies of the constituent components were also matched, showing no significant differences, ps > 0.05. Additionally, the word frequencies of the four visual words within the visual scene were statistically equivalent, F (3,136) = 2.19, p > 0.05.
In addition to the critical trials, four practice trials and 35 filler trials were added. In the filler trials, the spoken word
exactly matched one visual word on the display, with the other three serving as unrelated distractors (e.g., upon hearing 铃铛, /ling2dang0/, meaning bell, participants could find the identical visual word 铃铛 on the screen). Data from the filler
and practice trials were excluded from subsequent statistical analyses. Consequently, the formal experiment comprised a total of 74 trials.
Table 1 Semantic Relatedness, Stroke Count, and Frequencies for the Related Words
|
First-morpheme related word |
Second-morpheme related word |
Whole-word related word |
Unrelated word |
||
|
Semantic relatedness |
6.44(0.24) |
6.28(0.29) |
6.37(0.37) |
1.38(0.17) |
|
|
Word frequency |
9991.34(8783.16) |
11123.87(10610.40) |
7878.99 (8822.57) |
14164.11(13146.86) |
|
|
Morpheme frequency* |
825.60(1075.16) |
840.83(1079.31) |
- |
- |
|
|
Stroke* |
7.37(3.04) |
9.03(3.78) |
- |
- |
|
Note: * indicates the first or second morpheme of the spoken words; other items represent the semantically related visual words. Standard deviations are presented in parentheses.
Eye movements were recorded using an SR Research EyeLink 1000 Plus eye-tracker with a sampling rate of 1000 Hz. The visual scenes were displayed on a 21-inch LCD monitor (SONY Multiscan G520). During the experiment, participants viewed the screen binocularly, while the eye-tracker recorded monocular data from the participant’s right eye. The viewing distance from the participant’s eyes to the monitor was approximately 58 cm.
A standard 9-point calibration and validation procedure was performed prior to the formal experiment. We kept the maximum calibration error below 1° and the average validation error below 0.5°. A drift correction was executed at the beginning of each trial. Each trial started with a 600 ms blank screen, after which a display of four printed compound words appeared. The visual display was presented for a 200 ms preview period prior to the onset of the spoken target word (see Figure 2 for the trial sequence). Participants were instructed to judge whether the spoken target word was visually present on the screen by pressing the “F” key for “yes” and the “J” key for “no” (response keys were counterbalanced across participants). Before entering the formal experiment, participants completed practice trials to familiarize themselves with the procedure. The entire experimental session took approximately 20 minutes to complete.
Figure 2 Experimental Trial Sequence
To ensure data quality, participants with an overall behavioral error rate exceeding 5% were removed from further statistical analysis. Consequently, data from 4 participants were discarded, leaving a final sample of 32 participants for subsequent analyses. To capture the precise temporal dynamics of spoken compound word comprehension, we analyzed the eye-tracking data using two complementary approaches: First, following the time-interval approach (Salverda & Tanenhaus, 2010), the 0-2000ms epoch from spoken word onset was divided into consecutive 100-ms time windows. For each window, fixation probabilities were analyzed using Generalized Linear Mixed-Effects Models (GLMMs) (Jaeger, 2008). The unrelated condition served as the baseline. The experimental condition was specified as a fixed factor, while participants and items were entered as random intercepts. To resolve model convergence failures, we simplified the maximal random-effect structure by removing random slopes for participants and items. Convergence diagnostics confirmed that the revised model achieved stable convergence for all time-window analyses. All statistical modeling was conducted in the R environment (R Core Team, 2023) using the glmer function in the lme٤package (Bates et al., ٢٠١٥). p-values for the mixed models were calculated via the lmerTest package (Kuznetsova et al., ٢٠١٧). Second, to fully capture the continuous temporal dynamics of the data, we performed a Growth Curve Analysis (GCA) (Mirman et al., 2008). In the Level 1 submodel, individual fixation probability curves over time were modeled using second-order orthogonal polynomials, which included an intercept (overall mean fixation probability), a linear term (the overall rate of change), and a quadratic term (the curvature, reflecting a symmetric rise and fall). The Level 2 submodel included the fixed effects of the experimental conditions and random intercepts for participants and items. Regarding the GCA parameters, the intercept captures the overall time-averaged fixation probability, the linear term reflects the baseline slope of the fixation curves, and the quadratic term describes the rise-and-fall shape of the fixation curve (Mirman, 2017). While the psychological reality of higher-order polynomial terms (e.g., cubic or quartic) remains less clear, we thus focused exclusively on the linear and quadratic terms in the present study.
Figure 3 illustrates the dynamic changes in fixation probabilities for the first-morpheme related, second-morpheme related, whole-word related, and unrelated control conditions across consecutive 100-ms intervals from 200ms prior to the onset of the auditory target word to 2000ms post-onset. Table 2 summarizes the parameter estimates from the GLMMs within the significant time windows. Given that executing a saccadic eye movement typically requires approximately 150 to 200ms (Rayner, 1998), our results indicated that starting from the 200-ms time window, participants elicited significantly higher fixation probabilities for the first-morpheme related competitors than for the unrelated distractors in the early windows (i.e., 200-400ms and 700-800ms), as well as in a late window (1400-1600ms). However, no significant differences in fixation probabilities were detected between either the second-morpheme related or whole-word related conditions and the unrelated baseline from 0 to 1200ms. Starting from the 1200-ms window, fixations on the whole-word related condition became significantly greater than the unrelated condition. The second-morpheme related condition also attracted significantly more fixations than the unrelated baseline from 1300ms and persisted until 1800ms. Notably, within the 1400-1600ms, the fixation probabilities for the first-morpheme, second-morpheme, and whole-word related items were all significantly higher than those for the unrelated distractors.
Figure 3 Fixation probabilities across conditions
Note: The blue, green, and pink shaded areas represent the significant time windows for the first-morpheme, whole-word, and second-morpheme conditions relative to the unrelated baseline, respectively.
Table 2 GLMM Results Across Conditions Relative to the Unrelated Baseline
|
Time windows (ms) |
Intercept |
first-morpheme related condition |
second-morpheme related condition |
whole-word related condition |
||||||||
|
b |
SE |
z |
b |
SE |
z |
b |
SE |
z |
b |
SE |
z |
|
|
200 - 300 |
-1.141 |
0.06 |
-19.98*** |
0.154 |
0.08 |
1.93† |
-0.003 |
0.08 |
-0.04 |
-0.036 |
0.08 |
-0.45 |
|
300 - 400 |
-1.147 |
0.06 |
-20.06*** |
0.196 |
0.08 |
2.48* |
-0.033 |
0.08 |
-0.41 |
-0.007 |
0.08 |
-0.08 |
|
400 - 500 |
-1.157 |
0.06 |
-20.18*** |
0.130 |
0.08 |
1.63 |
0.039 |
0.08 |
0.48 |
0.055 |
0.08 |
0.68 |
|
700 - 800 |
-1.118 |
0.06 |
-19.70*** |
0.161 |
0.08 |
2.04* |
0.000 |
0.08 |
< 0.001 |
0.054 |
0.08 |
0.68 |
|
1200 - 1300 |
-1.137 |
0.06 |
-19.94*** |
-0.023 |
0.08 |
-0.28 |
0.013 |
0.08 |
0.161 |
0.147 |
0.08 |
1.86† |
|
1300 - 1400 |
-1.272 |
0.06 |
-21.51*** |
0.062 |
0.08 |
0.75 |
0.148 |
0.08 |
1.80† |
0.249 |
0.08 |
3.07** |
|
1400 - 1500 |
-1.473 |
0.07 |
-21.74*** |
0.172 |
0.09 |
1.99* |
0.255 |
0.09 |
2.98** |
0.412 |
0.08 |
4.90*** |
|
1500 - 1600 |
-1.676 |
0.08 |
-20.81*** |
0.173 |
0.09 |
1.88† |
0.251 |
0.09 |
2.77** |
0.456 |
0.09 |
5.15*** |
|
1600 - 1700 |
-1.889 |
0.10 |
-19.51*** |
0.136 |
0.10 |
1.38 |
0.231 |
0.10 |
2.38* |
0.408 |
0.10 |
4.32*** |
|
1700 - 1800 |
-2.226 |
0.13 |
-17.86*** |
0.198 |
0.11 |
1.84† |
0.215 |
0.11 |
2.00* |
0.411 |
0.10 |
3.94*** |
|
1800 - 1900 |
-2.535 |
0.15 |
-16.91*** |
0.095 |
0.12 |
0.79 |
0.130 |
0.12 |
1.09 |
0.404 |
0.11 |
3.54*** |
|
1900 - 2000 |
-2.925 |
0.18 |
-16.44*** |
-0.010 |
0.14 |
-0.07 |
0.155 |
0.13 |
1.16 |
0.285 |
0.13 |
2.19* |
Note: † p < 0.1, *p<0.05, **p<0.01, *** p<0.001.
Furthermore, we also compared the differences in fixation probabilities among the three critical semantic conditions, as summarized in Table 3. Results revealed that the fixation probability for the first-morpheme related condition was significantly higher than that for the whole-word related condition from 0 to 400ms, ps < 0.05. After 1200ms, this pattern reversed, with the fixation probability for the whole-word related condition becoming significantly greater than that for the first-morpheme related condition, ps < 0.05. The divergence between the whole-word related and second-morpheme related conditions emerged at 1400 ms, reaching a marginally significant level, p = 0.053, and subsequently achieved full significance from 1500 ms to 1900 ms, where whole-word items consistently attracted more fixations, ps < 0.05.
Table 3 GLMM Results for First- and Second-Morpheme Related Condition Relative to the Whole-Word Related Baseline
|
Time windows (ms) |
Intercept |
first-morpheme related condition |
second-morpheme related condition |
||||||
|
b |
SE |
z |
b |
SE |
z |
b |
SE |
z |
|
|
0 - 100 |
-1.354 |
0.06 |
-20.89*** |
0.174 |
0.08 |
2.09* |
0.018 |
0.09 |
0.21 |
|
100 - 200 |
-1.224 |
0.06 |
-20.97*** |
0.163 |
0.08 |
2.02* |
0.014 |
0.08 |
0.17 |
|
200 - 300 |
-1.177 |
0.06 |
-20.42*** |
0.190 |
0.08 |
2.38* |
0.033 |
0.08 |
0.41 |
|
300 - 400 |
-1.154 |
0.06 |
-20.14*** |
0.203 |
0.08 |
2.56* |
-0.026 |
0.08 |
-0.33 |
|
1200 - 1300 |
-0.990 |
0.06 |
-17.99*** |
-0.170 |
0.08 |
-2.14* |
-0.134 |
0.08 |
-1.70† |
|
1300 - 1400 |
-1.023 |
0.06 |
-18.46*** |
-0.187 |
0.08 |
-2.32* |
-0.101 |
0.08 |
-1.27 |
|
1400 - 1500 |
-1.061 |
0.06 |
-17.24*** |
-0.240 |
0.08 |
-2.93** |
-0.157 |
0.08 |
-1.94† |
|
1500 - 1600 |
-1.220 |
0.07 |
-16.58*** |
-0.283 |
0.09 |
-3.30** |
-0.205 |
0.08 |
-2.42* |
|
1600 - 1700 |
-1.480 |
0.09 |
-16.37*** |
-0.272 |
0.09 |
-2.95** |
-0.178 |
0.09 |
-1.96† |
|
1700 - 1800 |
-1.815 |
0.12 |
-15.37*** |
-0.213 |
0.10 |
-2.12* |
-0.196 |
0.10 |
-1.96* |
|
1800 - 1900 |
-2.131 |
0.143 |
-14.89*** |
-0.309 |
0.11 |
-2.76** |
-0.274 |
0.11 |
-2.46* |
|
1900 - 2000 |
-2.639 |
0.172 |
-15.31*** |
-0.296 |
0.13 |
-2.26* |
-0.131 |
0.13 |
-1.04 |
Note: † p < 0.1, *p<0.05, **p<0.01, *** p<0.001.
Based on the time-window analyses showing distinct significant effect patterns across time, we conducted Growth Curve Analysis (GCA) on two key intervals: 0-800 ms (only first-morpheme effects were significant) and 1200-2000 ms (all three semantic conditions showed significant effects). The 800-1200 ms interval contained no stable, significant semantic effects and was thus excluded from GCA. GCA for the 0-800ms time windows revealed a highly significant effect for the intercept term in the first-morpheme-related condition relative to the unrelated baseline (see Table 4 for statistical details; see Figure 4 for model fits). This finding indicates that the overall mean fixation probability for the first-morpheme-related condition was significantly higher than that for the unrelated condition across this entire early interval. However, no significant effects were observed for any terms in either the second-morpheme-related or whole-word related conditions.
Table 4 GCA Results Within the 0-800 ms Time Window Relative to the Unrelated Baseline
|
Condition |
Term |
b |
SE |
t |
p |
|
First-morpheme related |
Intercept |
0.027 |
0.004 |
6.2 |
< 0.001 |
|
Linear |
-0.005 |
0.012 |
-0.38 |
0.703 |
|
|
Quadratic |
-0.003 |
0.012 |
-0.26 |
0.792 |
|
|
Second-morpheme related |
Intercept |
0.000 |
0.004 |
-0.01 |
0.994 |
|
Linear |
0.002 |
0.012 |
0.14 |
0.886 |
|
|
Quadratic |
-0.001 |
0.012 |
-0.08 |
0.933 |
|
|
Whole-word related |
Intercept |
0.001 |
0.004 |
0.24 |
0.811 |
|
Linear |
0.012 |
0.012 |
1 |
0.319 |
|
|
Quadratic |
0.003 |
0.012 |
0.26 |
0.795 |
Figure 4 GCA Results of Fixation Proportions Across Conditions From 0 to 800 ms
For the late 1200-2000ms time window, the GCA revealed that the intercept terms for the first-morpheme, second-morpheme, and whole-word related conditions were all significantly higher than that for the unrelated condition (see Table 5 for statistical details; see Figure 5 for model fits). These results demonstrate that during the late stage of processing, the overall mean fixation probabilities for all three semantic conditions were significantly greater than those for the baseline, with the whole-word related condition exhibiting the strongest effect. Regarding the linear term, both the second-morpheme-related and whole-word-related conditions yielded significantly (or marginally significantly) higher slopes than the unrelated condition. This indicates that both the second-morpheme and whole-word related conditions showed a significantly faster rate of fixation accumulation within this window compared to the unrelated distractors. Crucially, on the quadratic term, only the whole-word related condition differed significantly from the unrelated condition, indicating that the fixation curve for the whole-word related condition captured a distinct rise-and-fall curvature relative to the unrelated condition.
Table 5 GCA Results Within the 1200-2000 ms Time Window Relative to the Unrelated Baseline
|
Condition |
Term |
b |
SE |
t |
p |
|
First-morpheme related |
Intercept |
0.026 |
0.009 |
2.93 |
0.003** |
|
Linear |
0.032 |
0.025 |
1.29 |
0.197 |
|
|
Quadratic |
-0.031 |
0.025 |
-1.25 |
0.211 |
|
|
Second-morpheme related |
Intercept |
0.046 |
0.009 |
5.19 |
< 0.001*** |
|
Linear |
0.057 |
0.025 |
2.27 |
0.023* |
|
|
Quadratic |
-0.022 |
0.025 |
-0.88 |
0.381 |
|
|
Whole-word related |
Intercept |
0.075 |
0.009 |
8.44 |
< 0.001*** |
|
Linear |
0.045 |
0.025 |
1.79 |
0.074† |
|
|
Quadratic |
-0.066 |
0.025 |
-2.64 |
0.008** |
Note: † p < 0.1, *p<0.05, **p<0.01, *** p<0.001.
Figure 5 GCA Results of Fixation Proportions Across Conditions From 1200 to 2000 ms
The present study used the printed-word VWP to systematically investigate the time course of semantic activation for the first morpheme, the second morpheme, and the whole word during spoken Chinese disyllabic compound word comprehension. Findings from both the time-window analysis and GCA demonstrated that during the early stage of spoken word input (0-800 ms), the semantic representation of the first morpheme was accessed first. During the mid-to-late stage of spoken word processing (1200-2000 ms), the second morpheme and whole-word semantics exhibited a parallel activation pattern; however, the whole-word-related condition elicited a significantly higher mean fixation probability than the other related conditions, indicating its dominance during the late stage of processing. These findings provide direct online eye-tracking evidence that clarifies both the temporal trajectory of semantic access and the relative dominance of processing pathways over time in spoken Chinese compound word recognition.
Specifically, within the 0-800 ms time window, only the first-morpheme condition showed significant semantic activation, whereas the whole-word and second-morpheme semantic effects had not yet emerged. This pattern is consistent with the findings of Shen et al. (2017), providing empirical support for the view that morphological decomposition occurs early during compound word processing. According to the Cohort Model of spoken word recognition (Marslen-Wilson, 1987), spoken input unfolds sequentially over time. Once the acoustic input of the first syllable ends, morphological decomposition initiates to retrieve the semantics of the first morpheme, thereby driving an early fixation deviation toward the first-morpheme-related condition. This finding is also consistent with the evidence reported by Holle et al. (2010), which suggests that at the earliest stage of compound processing, the cognitive system tends to perform morphological segmentation on temporally available auditory input. In Chinese, the high correspondence between syllables and morphemes provides salient boundary cues, which further facilitate lexical segmentation in continuous speech, enabling bottom-up morphological processing to occur at the early stage of lexical recognition (Tsang & Chen, 2014).
However, the results observed within the 1200-2000 ms time window challenge a strict serial decomposition model (Taft, 2003, 2004). The data revealed that during the late stage, not only did the second-morpheme related condition elicit significant fixation effects, but the whole-word condition also reached a strong level of significance. This indicates that, upon the auditory input of the second syllable, the bottom-up decoding pathway of the second morpheme and the top-down predictive pathway of the whole word unfold concurrently within the same temporal integration window. These results extend Libben’s (2006) maximization of opportunity principle to spoken Chinese, demonstrating that the comprehension system concurrently recruits both constituent-level and whole-word representations to optimize processing efficiency. This pattern diverges from findings in Western alphabetic languages. In the auditory processing of compound words in German, for example, Isel et al. (2003) reported a temporal precedence of whole-word access, with morphological decomposition effects emerging relatively late. In contrast, modern Chinese compound words are typically high-frequency and structurally concise, consisting of only two syllables. This compact structure leads to a high degree of temporal overlap between the whole-word access window and the second-constituent decoding window, thereby favoring an online retrieval strategy jointly driven by both constituent-level and whole-word-level pathways.
Crucially, although the late window was characterized by parallel activation, the pattern showed clear asymmetry and pathway dominance. The whole-word-related condition elicited the strongest effect among the three related conditions, whereas both the second morpheme and whole word showed significant or marginally significant linear growth during the late phase. In contrast, the first morpheme showed a significant intercept but no significant linear trend. This pattern may reflect the influence of homophonic competition in spoken Chinese (Zhou & Marslen-Wilson, 1995). Spoken Chinese contains numerous homophonic morphemes. An initial syllable thus activates many candidate entries, producing strong early activation but weak sustained effects. In the late stage, such competition may attenuate the contribution of first-morpheme activation, while processing increasingly relies on whole-word access and the integration of second-syllable information for disambiguation (Cheng, 2023).
Nevertheless, the current study is subject to several methodological limitations regarding its methodology and theoretical generalizability. First, the activation time course of constituent semantics may be modulated by factors such as semantic transparency and morphological neighborhood density, which should be systematically considered in future research. Second, the PW-VWP employed in the present study used orthographic materials in the visual display, which inevitably introduced orthographic information and may have triggered top-down orthographic feedback. Future studies should consider combining event-related potential techniques or using an object-based version of the visual world paradigm in a purely auditory environment, entirely devoid of orthographic artifacts, to further validate the temporal course and neural mechanisms underlying the parallel activation of second-constituent and whole-word semantics.
The present study investigated the semantic access mechanisms of Chinese disyllabic compound words in the auditory modality, providing a detailed account of the temporal dynamics of spoken word recognition. Crucially, the findings demonstrate that spoken Chinese compound word processing follows a morphological decomposition pathway during the early stage of lexical input. In this initial phase, the semantics of the first morpheme are activated earlier than others, showing clear temporal precedence. As the acoustic signal unfolds into the mid-to-late stage, processing shifts toward a parallel mechanism, characterized by the concurrent activation of bottom-up acoustic decoding of the second constituent and top-down whole-word access pathways. Within this late parallel activation phase, however, the top-down whole-word access pathway exerts dominance; the rapid extraction of holistic representations is pivotal for real-time semantic disambiguation as the auditory signal unfolds.
[1] Academia Sinica. (1998). Academia Sinica balanced corpus (Version 3) [Online database]. Chinese Knowledge and Information Processing Group.
[2] Bates, D., Mächler, M., Bolker, B. M., & Walker, S. C. (2015). Fitting linear mixed-effects models using lme4. Journal of Statistical Software, 67(1), 1-48.
[3] Bertram, R., & Hyönä, J. (2003). The length of a complex word modifies the role of its constituents. Journal of Memory and Language, 48(3), 615-634.
[4] Caramazza, A., Laudanna, A., & Romani, C. (1988). Lexical access and inflectional morphology. Cognition, 28(3), 297-332.
[5] Cheng, H. (2023). A study on representational and morpheme position priority in auditory Chinese compound word processing [Doctoral dissertation]. Zhejiang University.
[6] Giraudo, H., & Grainger, J. (2000). Effects of prime word frequency and cumulative root frequency in masked morphological priming. Language and Cognitive Processes, 15(4-5), 421-444.
[7] Giraudo, H., & Grainger, J. (2001). Priming complex words: Evidence for supralexical representation of morphology. Psychonomic Bulletin & Review, 8(1), 127-132.
[8] Grainger, J., & Jacobs, A. M. (1996). Orthographic processing in visual word recognition: A multiple-readout model. Psychological Review, 103(3), 518-565.
[9] Holle, H., Gunter, T. C., & Koester, D. (2010). The time course of lexical access in morphologically complex words. NeuroReport, 21(5), 319-323.
[10] Hyönä, J., & Pollatsek, A. (1998). Reading Finnish compound words: Eye movements in the identification of standard and transected compounds. Journal of Experimental Psychology: Human Perception and Performance, 24(6), 1612-1627.
[11] Isel, F., Gunter, T. C., & Friederici, A. D. (2003). Prosody-assisted head-driven access to spoken German compounds. Journal of Experimental Psychology: Learning, Memory, and Cognition, 29(2), 277-288.
[12] Jaeger, T. F. (2008). Categorical data analysis: Away from ANOVAs (transformation or not) and towards logit mixed models. Journal of Memory and Language, 59(4), 434-446.
[13] Juhasz, B. J., Starr, M. S., Inhoff, A. W., & Placke, L. (2003). The effects of morphology on the processing of compound words: Evidence from naming, lexical decisions and eye fixations. British Journal of Psychology, 94(2), 223-244.
[14] Kuperman, V., Schreuder, R., Bertram, R., & Baayen, R. H. (2009). Reading polymorphemic Dutch compounds: Toward a multiple route model of lexical processing. Journal of Experimental Psychology: Human Perception and Performance, 35(3), 876-895.
[15] Kuznetsova, A., Brockhoff, P. B., & Christensen, R. H. B. (2017). lmerTest package: Tests in linear mixed effects models. Journal of Statistical Software, 82(13), 1-26.
[16] Libben, G. (2006). Why study compounds? An overview of the issues. In G. Libben & G. Jarema (Eds.), The Representation and Processing of Compound Words (p. 1-21). Oxford University Press.
[17] Marslen-Wilson, W. (1987). Functional parallelism in spoken word-recognition. Cognition, 25(1-2), 71-102.
[18] Mirman, D. (2017). Growth Curve Analysis and Visualization Using R. Chapman and Hall/CRC.
[19] Mirman, D., Dixon, J. A., & Magnuson, J. S. (2008). Statistical and computational models of the visual world paradigm: Growth curves and individual differences. Journal of Memory and Language, 59(4), 475-494.
[20] Peng, D. L., Liu, Y., & Wang, C. (1999). How is access representation organized? The relation of polymorphemic words and their morphemes in Chinese. In J. Wang, A. W. Inhoff, & H. C. Chen (Eds.), Reading Chinese Script: A Cognitive Analysis (p. 65-89). Lawrence Erlbaum Associates.
[21] Peng, D., Li, Y., & Liu, Z. (1994). Identification of the Chinese two-character word under repetition priming condition. Acta Psychologica Sinica, 26(4), 393-400.
[22] Pollatsek, A., Hyönä, J., & Bertram, R. (2000). The role of constituents in reading Finnish compound words. Journal of Experimental Psychology: Human Perception and Performance, 26(2), 820-833.
[23] Pollatsek, A., Reichle, E. D., & Rayner, K. (2011). E-Z Reader 10: An updated model of eye-movement control in reading. Psychonomic Bulletin & Review, 18(4), 656-666.
[24] R Core Team. (2023). R: A language and environment for statistical computing [Computer software]. https://www.r-project.org/
[25] Rayner, K. (1998). Eye movements in reading and information processing: 20 years of research. Psychological Bulletin, 124(3), 372-422.
[26] Salverda, A. P., & Tanenhaus, M. K. (2010). Tracking the time course of orthographic information in spoken-word recognition. Journal of Experimental Psychology: Learning, Memory, and Cognition, 36(5), 1108-1117.
[27] Schreuder, R., & Baayen, R. H. (1995). Modeling morphological processing. In L. B. Feldman (Ed.), Morphological Aspects of Language Processing (p. 131-154). Lawrence Erlbaum Associates.
[28] Shen, W., Qu, Q., Ni, A., Zhou, J., & Li, X. (2017). The time course of morphological processing during spoken word recognition in Chinese. Psychonomic Bulletin & Review, 24(6), 1957-1963.
[29] Taft, M. (2003). Morphological representation as a correlation between form and meaning. In E. Assink & D. Sandra (Eds.), Reading Complex Words (p. 113-137). Kluwer Academic Publishers.
[30] Taft, M. (2004). Morphological decomposition and the reverse base frequency effect. The Quarterly Journal of Experimental Psychology Section A, 57(4), 745-765.
[31] Taft, M., & Forster, K. I. (1975). Lexical storage and retrieval of prefixed words. Journal of Verbal Learning and Verbal Behavior, 14(6), 638-647.
[32] Taft, M., Huang, J., & Zhu, X. (1994). The influence of character frequency on word recognition responses in Chinese. In H. W. Chang, J. T. Huang, & C. W. Hue, et al. (Eds.), Advances in the Study of Chinese Language Processing (p. 59-73). National Taiwan University.
[33] Tsang, Y. K., & Chen, H. C. (2014). Activation of morphemic meanings in processing opaque words. Psychonomic Bulletin & Review, 21(5), 1281-1286.
[34] Zhang, B., & Peng, D. (1992). Decomposed storage in the Chinese lexicon. In H. C. Chen & O. J. L. Tzeng (Eds.), Language Processing in Chinese (pp. 131-149). North-Holland.
[35] Zhou, X., & Marslen-Wilson, W. (1994). Words, morphemes and syllables in the Chinese mental lexicon. Language and Cognitive Processes, 9(3), 393-422.
[36] Zhou, X., & Marslen-Wilson, W. (1995). Morphological structure in the Chinese mental lexicon. Language and Cognitive Processes, 10(6), 545-600.
[37] Zou, L., Packard, J. L., Xia, Z., Liu, Y., & Shu, H. (2019). Morphological and whole-word semantic processing are distinct: Event related potentials evidence from spoken word recognition in Chinese. Frontiers in Human Neuroscience, 13, 133.