Directives form a constellation of communicative functions with the shared goal of steering the attention, actions, thoughts and plans of the recipient. They seek to trigger a change in their attention, actions, thoughts or plans (e.g., Put your phone down) or take more of a facilitative, supportive function in which a typically more experienced other guides the attention, actions, thoughts or plans of the recipient. Directives represent a ‘world to word direction of fit’ (Searle, 1983) or, even more succinctly, directives are ‘designed to get someone else to do something‘ (Goodwin, 2006). Since directives involve controlling the actions and thoughts of the recipient, they require consideration of politeness norms, social status, societal expectations, context, activity and culture. For this reason, directives can appear in a range of forms (e.g., Ervin-Tripp et al., 1990; Gaskins and Frick, 2023). For example, in English directives can take forms ranging from the most complex of yes/no question, such as Would it be possible for you to get the report to me by tomorrow?, to a single word utterance, such as Scalpel!. Despite their centrality to communication in any social context, directives are often viewed as the poor cousin of the more cognitively sophisticated declarative. Directives are not species specific (Gómez, 2007; Tomasello, 2007), they pattern less strongly with other cognitive skills such as theory of mind (Camaioni et al., 2004) and, from a language development perspective, were traditionally viewed as less facilitative with regard to caregiver speech as input than declaratives (e.g., Nelson, 1973; Tomasello & Farrar, 1986), though this claim has been revisited in later work (Pine, 1992; Rantalainen et al., 2022).
A more detailed consideration of directives, not only in terms of their production, but also the types of responses that they trigger, provides a rich arena for the study of caregiver-child interaction (Filipi, 2018; Gaskins and Frick, 2023; Goodwin and Cekaite, 2018). The directive-response sequence comprises a key adjacency pair (Schegloff, 2007) which emerges early in interaction between caregiver and child (e.g., Filipi, 2018; Gaskins and Frick, 2023). Adjacency pairs play a crucial role in linguistic interaction, as the form of the first pair part (FPP) projects the type of response required in the second pair part (SPP) (Schegloff and Sacks, 1973). Unlike other forms of adjacency pairs (e.g., information questions), however, the response (or SPP) can take the form of a verbal or non-verbal action. For example, the directive Pass the salt could be responded to appropriately by use of an acknowledgement token (e.g., Okay) along with the appropriate action, or just the action alone (i.e., passing the salt). The study of directives then enables us to explore why children produce linguistic action in contrast to how. That is, directives provide the perfect context for exploring and understanding children’s selection of social action. In the current study, we focus on Now directives, defined as directives which demand a change to activity or state in the here and now (Goodwin, 2006; Szczepek Reed et al., 2013).
Child language development research is a complex domain with differences in theory, methodology and even assumptions about the end state of the developmental process, leading to a quite fragmented research landscape. Language development is often viewed in isolation from communication more broadly, and the issue of why, given the choice, children use linguistic responses as opposed to non-verbal action responses and individual variations within this domain, has been overlooked. The question may sound mundane but, to paraphrase Sacks et al., (1974), it is the mundane questions that usually come up with the most interesting answers. In this paper we combine usage-based and language socialisation approaches to explore the patterns of, and factors relating to, children’s directive response types.
Researchers taking a socio-interactionist or usage-based approach to language development highlight the centrality of linguistic interaction in all aspects of the process and document the role of eye gaze, gesture and complex social skills such as role reversal (Tomasello, 2003) in language use and development. Even though, ‘language and interactional skill are learned hand-in-hand’ (Casillas et al., 2016) , arguably, our analyses of language development have a tendency to decouple these domains. For example, researchers working in usage-based approaches present detailed accounts of how children gradually extract linguistic structure from interaction with more experienced others through the use of cognitive processes such as analogy, categorisation and intention-reading. These studies typically explore how children learn the formal aspects of language such as vocabulary and syntax; in other words, the question relates to how children learn the structural components of language as opposed to why they choose to use these linguistic devices. Directives provide an interesting case study in this respect. Studies suggest that young children have a bias towards producing action responses in the early stages of development and that these are accepted as meaningful communicative bids by their interlocutors (e.g., Shatz, 1978). Communicative contexts which involve a linguistic response are subsequently shaped through interactive sequences during which the caregiver suggests that ‘something more is expected’ (Shatz, 1978, p. 277). In addition to the signals presented by the caregiver during joint attention interactions, it could also be suggested that children learn how to respond to directives through interaction, by monitoring, processing and subsequently using the types of directive-response patterns found during caregiver-child interaction. This would include storage and usage of frequently used construction types and the associated response types.
Studies in the field of language socialisation share an emphasis on the role of interaction in terms of communicative development but focus more on the ways in which language use contributes to social action and interaction more broadly. Goodwin and Cekaite (2018) present an ethnographic, embodied account of directive sequences in family settings with data collected from middle-class families in the United States and Sweden. Their detailed account demonstrates the ways in which children are socialised into embodied practices of action which centre around the everyday activities of family life. Directive-response sequences are achieved by means of multimodal actions including touch, eye gaze, sensitivity to bodily orientation, gesture and, often but not always, linguistic action. These local actions are situated within frequently occurring everyday events or routines (e.g. Bruner, 1973, 1975; Goodwin and Cekaite, 2018), which also shape the ways in which directive sequences are conducted.
Gaskins and Frick (2023) present a multimodal analysis of directive sequences produced in a bilingual triadic interactional context involving a child aged 1;8-2;4, their primary caregiver and their grandparent. The analysis demonstrates the child’s use of both verbal and non-verbal actions, namely gesture and eye gaze, in response to their caregivers’ directive utterances, as well as the value of engaging children in adult-to-adult conversational sequences, particularly children being raised in heritage language homes. De León (2017) investigates correctional directive sequences used within enskillment activities in Mayan families. Their analysis identifies the presence of both linguistic and non-linguistic modalities in directive sequences. Importantly, the child’s agency plays a key explanatory role in the analysis, as the process of enskillment is underpinned by the role of the child in exploring a new task under the guidance of a more experienced other. That is, in de León’s data, directives are not always – in fact are rarely – used to provide explicit instruction, but instead to ‘retune perspective and attention’ (de León, 2017) and, consequently, call for action as opposed to words in their responses.
The aim of the current study is to investigate the ways in which children respond to directives, and the factors associated with their response choice. Specifically, we investigate directive-response sequences in the format of first pair part (FPP) and second pair part (SPP) adjacency pairs. The study is based on longitudinal corpus data consisting of naturalistic caregiver-child interaction recorded during a range of activities. We begin with a quantitative structural account of directive-response sequences across the corpus, focussing on the relationship between directive construction type and response type. We then present a multimodal account of directive-response sequences in two of the children from the corpus in order to examine the contribution of multimodal behaviours and context in the child’s responses to directives. We bring together the results of the mixed-methods study to provide a holistic account of directive-response type sequences and factors associated with response type selection.
The current study is based on data from the New England corpus (Ninio et al., 1994), accessed from the CHILDES website (https://talkbank.org/childes/access/index.html). The corpus comprises a rich dataset of naturalistic, longitudinal data from 52 caregiver-child dyads at the beginning of the data collection period. The dyads were recruited from the New England area in 1980s and included an equal number of girls and boys recruited from across the socio-economic spectrum. Children were video-recorded while engaged in four activities with their caregivers (book reading, drawing, puppet play and toy play) at ages 14, 22 and 32 months. Direct observations from the video recordings were a central component of the analyses. Each recording session lasted ten minutes. The current study is based on the 32-month sample, of which data from 41 caregiver-child dyads were available.
The New England corpus is fully transcribed in CHAT format (the standard transcription conventions used on the CHILDES system) and also includes communicative intent coding by means of the Inventory of Communicative Acts – abridged (INCA-A). INCA-A is a detailed coding taxonomy consisting of two levels of categorisation: the interchange level, which categorises the general interactive intent of an interactive segment (e.g., discussing a joint focus of attention or negotiating the immediate activity), and the speech act level. The full INCA-A coding taxonomy can be found in the Appendices. For the purposes of the current study, we focussed on sequences within which the first pair part (FPP) took the form of a linguistically produced directive. In INCA-A, these are represented at the speech act level by the category ‘Request, propose, suggest’ (RP). For context, the category ‘Request, propose, suggest’ (RP) is typically embedded within one of two interchange levels, referred to in INCA-A as ‘Negotiating the Immediate Activity’ (NIA) and ‘Directing the Hearer’s Attention’ (DHA). Directives embedded in both interchange types were included in the present study. Examples of each category combination are shown in 1a and 1b.
1a. Negotiating the Immediate Activity (NIA): Request, propose, suggest (RP)
Caregiver: Sit down.
1.b Directing the Hearer’s attention (DHA): Request, propose, suggest (RP)
Caregiver: Look at the moon and the stars!
For Analyses 1 and 2, we categorised all FPP caregiver directives and child directives into six basic construction types based on Cameron-Faulkner et al., (2003), as displayed in Table 1. Categorisation was based on the transcription files. The responses were coded according to the taxonomy displayed in Table 2. In Analysis 1, statistical analyses were conducted to examine whether there were differences between caregivers and children in their use of different types of directive constructions. Because the frequency data for directive constructions were non-parametric and did not meet the assumption of normality required for parametric tests, the Mann-Whitney U test was used as an appropriate non-parametric alternative to the independent sample t test for comparing two independent groups. In Analysis 2, statistical analysis was conducted to examine the main effects of directive construction type, SPP participant type (caregiver vs child) and their interaction on the response type (verbal vs non-verbal). Logistic regression was chosen because the dependent variable was binary (verbal vs non-verbal), making it a suitable for modelling the effects of directive construction type and SPP participant type on the response type.
Table 1. Directive construction types
| Directive codes | Definition | Example |
|---|---|---|
| Fragment | Any utterance (1) without a subject and predicate and (2) not beginning with a verb | In the box |
| Imperative | Imperative – subjectless construction beginning with a verb | Sit down Pass the salt |
| Let(s) x | Any directive starting with Let(’s) | Let’s sit down Let me do it |
| Subject verb (x) | Sentence with a subject, verb and any following constituents | You hold the crayon with Mummy |
| Yes-no question | Yes-no question | Shall we draw a picture? Can you open the box? |
| Wh-question | Wh-question | Why don’t you put that away? |
Table 2. Response types
| Response codes | Definition | Example (response to directive shown in bold) | ||||||
|---|---|---|---|---|---|---|---|---|
| Nonverbal | Response with a non-linguistic action only |
|
||||||
| Verbal | Response with a linguistic action |
|
We excluded all directive-response sequences in which both turns were taken by the same person (e.g., the caregiver produced both the directive and response), sequences which could not be fully transcribed (e.g., those denoted by xx or xxx in the transcripts) and sequences in which the intent or meaning of the response was not clearly related to the directive.
Analysis 3 focusses on the directive-response adjacency pairs of two caregiver-child dyads from the corpus and was based on video recordings available on CHILDES. The selection of the dyads was based on overall frequency of caregiver directives, with one caregiver being at the high range of directive frequency and the second at the lower end of directive frequency. We focussed our analysis on directive-response sequences in which the caregiver produced the directive in FPP and the child responded in SPP. In this analysis, the aim was to follow a talk-in-interaction approach in the spirit of Goodwin and Cekaite (2018), Goffman, (1983), and Sacks, Jefferson and Schegloff (e.g., Sacks et al., 1974). The analyses presented are based on the transcriptions with additional contextual notes and observations added from the video recordings.
In the current study, we investigate patterns of caregiver-child interaction found in directive-response sequences. Our results open with the quantitative structural account of directive-response sequences and then move on to a multimodal account of directive-response sequences in two of the children from the dataset.
Caregivers produced 180 directives overall (M = 5, SD = 5.01) and the children produced 207 (M = 6.27, SD = 4.03). Figure 1 displays the frequency of construction types produced by the caregivers and by the children. Statistical analysis indicated a significant difference between caregivers and children in the use of Fragments (U = 23, p = 0.045), Imperatives (U = 660.5, p = 0.018), SV(X) (U = 149.5, p = 0.132) and Questions (U = 153.5, p = 0.017), but not in the use of Let’s (x) constructions (U = 139.5, p = 0.382).

Figure 1. Average number of construction types in the caregivers’ and the children’s directives
There was a clear asymmetry between the caregivers and children with regard to directive construction type. While Imperatives comprised the most frequently used construction type by caregivers, SV(X) constituted the most frequently used construction by the children.
There was also a clear asymmetry with regard to response type, with caregivers producing verbal responses more frequently than non-verbal, and children producing more non-verbal responses than verbal (B = -3.31, SE = 0.30, z = -11.14, p <0.001).
Figure 2. Average number of non-verbal and verbal response types for the caregivers and children
Findings from the logistic regression analysis indicated that there was no effect of directive construction type (X2 (4) = 2.60, p = 0.628) or the interaction between directive construction type and SPP participant (X2 (4) = 5.29, p = 0.259) on response type. The main effect of SPP participant was significant (X2 (1) = 165.22, p < 0.001), suggesting children used more non-verbal responses than verbal responses in comparison to their caregivers. This indicates that the children’s selection of response type is not based on the linguistic construction used in the caregiver FPP of the directive-response sequence.
All caregivers and most children produced directive and response adjacency pairs frequently in the corpus. Asymmetry between the caregivers and children with regard to the types of directive constructions used and response types was attested. There was no relationship between construction type and response type, which is surprising given the primacy of adjacency pairs in interaction-focussed analyses. The results of Analysis 2, therefore, indicate that the formal features of the adjacency pairs, namely the grammatical construction used to express the directive FPP, does not appear to be the driving force behind the children’s response type. Also, the predominance of verbal responses in the caregivers’ SPP indicates that caregiver-child modelling in general also cannot account for the selection of response type by the children. The quantitative analysis indicates that children do not learn everything about language through language and that learning how to co-participate in linguistic interaction involves more than storage of high frequency linguistic adjacency pairs. Additional factors need to be taken into account to understand how children learn to fulfil their role as co-participants. In the next section, we qualitatively analyse the directive-response sequences of two caregiver-child pairs to examine the roles of embodied interaction and context in the performance of directive-response pairs.
Interaction between caregivers and children is a deeply embodied activity, which is situated within the material world (Goodwin, 2006; Goodwin and Cekaite, 2018). Therefore, embodied actions (e.g., touch, eye gaze, gesture, physical orientation) need to be considered alongside the affordances of objects in the environment (Rodríguez et al., 2018) and the expectations associated with the activity itself (Tarplee, 1993). Through this approach, the production of directive-response adjacency pairs can be viewed as authentic co-production of meaningful social action. In the current analysis, we consider the range of work (Filipi, 2018) achieved by the caregivers’ directives, namely calls for action (e.g., directing the child to engage in a specific action), correctives (e.g., shaping the child’s current action to fit more appropriately with the desired outcome of the action) and attentional directives (e.g., directing the child to attend to a feature of the environment). We then examine the type of response produced by the child and the extent to which the response type is considered acceptable by the caregiver through the consideration of the third pair part (3PP) of the exchange. The first pair part (i.e., caregiver directive) and second pair part (i.e., the child’s response) are in boldface throughout. To reflect the substantial variation across the sample, the analysis focusses on two caregiver-child pairs with contrasting directive frequencies; one dyad which included few directives (dyad 1: New England corpus file 31) and, in contrast, the dyad with the highest number of caregiver directives in the sample (dyad 2: New England corpus file 36).
Caregiver 1 produces nine directives which trigger SPP responses from their child. Most of the attentional directives took place during the book reading activity, while most of the action directives were produced during the puppet, drawing and toy play activities. One of the action directives and the corrective directive were produced during the transition between the drawing and toy play activity. The child produced an almost equal number of verbal and non-verbal responses. Raw frequency counts of directive type and child response type are displayed in Table 3. In the following section, we discuss three extracts and examine the extent to which embodied features of the interaction (i.e., eye gaze, touch, bodily orientation), agency and activity type reflect the response type of the child. We also examine the nature of the 3PP as a reflection of the caregiver’s acceptability of the child’s response. For each extract, we begin with a contextual summary, the extract itself and then its detailed analysis.
Table 3. Raw frequency of directive type and response types produced by dyad 1
| Caregiver directive type | Child response type | |
|---|---|---|
| verbal | non-verbal | |
| action | 2 | 2 |
| corrective | 0 | 1 |
| attention | 2 | 2 |
Extract 1. Transition between the Book Reading and Puppet Play Activity. In Extract 1, the caregiver and child have just put away the book from the book reading activity and have picked up the next activity box from a table/shelf in the corner of the room. The caregiver has picked up the box (labelled 2 puppets), which was a little too high for the child to reach comfortably. The caregiver takes the puppet activity box down from the table/shelf and addresses the child with the utterance ‘Read it’ while moving the box to the child’s eye level (line 3).
Caregiver: Do you know what this says?
(child’s gaze follows the box as the caregiver brings it the child’s eye level and smiles)
Caregiver: Read it. ((caregiver holds box in front of child))
Child: arhh ((vocalises with a letter-type sound produced with rising intonation)).
(Caregiver places the box on the floor and sits down behind it. Child sits down next to the caregiver touching the box as they sit)
Child: This says {.} this one.
Caregiver: Yeah, this is the one.
During the transition between activities, the caregiver and child are in a relatively distal configuration, as they are standing on either side of the table on which the box is located. The caregiver’s face is out of the camera shot so eye gaze cannot be ascertained, but the child’s gaze is fixed firmly on the box held by the caregiver (which is out of shot in the image above) as the caregiver moves it from the table. The distal orientation and the child’s maintenance of eye gaze to the object contrast to many of the other directive sequences which have a closer orientation and joint attentional eye gaze reminiscent of the intertwining patterns identified in Goodwin and Cekaite (2018). The caregiver uses an imperative construction to express an attentional directive as they prompt the child to read the words on the new activity box. The directive is preceded by a yes/no question (line 1), which draws attention to the words on the box. Both utterances are accompanied by the movement of the box into the child’s field of vision. As this is the second activity in the session, the child is aware that the words on the box relate to its contents but is not able to read the words. Both the caregiver’s directives involve verbs relating to verbal action (Do you know what this says? Read it.) and this, in itself, prompts the child for a linguistic response. However, it is also worth noting that, other than moving the box within eye level of the child, no other actions (e.g., pointing at the words, giving the box to the child, moving close to the child) are produced by the caregiver. In summary, the interaction appears distal in orientation and, also, in terms of action. The child responds with a verbal action (line 4) which sounds akin to a letter-sounding out strategy – a form of reading routine often practiced with young pre-literate children – though the letter sound does not fit with the word on the box. Thus, all aspects of the sequence (distal orientation, verbs of speech, movement of the object within the child’s gaze but not hands) call for a linguistic response from the child, even though the accurate production of the response is beyond the abilities of the child. The child follows up their SPP with a recasting of the response in the form of ‘That says {.} This one’, which is accepted by the caregiver as an appropriate turn in the sequence signalled by the repetition of the child utterance (line 7).
Extract 2. Getting to Know the Puppets. In Extract 2, the caregiver and child are engaged in the puppet play activity. They have just opened the activity box to find two puppets – a bird and the Sesame Street character Cookie Monster. The box is between but slightly in front of the caregiver and child. The caregiver has the Cookie Monster puppet on their hand and has just been looming over to the child playfully with the puppet. The child has the bird puppet on their hand.
Child: What’s that? ((child looks down at the puppet on their hand))
(child takes off the puppet, puts it back in the box and maintains gaze on puppet)
Caregiver: That’s a bird ((exaggerates a lean over to look at the puppet in the box)).
Caregiver: Lemme see that one. ((caregiver leans back while continuing to look at the puppet in the box)).
(child looks to the caregiver, smiles and then looks down to the bird puppet and makes an initial attempt to put the puppet back on their hand in response to the caregiver’s directive)
((caregiver leans towards child with Cookie monster puppet)) Here I come!
(child drops bird puppet back in the box, and smiling, looks towards and touches the Cookie monster puppet)
Extract 2 begins with the child putting the bird puppet on their hand with eye gaze directed towards the puppet. The caregiver is simultaneously leaning over to the child and crouching slightly to orient at the child’s eye level, and trying to get the child’s attention by extending their hand with the Cookie Monster puppet on it. The caregiver sits back and retracts the puppet as the child, while still looking at their bird puppet, says ‘What’s that’ in a playful and animated manner (line 1). The caregiver picks up on the child’s question (which appears self-directed) and eye gaze to the puppet as an opportunity to engage in joint attention. The caregiver responds to the child’s question (line 3) and then produces a directive which re-engages the child with the bird puppet (line 5). The example seems to be construed by the caregiver as a missed opportunity for joint engagement, which they then rectify through the use of an attentional directive. During the extract, there is considerable movement on the part of the caregiver in terms of leaning, crouching and extending the puppet towards the child but, as the directive is produced, their eye gaze moves to the child’s face and their body position becomes relaxed but upright, signalling a break in the puppet play activity and the requirement of some form of response from the child. The child responds by means of a non-verbal action (line 5) and this is then followed by a move-on turn by the caregiver (line 6) as they return to focussing attention on the Cookie Monster puppet. The presence of the move-on turn and consequent absence of a pursuit of response indicates that the child’s non-verbal turn is considered acceptable in the directive-response sequence.
Extract 3. Tidying away the Puppets - Closing the Box. In this example, dyad 1 have just finished the drawing activity. The completion of the activity is signalled by the child putting the paper in the activity box. The caregiver is seated next to the child and supervising the child’s tidying away of the paper and crayons. On this occasion, the caregiver produces an action directive to remind the child to put the lid back on the open activity box before putting it away and starting the next activity. The child has stood up and is looking down at the box, which is positioned slightly in front of the dyad. The caregiver is seated next to the box with their right arm resting on their leg near to the box.
Caregiver: There’s only one box left ((gazes to the paper and crayons box in front of them))
(child stands up and looks down at the paper and crayons box)
Caregiver: The cover? ((caregiver glances briefly to the box cover and taps it gently))
Child: The cover. ((child looks to the cover, crouches down, picks it up and attempts to put it on the box)).
Caregiver: The cover. The cover. (watches as child successfully puts the cover on the box).
In Extract 3, the caregiver and child have just completed a joint activity (drawing), which has been signalled as completed by the child’s action of beginning to put the paper and crayons away. The caregiver has accepted this action and is monitoring the child from their original sitting position as the child stands up. The child’s change in spatial orientation signals a break in the closely intertwined activity of drawing. Throughout the sequence, both caregiver and child gaze to the box containing the paper and crayons. The caregiver momentarily shifts their gaze as they look to the box cover, taps it gently and produces the action directive (line 3). The directive is thus signalled by means of conjoined embodied and linguistic action at a relatively distal orientation. The child responds to the directive with a verbal response which recycles the caregiver’s directive and is produced with falling intonation. The caregiver follows up the child’s SPP with an acceptance 3PP comprising the child’s response, but with corrective-style stress on the definite article. The 3PP therefore serves, on the one hand, to display acceptance of the verbal response as an appropriate response type but, on the other hand, also models the caregiver’s view of the exact form that the verbal response should take. It is interesting to note a number of differences between Extracts 2 and 3. Firstly, in contrast to Extract 2 there is no mutual eye-gaze between the co-participants and, secondly, there are differences in bodily orientation, as, in Extract 3, the caregiver is sitting down and the child is standing while, in Extract 2, both are sitting down side by side. Thus, Extract 3 indicates a more distal orientation between caregiver and child, which appears more similar to Extract 1 than 2.
Caregiver 2 produces thirteen directives that receive SPP responses from the child. Six of these are action directives, five are attentional directives and two are corrective directives. Most of the directives receive a non-verbal response from the child (see Table 4).
Table 4. Raw frequency of directive type and response types produced by dyad 2
| Caregiver directive type | Child response type | |
|---|---|---|
| verbal | non-verbal | |
| action | 1 | 5 |
| corrective | 0 | 2 |
| attention | 2 | 3 |
Six of the directives occurred in transitions between activities (four action and two corrective directives), while seven are part of the activities themselves (book reading, two attentional directives; drawing, three action directives; toy play, two attentional directives). We have selected two adjacency pairs from the toy play activity. This activity involves a toy house that opens up to reveal the interior rooms (kitchen, bedrooms, …) and contains loose items such as beds. The child explores the house and plays with it. The two extracts included here illustrate directive pairs with a non-verbal response and a verbal one. We again focus on the embodied features of the interaction (i.e., eye gaze, touch, bodily orientation), agency and activity type, as well as comment on the 3PP to gauge the caregiver’s acceptability of the child’s response.
Extract 4. Playing with the House – the Garage. In Extract 4, the caregiver and child have just started exploring the toy house. The child is seated in front of the toy house with the caregiver sitting next to the house with their legs folded next to them and their upper body oriented towards the child. The caregiver has asked the child to identify what the toy is, upon which the child not only identifies the object as a toy house, but also adds that it includes a garage. The caregiver then initiates a closer exploration of the garage. After failing to open the garage door themselves, the child instructs the caregiver to do it.
(child tries to open garage door but fails.)
Child: Open it.
(child leans back watching the caregiver’s hand open the garage)
((caregiver opens the garage, which makes a popping noise; both caregiver and child are looking at the garage)) Caregiver: Oh!
((caregiver moves their hand back, sits up straight and moves a little to the right, away from the toy house)) Caregiver: Look what’s in there. ((caregiver shifts their gaze from the garage to the child))
(child moves forward and reaches into the garage; caregiver turns and leans forward slightly to be able to look at the garage)
(child pulls a little car out of the garage; both child and caregiver shift their gaze to the car when the child moves it away from the toy house)
In Extract 4, the caregiver and the child alternate between performing actions and observing the actions of the other with regard to the object of joint attention, the toy house. Agency changes between them and the action as a whole is co-produced. Each of them makes physical space for the other to perform actions centred around the toy house, and requests them by directives (child ‘open it’ (line 2) and caregiver ‘look what’s in there’ (line 5)), creating a sequence of conjoined verbal and physical directives. Before the caregiver-initiated directive sequence, both caregiver and child are looking at the garage. The caregiver uses an attentional directive, ‘look what’s in there’ (line 5). The child does not give a verbal response but moves forward and reaches into the garage (line 6). During the directive-response sequence, the caregiver’s eye gaze shows a minimal shift from the garage to the child and back. When the child reaches into the garage, the caregiver leans forward to be able to look into the garage. 3PP is constructed through eye gaze and continuation of the intertwined activity. This extract is typical of many interactions between caregiver and child 2 during the activities. The caregiver and child are closely engaged in the activity with eye gaze, bodily orientation and action forming a seamless choreography.
Extract 5. Playing with the House – the Front Door. Extract 5 contrasts with Extract 4 as one of the small number of directive sequences in which the child in dyad 2 gives a verbal response. This extract occurs roughly a minute and a half later than the interaction in Extract 4 during the toy play activity. The caregiver is in the same physical position – sitting next to the toy house with their legs folded next to them oriented towards the child. The child is oriented towards the caregiver, with their feet touching the caregiver’s knees. The child has just closed up the toy house and looked up at the caregiver. The child has one hand in their lap and the other hand is away from the toy house but not visible, as it is behind their leg or body.
(caregiver leans forward and turns, shifting their gaze to the house)
((caregiver moves their hand to open the front door of the toy house)) Caregiver: Look at the front door. (child’s gaze follows caregiver’s hand to stop on the door)
(caregiver closes and opens the front door multiple times and rings the bell)
Child: Yeah. (child tilts head to get a better look at the door, while the caregiver continues to play with the door)
During the interaction in Extract 5, the child is not physically engaging with the toy house. They are sitting back, while the caregiver leans forward, moving in front of the child to touch the front door of the toy house (line 1). As in Extract 4, the caregiver utters an attentional directive, ‘look at the front door’ (line 2). However, unlike in Extract 4, the directive is accompanied by the physical manipulation of the target of the attention – the front door – by the caregiver themselves. The caregiver continues to manipulate the door after the directive is uttered (line 3). The caregiver’s gaze is fixed on the front door. The child’s gaze follows the caregiver’s hand movement towards the door and observes the initial manipulation of the door (line 2). Different from Extract 4, there is no opportunity for the child to respond non-verbally by taking over the action. Because of the unchanged physical position of the caregiver, there is no opportunity for the child to join the physical action. The child instead verbally responds, ‘yeah’, acknowledging engagement with the directive (line 3). The verbal response is accompanied by a non-verbal response, in which the child physically reorients themself to look closer at the object of attention. There is no obvious 3PP.
Child 2 only rarely gives a verbal response in the caregiver-initiated direction sequences. In this particular example, the less usual, verbal response may be the result of the specific situation in which the child is not (invited to become) physically involved in the action. The caregiver’s attention and gaze are on the front door. It is possible that the child is aware of the fact that the caregiver is unable to ‘see’ a non-verbal response, as the caregiver is looking away from the child. Even though both caregiver and child have a joint focus of attention and are involved closely in the activity, the caregiver’s gaze and orientation do not allow for the intertwined action seen in joint play in Extract 4. A different way of looking at the contrast between the two extracts is that in Extract 4 we are seeing a common play routine between caregiver and child 2, in which a non-verbal response is an established part of the routine. In Extract 5, by contrast, the deviation from the routine prompts a different, verbal response.
Extract 6. Closing the Box with the Drawing Materials. The final extract we have selected for dyad 2 closely resembles Extract 3 involving dyad 1. The extract takes place towards the end of the session, and the child has had enough and wants to go home. The child has opened the box with drawing materials again but swiftly proceeds to put the paper back in the box and, subsequently, tries to close the box. However, the child has trouble fitting the cover in the correct way to do so. The caregiver is sitting on the floor with their legs folded and is oriented toward the child, who is some distance away. In between the caregiver and the child, scanning from the left – where the caregiver is sitting with only their knees and part of their face in the camera shot – to the right, are the open toy house and the box containing the paper and crayons. The toy house and box are slightly more towards the back, with the child sitting in front of the right-most corner of the box with their left leg in front of the box and their right leg underneath them. The child has put the paper back in the box and has grabbed the cover, which was lying on the floor on their right. The child has moved the cover on top of the box but with the long side of the rectangular cover facing towards themselves, while the box is facing them with the short side. The child attempts to push down the cover and get up on their left leg to be able to push down harder, but the cover does not fit. The caregiver is observing the action.
Child: Mommy! ((child looks at caregiver))
((child looks back at box, leans forward and pushes the cover down harder.)) Child: I wanna go.
(child gets up on both feet, stands up, lifts up the cover and moves it down again.)
Caregiver: Turn it around the other way.
(child with some effort turns around the cover 90 degrees and presses it down again, while squatting down.)
(child keeps trying to fit and push down the cover unsuccessfully with both hands.)
Caregiver: The other way honey.
In Extract 6, there is more physical distance between the caregiver and the child than in Extracts 4 and 5 – they are at opposite ends of the shot frame. After calling out to their mother, the child’s attention and gaze stay fixed onto the box. The caregiver is observing the action in a more distant way, letting the child try to fit the cover. The caregiver does not immediately respond to the child-initiated directive ‘Mommy!’ (line 1). After several attempts, the caregiver intervenes with the corrective directive ‘turn it around the other way’ (line 4). The child immediately responds non-verbally by starting to turn the cover, whilst maintaining their gaze on the box (line 5).
The interaction in this extract is different from Extracts 4 and 5 in that caregiver and child are not positioned closely and involved in a joint activity. The orientation is more distal and only the child is physically engaged in the closing of the box. The child’s action is ongoing and the function of the caregiver’s corrective directive is to further the action by re-directing it. As the non-verbal response from the child does not result in a successful outcome – the cover is turned 45 degrees too far – the caregiver follows up with a further corrective directive (3PP).
The FPP-SPP pattern illustrated in this extract is also found elsewhere in the recording of this dyad in activities where caregiver and child are more closely oriented. The caregiver uses directives to guide the child through the activity and uses attention and action directives to prompt the child to move on to a next action in an activity sequence or to better perform an action, e.g., in the drawing activity the child is drawing around their hand and the caregiver instructs to ‘color all around it’. These directive sequences are reminiscent of the ‘fine-tuning’ behaviours that are central in the enskillment process (e.g., Ingold, 2000; De Léon, 2017; Goodwin, 2007) in that the sequences draw the child’s attention to the specific aspects of the task and provide cues for successful completion of the task at hand. The directives do not call for a verbal SPP as such; the child’s response is non-verbal and consists of performing or continuing with the action. Furthermore, the caregiver’s response indicates that the child’s attention to the activity itself is an acceptable response to the enskillment episode. The higher frequency of directives sequences for dyad 2 appears to be due to the use of directives with this action-supporting function.
The aim of the study was to present a multi-methods approach to the analysis of directive sequences between caregivers and their three-year-old children. The emphasis of the analysis was placed on the nature of children’s responses in the second pair part of the directive sequence. We investigated the extent to which the linguistic construction of the directive affected the response type produced by the children, and then moved on to conduct a more detailed talk-in-interaction account of caregiver-child sequences to identify factors that determine the child’s response. The analyses indicate the strength of taking both approaches. In the discussion below, we focus on two inter-related findings, namely, linguistic asymmetry between caregiver and child directives and responses, and the contribution of bodily orientation, eye gaze and the nature of the activity or routine in the co-construction of caregiver-child directive sequences.
In the construction-based analysis of directive sequences we presented a quantitative analysis of the types of linguistic structures used to express directives in both caregivers and children, the response types and the extent to which directive construction type corresponded with response type. Given the centrality of adjacency pairs in everyday communication, we may predict that certain directives expressed as questions may pattern more closely with verbal responses than imperatives. Indeed, in a usage-based approach to development, it would be expected that children learn both linguistic information and patterns of usage directly through interaction.
Overall, the results present a clear asymmetry in directive construction types and response types in the turns of caregivers and their children and, also, a lack of correlation between construction type and response type. Caregivers and children used a wide range of constructions in order to express the first pair part of the directive sequence. Questions and imperative constructions constituted the most frequent construction types for the caregivers, while subject-verb and fragment constructions were more frequent for the children. We suggest that caregivers’ use of questions reflects their attempts to give the children agency in the directive sequences and co-construct the activity as opposed to directing. This could be viewed as a specific way of socialising children into a more active role in social interaction. Children’s use of subject-verb constructions, however, highlights their use of directives to demand a change in the state of affairs, with utterances such as I want (to) x or I need (to) x being typical in the children’s sample. This reflects the power dynamic between the co-participants; while the child’s role is to express what they want out of the activity, the caregiver’s role is to shape and socialise the child’s ways of being within a particular activity.
The absence of a relationship between questions expressed as directives and children’s response type is instructive in a further way. It indicates that children are basing their response on the function of the FPP as opposed to its grammatical form, which has been established to strongly prompt a yes/no answer (e.g., Filipi, 2018).
Asymmetry is also attested in the response-type analysis, with caregivers overwhelmingly producing verbal responses to their children’s directive, and children being more likely to produce non-verbal responses. This asymmetry may reflect the different roles and level of agency ascribed to the caregiver and child. While the child has agency in terms of producing directives, the caregiver has the power to either approve or reject these propositions and this is done explicitly through language. The caregiver’s directives, in contrast, have the purpose of directing or guiding the activity. Therefore, non-verbal responses (i.e., following the directive) are used and, as we show in the qualitative analysis, accepted as appropriate second pair parts by the caregivers.
Accordingly, the lack of correlation between directive type and response type is not at all surprising. While the sequences can be categorised as ‘directive-response’ adjacency pairs, the children’s use and formulation of these sequences is not simply a mirror image of their caregivers’ productions. It is motivated by an understanding of the purpose of the directives and differences in agency and role.
The talk-in-interaction analysis of the directive sequences supports these findings and further illustrates how the situated and embodied nature of the interactions shapes the verbal and non-verbal actions of both co-participants. In particular, the analysis highlights the role of bodily proximity and eye gaze, on the one hand, and of the nature of the overarching activities and routines, on the other, in determining children’s responses to their caregivers’ directives. While it may be tempting to consider the contribution of each feature and action in isolation, the data clearly lead to an analysis in which all aspects of the directive sequences are interconnected, intertwined and, consequently, framed as a constellation of actions.
Many of the extracts show alignment of bodily orientation, eye gaze and linguistic action. For example, in Extract 1, the caregiver’s directive is produced within a relatively distal orientation from the child. The directive itself contains explicit prompts for a linguistic response and is accompanied by the caregiver moving the box within the child’s field of vision. Together the actions, both linguistic and non-linguistic, call for a linguistic response. In contrast, in Extract 4, the caregiver and child are positioned close together and deeply engaged in a joint attentional state. The intertwined context provides the opportunity for a non-verbal response, which is accepted by the caregiver. In Extract 2, in which the caregiver interrupts their own play with the Cookie Monster puppet, the caregiver not only utters an attentional directive but shifts their eye gaze to the child’s puppet and reorients their body. The caregiver’s use of bodily proximity, eye gaze and linguistic action creates a multimodal, laminated (Goodwin, 2013) structure, which does not always require or call for a linguistic second pair part.
The situated nature of the activity also contributes to the ways in which the caregiver-child dyads play out the directive-response sequences. For example, the directives found in the book reading activity (not included in Analysis 3 due to their very specific nature) tended to take the form of attentional premoves for subsequent labelling sequences (cf, Tarplee, 1993) and, consequently, triggered attention to the book as opposed to a verbal response. Another example can be found in Extract 4 where it is not the form of directive, but its situated nature that triggers the response. There the caregiver uses attentional directives for ‘playing with the house’ where, because the intended responses are to play with, action directives might have been expected. The extract suggests that the child knows what is expected as a response, that is, to physically engage with the element the caregiver is drawing attention to. The directive sequence is embedded in a play routine which constitutes shared knowledge between this caregiver and child.
On the basis of the extracts discussed, we can present some suggestions about what may influence the choice of verbal versus non-verbal response. Firstly, as the discussion of Extracts 1 and 4 already shows, the choices of response appear to attune to what is appropriate for the overarching activity or routine. Extract 5 provides an interesting perspective on this. There the caregiver directs the child to attend to the front door of the toy house but does not move out of the way and keeps touching the door themselves. With the option of a non-verbal response removed and a deviation from the routine, the child provides a verbal acknowledgement.
Secondly, directives with a distal orientation where primary agency is ascribed to the caregiver may be more likely to invite a verbal response, as opposed to those in which the two participants are closely intertwined. Extracts 3 and 4 illustrate the opposing end of the scale. In Extract 3 the caregiver used a fragment directive (The cover) to prompt the child to place the cover on the activity box before putting it away. The action of putting the cover on the box was a relatively novel action and eye gaze of both the caregiver and child was solely directed towards the box and the cover. In this instance, the child responded both by performing the appropriate action but also by providing a verbal response (recycling of the caregiver’s phrase the cover). In Extract 4, by contrast, the caregiver and child are positioned close together and exhibit a strong degree of triadic joint attention, as they look to each other and to the toy throughout the extract. Agency is shared between caregiver and child as they co-construct an exploration activity based on the affordances of the toy. The intertwining of caregiver and child results in a non-verbal action response by the child.
Thirdly, directives which serve to guide and facilitate the activity may be more likely to receive a non-verbal response. For example, in Extract 6, the caregiver provides a corrective directive that ‘retunes’ the child’s attempts to put the cover on the box. In this type of enskillment situation in our extracts, it was, interestingly, possible to have distinct eye gaze focus on the object/event in question alongside physical distance between the participants. This indicates the importance of considering the multimodal context of the interaction between caregiver and child when accounting for the response type of the child.
The observation that the first pair part of the directive understood in all its multimodal and contextual complexity does not always require or call for a linguistic second pair part then contrasts with the notion that learning how to co-participate in conversation is a linear process whereby non-linguistic action is superseded by linguistic action in all contexts and based on the structural nature of the adjacency pair. The current analyses indicate instead that learning how to respond to directives involves socialisation into the situated, embodied, multimodal orchestration of social activity.
Directive sequences play a key role in everyday interaction. An important consideration is how children learn to co-participate in these sequences. Our analysis demonstrates that, while both caregivers and children produced both directives and responses to directives, there was a clear asymmetry in the types of strategies used, and that the selection of construction and response type was not a consequence of input frequency but, instead, a reflection of social roles, expectations and agency of caregiver and child. The linguistic action is played out in a situated and embodied context, where non-verbal as well as verbal responses can be called for. Therefore, in order to understand how children learn from experienced others, we need to look beyond the suggestion that children learn only and directly from the language that they hear and consider the role of embodied and situated interactional strategies, agency and power relations. These strategies and relations are deeply embedded within everyday caregiver-child routines and consequently shape both the child’s development and patterns of caregiver interaction.
Bruner, J. S. (1973). Organisation of early skilled action. Child Development, 44, 1–11. https://doi.org/10.2307/1127671
Bruner, J. S. (1975). From communication to language - A psychological perspective. Cognition, 3(3), 255–287. https://doi.org/10.1016/0010-0277(74)90012-2
Camaioni, L., Perucchini, P., Bellagamba, F., & Colonnesi, C. (2004). The role of declarative pointing in developing a Theory of Mind. Infancy, 5(3), 291–308. https://doi.org/10.1207/s15327078in0503_3
Casillas, M., Bobb, S. C., & Clark, E. V. (2016). Turn-taking, timing, and planning in early language acquisition. Journal of Child Language, 43(6), 1310–1337. https://doi.org/10.1017/S0305000915000689
de León, L. (2017). Emerging learning ecologies: Mayan children’s initiative and correctional directives in their everyday enskillment practices. Linguistics and Education, 41, 47–58. https://doi.org/10.1016/j.linged.2017.07.003
Ervin-Tripp, S., Guo, J., & Lampert, M. (1990). Politeness and persuasion in children’s control acts. Journal of Pragmatics, 14(2), 307–331. https://doi.org/10.1016/0378-2166(90)90085-R
Filipi, A. (2018). Making knowing visible: Tracking the development of the response token yes in second turn position. In S. Pekarek Doehler, J. Wagner, & E. González-Martínez (Eds.), Longitudinal studies on the organization of social interaction (pp. 39–66). Palgrave Macmillan UK.https://doi.org/10.1057/978-1-137-57007-9_2
Gaskins, D., & Frick, M. (2023). Embodiment in directive sequences: The case of triadic interactions in a Polish-English bilingual family. International Journal of Bilingualism, 27(1), 122–142. https://doi.org/10.1177/13670069221078334
Goffman, E. (1983). The interaction order: American Sociological Association, 1982 Presidential Address. American Sociological Review, 48(1), 1–17. https://doi.org/10.2307/2095141
Gómez, J.-C. (2007). Pointing behaviors in apes and human infants: A balanced interpretation. Child Development, 78(3), 729–734. https://doi.org/10.1111/j.1467-8624.2007.01027.x
Goodwin, C. (2013). The co-operative, transformative organization of human action and knowledge. Journal of Pragmatics, 46(1), 8–23. https://doi.org/10.1016/j.pragma.2012.09.003
Goodwin, M. H. (2006). Participation, affect, and trajectory in family directive/response sequences. Text & Talk, 26(4–5), 515–543. https://doi.org/10.1515/TEXT.2006.021
Goodwin, M., & Cekaite, A. (2018). Embodied family choreography: Practices of control, care, and mundane creativity. Routledge. https://doi.org/10.4324/9781315207773
Nelson, K. (1973). Structure and strategy in learning to talk. Monographs of the Society for Research in Child Development, 38(1/2), 1–135. https://doi.org/10.2307/1165788
Ninio, A., Snow, C., Pan, B., & Rollins, P. (1994). Classifying communicative acts in children’s interactions. Journal of Communications Disorders, 27, 157–188. https://doi.org/10.1016/0021-9924(94)90039-6
Pine, J. M. (1992). Maternal style at the early one-word stage: Re-evaluating the stereotype of the directive mother. First Language, 12(35), 169–186. https://doi.org/10.1177/014272379201203504
Rantalainen, K., Paavola-Ruotsalainen, L., & Kunnari, S. (2022). Maternal responsiveness and directiveness in speech to 2-year-olds: Relationships with children’s concurrent and later vocabulary. First Language, 42(1), 81–100. https://doi.org/10.1177/01427237211049585
Rodríguez, C., Basilio, M., Cárdenas, K., Cavalcante, S., Moreno-Núńez, A., Palacios, P., & Yuste, N. (2018). Object pragmatics: Culture and communication - the bases for early cognitive development. In A. Rosa & J. Valsiner (Eds.), The Cambridge handbook of sociocultural psychology (2nd ed., pp. 223–244). Cambridge University Press. https://doi.org/10.1017/9781316662229.013
Sacks, H., Schegloff, E. A., & Jefferson, G. (1974). A simplest systematics for the organization of turn-taking for conversation. Language, 50(4), 696–735. https://doi.org/10.2307/412243
Schegloff, E. A. (2007). Sequence organization in interaction: A primer in conversation analysis (Vol. 1). Cambridge University Press. https://doi.org/10.1017/CBO9780511791208
Schegloff, E., & Sacks, H. (1973). Opening up closings. Semiotica, 8(4), 289–327. https://doi.org/10.1515/semi.1973.8.4.289
Searle J. R. (1983). Intentionality: An essay in the philosophy of mind. Cambridge University Press. https://doi.org/10.1017/CBO9781139173452
Shatz, M. (1978). Children’s comprehension of their mothers’ question-directives. Journal of Child Language, 5(1), 39–46.https://doi.org/10.1017/S0305000900001926
Szczepek Reed, B., Reed, D., & Haddon, E. (2013). NOW or NOT NOW: Coordinating restarts in the pursuit of learnables in vocal master classes. Research on Language and Social Interaction, 46(1), 22–46. https://doi.org/10.1080/08351813.2013.753714
Tarplee, C. (1993). Working on talk: The collaborative shaping of linguistic skills within child-adult interaction (uk.bl.ethos.333751) [Doctoral dissertation, University of York]. https://etheses.whiterose.ac.uk/id/eprint/9808/
Tomasello, M. (2003). Constructing a language: A usage-based theory of language acquisition. Harvard University Press. https://doi.org/10.2307/j.ctv26070v8
Tomasello, M. (2007). If they’re so good at grammar, then why don’t they talk? Hints from apes’ and humans’ use of gestures. Language Learning and Development, 3(2), 133–156. https://doi.org/10.1080/15475440701225451
Tomasello, M., & Farrar, M. (1986). Joint attention and early language. Child Development, 57(6), 1454–1463. https://doi.org/10.2307/1130423
The corpus data are available on CHILDES, and the analysis code is available on OSF at https://osf.io/f7cnh
The study is based on corpus data from the CHILDES website and is freely available to researchers at https://childes.talkbank.org.
Thea Cameron-Faulkner conceived of the study and with Tine Breban contributed design, data analysis and writing. Ava Gordon categorized the data for analysis one and Chen Zhao conducted the statistical analyses. All authors revised and approved the manuscript and agree to be accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved.
We are most grateful for the work and dedication of Barbara Alexander Pan (1951-2011) and Catherine Snow in creating the New England corpus and making it accessible for language researchers worldwide. We also thank the caregivers and children for participating the corpus. We dedicate this paper to the memory of Barbara Alexander Pan.
License
Language Development Research (ISSN 2771-7976) is published by TalkBank and the Carnegie Mellon University Library Publishing Service. Copyright © 2026 The Author(s). This work is distributed under the terms of the Creative Commons Attribution-Noncommercial 4.0 International license (https://creativecommons.org/licenses/by-nc/4.0/), which permits any use, reproduction and distribution of the work for noncommercial purposes without further permission provided the original work is attributed as specified under the terms available via the above link to the Creative Commons website.