Abstract
This study integrates two lines of research: technologies as tools and technologies as social beings, under the theoretical framework of dynamic systems, to investigate the reciprocal dynamics between functional use and relational use of artificial intelligence (AI) voice assistants, and the mediating roles of self-disclosure and privacy concerns. A two-wave longitudinal survey was conducted among 354 AI voice assistant users across 2 months. Factor analysis results supported the conceptualization and operationalization of functional use and relational use of voice assistants. Results from the cross-lagged panel model confirmed that functional use and relational use reinforced themselves over time, respectively. Relational use increased subsequent functional use, and relational use reinforced itself through self-disclosure. Surprisingly, functional use did not increase subsequent relational use; instead, longitudinal mediation analysis showed that functional use reduced subsequent relational use due to the lack of self-disclosure. Furthermore, while self-disclosure increased subsequent privacy concerns, privacy concerns did not reduce subsequent self-disclosure.
Keywords
Introduction
Artificial intelligence (AI) voice assistants, also called conversational agents or chatbots, employ dialogue systems to enable human-like text or voice conservations with users (Oh et al., 2021; Poushneh, 2021). Powered by the advancement of artificial intelligence (AI), deep learning, and natural language processing, AI voice assistants are no longer just characters in science fiction movies; instead, they are embedded in a variety of products such as smartphones (mobile applications) and smart speakers in consumers’ homes, and they are becoming integral to our daily lives (Poushneh, 2021). According to a recent global market report (Statista, 2021), there were about 112 million voice assistant users in the United States in 2019, and by 2021, the number had reached 132 million, accounting for at least one-third of the US population.
How do people interact with their AI voice assistants? Two lines of research have shed light on human–AI interaction. One considers technologies primarily as tools to perform various functions, and technologies need to be adopted and used in order to achieve their utilitarian functions. Theoretical frameworks in this stream include the diffusion of innovation theory (Rogers, 2003), the domestication theory (Silverstone and Haddon, 1996), and the more recent theory of the six-acceptance-phase model of interactive technologies (De Graaf et al., 2018). Another line of research, however, assumes that humans treat technologies as social beings and develop personal relationships with them (e.g. Croes and Antheunis, 2021). This line of research is primarily guided by the computers as social actors (CASA) paradigm (Reeves and Nass, 1996), with an abundance of research building on this paradigm (Nass and Moon, 2000; Reeves and Nass, 1996 for reviews, see Fox and Gambino, 2021).
In this current study, we propose to theoretically integrate these two lines of research under the framework of dynamic systems (Wang and Tchernev, 2012; Xu and Wang, 2021), to understand the dynamic processes of human–AI interaction over time. Specifically, we examine how users form relationships with AI voice assistants: how the functional use, that is, using voice assistants as tools to perform different functions, evolves into relational use, that is, using voice assistants as social companions and establishing bonding. Building on existing research of social penetration theory and human-chatbot interaction, we consider self-disclosure and privacy concerns as the key mechanisms in this dynamic relationship.
Accordingly, we conducted a longitudinal two-wave study across 2 months to examine how users interact and develop relationships with their voice assistants in real life. The human-AI relationship is not a one-shot experience; it is characterized by familiarity established through multiple interactions, which indicates that cross-sectional studies may not be well-suited for examining a human-like relationship. However, existing longitudinal studies on human–AI interaction mainly rely on qualitative methods, for example, interviews (Lee et al., 2020), or descriptive analysis (Croes and Antheunis, 2021). They did not provide a formal, quantifiable examination of the dynamic, iterative feedback effects or reciprocal effects of different user patterns over time. In addition, some longitudinal studies required participants to interact with an assigned chatbot on specific days for a very limited time (e.g. 5 minutes; Croes and Antheunis, 2021), which limited their ecological validity and generalizability. To fill these gaps, the current study examines how people engage with their voice assistants of their choice in real life over time.
Literature review
Functional use and relational use of AI voice assistants
In the current study, we conceptualize using technology as tools as the function use (De Graaf et al., 2018; Rogers, 2003; Silverstone and Haddon, 1996), and using technology as social beings as the relational use (Croes and Antheunis, 2021; Fox and Gambino, 2021; Reeves and Nass, 1996). Functional use considers the devices primarily as tools to perform various tasks and offer functional assistance. Popular AI voice assistants, for example, Apple’s Siri, Google Assistant, and Amazon’s Alexa, are designed and trained to perform well in specific domains, such as searching or navigation, and to accomplish clearly defined tasks. For most users, their interactions with AI voice assistants typically involve a great amount of task-oriented usage (Cho et al., 2019; De Graaf et al., 2018), such as checking weather or news, setting up alarms or appointments, ordering food, playing music, and so on. An increasing number of studies have shown that the functional use of voice assistants improves users’ task efficiency (Mahajan et al., 2021), life quality (Oh et al., 2021), and work performance (Kim et al., 2020).
Relational use of AI voice assistants, however, considers devices to be social actors and users can form attachments and bonds with the device. From this perspective, the technology exceeds its functional purpose and becomes a social companion (De Graaf et al., 2018), and users develop relationships with voice assistants as if they were humans (Reeves and Nass, 1996). There are two mechanisms explaining why users would treat voice assistants as social beings and form relationships with them. First, human beings have the fundamental need to socialize, feel connected, and develop relationships, even if that relationship is with a machine or a voice assistant (Reeves and Nass, 1996). According to the CASA paradigm (Reeves and Nass, 1996), human brains do not automatically distinguish human–machine interactions from human–human interactions, so they treat technologies as social beings. A decade of research building on the CASA paradigm has garnered support for this assumption that it is possible for humans to form relationships with a computer or machine (Gambino et al., 2020; Nass and Moon, 2000; Reeves and Nass, 1996).
Second, the technological affordances of AI-based technologies may facilitate relational interaction. For example, the social cues of voice assistants, such as voice, tone, humor, personality, a gender identity, and a name, are engineered to engender naturalness and trust (Humphry and Chesher, 2021; Obinali, 2019; Phan, 2019) and lead people to categorize them as “worthy of social responses” (Nass and Moon, 2000: 83). Visual anonymity also ensures that people feel safer and more comfortable with personal information, which stimulates relationship formation (Croes and Antheunis, 2021). Moreover, with advanced machine-learning algorithms, AI voice assistants are able to process complex data both in real time and over time to provide users with personalized feedback and conversations (Shumanov and Johnson, 2021), leading to greater user engagement and facilitating the formation of relational usage.
Importantly, functional use and relational use of voice assistants are two bivariate processes instead of a single bipolar process. That is, the two usage patterns are independent, although they may be correlated. For example, users may have both functional use and relational use of their voice assistants at the same time. Research on the temporal order of functional use and relational use has yielded mixed findings. Some studies proposed that relational development should be followed by functional use (Cho et al., 2019; De Graaf et al., 2018), whereas findings from longitudinal studies suggested that the social processes associated with relationship formation decreased over time (e.g. Croes and Antheunis, 2021). In the current study, we assume that these two usage patterns can coexist; they ebb and flow, and mutually influence each other over time, based on the dynamic systems’ perspective.
The dynamic processes of human–AI interaction
We propose to theoretically integrate the aforementioned research under the framework of dynamic systems (Wang and Tchernev, 2012; Xu and Wang, 2021) to understand the processes of functional use, relational use, self-disclosure, and privacy concerns over time. The dynamic system perspective originated in mathematics and engineering, and has been introduced by Wang and colleagues to explain the dynamics of communication processes, especially human–media interactions (Wang and Tchernev, 2012; Wang et al., 2011). This dynamic perspective resonates with the proposition in social penetration theory that social relationships involve dynamically evolving elements as a unified system (Altman and Taylor, 1973).
Specifically, two key features define a dynamic system: the feedback effect and reciprocal causality (Wang et al., 2011). Feedback effects refer to the self-reinforcing property of human cognition and behavior. For example, an individual’s pattern of using a voice assistant is likely to reinforce itself over time. This self-sustaining property of technology usage resonates with propositions in the reinforcing spiral of media use (Slater, 2007), and in the selective exposure paradigm (Knobloch-Westerwick, 2015). In fact, this history-dependent, self-sustaining property is the defining feature of a dynamic system (Buzsàki, 2006). Longitudinal studies of media usage have accumulated evidence of such reinforcement patterns in media multitasking (Xu et al., 2019), media choices (Knobloch-Westerwick, 2015), and exposure to misinformation (Tan et al., 2015). Applying the feedback effects to understand user interaction with AI voice assistants, we propose:
Hypothesis 1 (H1). Functional use and relational use of voice assistants should reinforce themselves over time, respectively.
In addition to the feedback effects, another fundamental feature of dynamic media processes is reciprocal causality. Reciprocal causality refers to the mutual influence between elements (Wang et al., 2011), and evidence of reciprocal causality in daily media use was demonstrated in longitudinal studies (Xu et al., 2019; Xu and Wang, 2021). Applying the reciprocal causality in the context of human–AI interaction, relational use would not replace subsequent functional use of AI; instead, it may further strengthen it. It is likely that the increase in bonding and affinity toward the voice assistant would translate into increased functional usage of it. Therefore, we proposed:
Hypothesis 2 (H2). Relational use of AI voice assistants should reinforce subsequent functional use.
Building on the reciprocal causality between relational use and functional use, it is also possible that after a period of functional use, people may feel a functional dependency on the technologies (Karapanos et al., 2009), and the technologies become part of their life (De Graaf et al., 2018). As functional interaction evolves, “the technology exceeds its functional purpose and becomes a personal object as people get emotionally attached to it.” (De Graaf et al., 2018: 2586). Thus, the functional use should increase subsequent relational use. But what is the mechanism of the transition between functional use and relational use? We then turn to social penetration theory, a theory on relational development in human relationships.
The role of self-disclosure
Social penetration theory proposes the process of bonding moves a superficial relationship to an intimate relationship, or an intimate relationship to a deeper relationship. Such relational penetration is accomplished through self-disclosure (Altman et al., 1981; Gibbs et al., 2006). Self-disclosure is defined as the act of communicating information about the self to another (Gibbs et al., 2006; Wheeless and Grotz, 1976). According to social penetration theory, self-disclosure plays a key role in the development of personal relationships as it fosters closeness, intimacy, and fondness (Taylor and Altman, 1987). The role of self-disclosure in relationship building can occur in various types of human-human relationships, for example, friendships (Valkenburg and Peter, 2009) and romantic relationships (e.g. Gibbs et al., 2006), in both mediated and face-to-face settings (Carpenter and Greene, 2015).
Increasing numbers of studies on human–AI interaction have studied the important role of self-disclosure. Previous research suggested that voice assistants had the advantage of anonymity and being non-judgmental to promote user self-disclosure (Ma et al., 2016), and people were more willing to share information with a virtual assistant than to human counterparts (Lucas et al., 2017). Another study found that self-disclosure to AI chatbots contributed to a stronger bonding over time (Lee et al., 2020). Evidence from an experimental study also supported the idea that people felt socially connected to the chatbot after emotional disclosure to it (Ho et al., 2018). A more recent qualitative study also revealed that users’ conversations with social chatbots moved from the sharing of superficial information to self-disclosure, and their relationship with the chatbot also strengthened (Skjuve et al., 2021). Taken together, greater functional interaction may lead to greater relational interaction of AI voice assistants, primarily through self-disclosure (Lee et al., 2020; Skjuve et al., 2021). Relationship development is also a reinforcing process, and the growth of social bonds is achieved by disclosing more intimate personal information (Altman et al., 1981; Taylor and Altman, 1987). Thus, the relational use of voice assistants is likely to reinforce itself through self-disclosure. Therefore, we propose:
Hypothesis 3 (H3). Functional use should increase subsequent relational use through self-disclosure to AI voice assistants.
Hypothesis 4 (H4). Relational use should increase subsequent relational use through self-disclosure to AI voice assistants.
The role of privacy concerns
The notion of privacy was added to the original social penetration theory by Altman et al. (1981) to explain the decrease of self-disclosure and the social de-penetration process. They argued that self-disclosure would not necessarily continue to deepen throughout a relational process. Rather, relationship parties might at times feel the need to step back and reduce self-disclosure due to privacy concerns (Altman et al., 1981). In the context of human–AI interaction, voice assistants constantly gather an enormous amount of personal information to provide more personalized user experiences (Pal et al., 2020), and personal information may be used by chatbot developers without user consent (Martin, 2018). Due to the increasing cases of data security breaches in recent years (Ayaburi and Treku, 2020), users may have privacy concerns with voice assistants and become cautious about self-disclosure to them (Foehr and Germelmann, 2020). Recent research showed that privacy concerns had a strong negative influence on users’ trust of voice assistants (Dinev et al., 2016), which may in turn lead to less use of voice assistants. Hence, it is possible that self-disclosure to voice assistants may increase users’ concerns over their privacy, which, in turn, may reduce subsequent self-disclosure.
Interestingly, research also has found that privacy concerns do not necessarily reduce or affect self-disclosure. Researchers have long noticed that although people are aware of the data security issues and concerned about their privacy, they would continue using technologies and self-disclosing to them (Zeng et al., 2020). This mismatch between users’ perceptions about privacy and actual behavior is popularly known as the “privacy paradox” (Kokolakis, 2017). Recent research on AI showed this “privacy paradox” among voice assistant users. For example, Vimalkumar et al. (2021) did not find a significant relationship between perceived privacy concerns and behavioral intention to adopt AI voice assistants. Hence, it is also possible that privacy concerns may not reduce self-disclosure, contradicting Altman et al.’s (1981) prediction that the social de-penetration came as a result of privacy concerns. Given the competing evidence regarding privacy concerns, we ask as follows:
Research Question 1 (RQ1). Will self-disclosure to AI subsequently increase levels of privacy concerns, which in turn are associated with decreased self-disclosure?
Method
The study was approved by the Institutional Review Board prior to the start of data collection. Data for this study came from a two-wave panel survey of online panel members recruited by Prolific, a professional survey platform designed for academic and marketing researchers (Palan and Schitter, 2018). The panel members were a national sample of US residents aged 18 or older who reported use of smart assistants in the prerequisite questionnaire collected by the survey company. Participants were paid US$5 for completing the surveys, and they provided consent online. The first wave of data collection took place in the last week of July 2021 with a total of 527 participants. Wave 2 data collection was in the last week of August 2021, with 367 returning participants (69.64% retention rate). Among these participants, 13 failed the attention check questions and were excluded from the final analysis, resulting in a total of 354 participants in the data set. Supplemental Document Table 1 provides the demographics of the sample.
Measures
All questions included in the survey were randomized to minimize order effects.
AI voice assistants
We began the survey by asking respondents to report the AI voice assistants they had used most often over the past 30 days from a list of 18 AI voice assistants. 1 The most frequently chosen choices were Alexa (Amazon Echo)—20.34%, Siri (iPhone)—19.21%, Alexa (Amazon Echo Dot)—18.08%, Google Home—14.69%, Google Assistant (phone)—14.12%, and the other 13 voice assistants combined accounted for 13.56%.
Functional use of AI voice assistant
We developed a set of 10 items, based on the existing literature (De Graaf et al., 2018; Kory-Westlund, 2019), to measure the functional use of voice assistants on a five-point scale (1 = strongly disagree, . . . 5 = strongly agree). An example item is, “my voice assistant improves my work/life performance.” Supplemental Document Table 2 provides the complete sets of questions for this scale. Scores of the final seven selected items, based on the results of factor analysis, were averaged to create an index (T1: M = 4.02, SD = 0.71, α = .87; T2: M = 4.04, SD = 0.71, α = .89).
Relational use of AI voice assistants
A set of 12 items, based on the previous literature (Birnbaum et al., 2016; Ho et al., 2018), were developed to measure the relational use of voice assistants on the same 5-point scale of functional use. An example item is “My voice assistant keeps me company when I am alone.” Supplemental Table 2 provides the full set of questions. Scores of the final 11 selected items, based on the results of factor analysis, were averaged to create an index (T1: M = 2.29, SD = 0.88, α = .93; T2: M = 2.33, SD = 0.99, α = .94).
Self-disclosure to AI voice assistant
This measure was adopted from the self-disclosure to human scale (Gibbs et al., 2006; Wheeless and Grotz, 1976), specifically, the amount of self-disclosure subscale. It was measured by asking participants the behavioral frequency regarding four statements: (1) discuss your feelings with voice assistants, (2) talk about yourself for fairly long periods at a time with voice assistants, (3) talk about yourself with voice assistants, and (4) express your personal belief and opinions with voice assistants, on a five-point scale (1 = never, . . . 5 = always). Scores were averaged to create an index (T1: M = 1.32, SD = 0.74, α = .95; T2: M = 1.33, SD = 0.72, α = .95).
Privacy concern
We measured participants’ degrees of privacy concern regarding voice assistants at Wave 2, because this question may sensitize the participants at Time 1 and influence their pattern of usage between Time 1 and Time 2, especially those who initially did not think about privacy issues. A one-item scale of privacy concern from Manikonda et al. (2018) was used to measure perceived privacy concerns regarding voice assistants: “How much are you concerned about the security of personal information while interacting with your voice assistants?” on a five-point scale from 1 (not at all) to 5 (very much) (M = 2.80, SD = 1.23).
Analytic approach
We used factor analysis to confirm the new scales of functional use and relational use, with the two-wave data set (T1 = 354 and T2 = 354) as a cross-validation strategy (Bandalos and Finney, 2018). We first conducted item analysis and exploratory factor analysis (EFA) using the full set of functional use and the relational use items on the first wave sample. Maximum likelihood factoring with an Oblimin rotation was used for examining item loadings. Based on the results of EFA, we performed confirmatory factor analysis (CFA) on the second wave data. If the model fit was not satisfactory for the measurement model, we examined the medication indices and correlated errors for item correlation within the same factor.
To test the hypotheses and research questions proposed, we performed a cross-lagged panel model using the structural equation model (SEM) (Cole and Maxwell, 2003; Little, 2013). The cross-lagged panel model is robust to reverse causation and causal effects. In the cross-lagged panel model, we estimated the effects of lagged predictor variables (Xt-1) on the outcome variables (Yt) to make causal inference, while controlling for the effects of the lagged outcome variable (Yt-1). The inclusion of the lagged outcome variable allowed us to make causal inferences based on Granger causality, which proposes that the value of a predictor should improve the prediction of the current value of the dependent variable in controlling for the effect of the past dependent variable (Granger, 1980). For the mediation relationships proposed, we conducted a longitudinal mediation analysis with bootstrapping simulations (Cole and Maxwell, 2003). Longitudinal mediation analysis with two-wave data relies on the stationarity assumption that the estimate of the relationship between the mediator and the outcome variable would hold true if more data were collected at a subsequent time point. Despite this caveat, this approach has been widely used in various disciplines (Cole and Maxwell, 2003).
Results
Factor analysis
As noted above, we started with EFA with the first wave of data (n1 = 354). We first examined the correlations of all the items (see Supplemental Documents Table 3) as well as their distributions. We found that one item within the functional use scale, “FU9: My voice assistant helps monitor my health status, e.g., heartbeat, blood pressure,” was weakly correlated with all other items and its distribution was skewed to the right. Upon checking the frequency distribution of this item response, we found that most participants in the sample selected strongly disagree on this item, meaning most of them did not use the health monitoring function of voice assistants. We suspected that this item might lead to a third factor in factor analysis.
Results of eigenvalues and the parallel analysis confirmed our speculation: they suggested the retention of three factors, which is different from the original proposal of two factors (functional use and relational use). Upon reviewing the item loadings, we found that FU9 discussed above heavily loaded on the third factor. Based on the theoretical consideration of the two-factor scale we originally proposed, we decided to stick with the two-factor scale, despite the presence of one problematic item. However, we still included the problematic item FU9 in the two-factor EFA results for the sake of transparency, so that readers can see the descriptive statistics and factor loadings of FU9 (see Supplemental Documents Tables 3 and 4).
Next, to determine which items to retain for each factor, we first looked for items with a strong loading on one factor and then considered cross-loadings. In addition to FU9, two items were removed from the functional use scale due to weak loadings of less than .4 (Boateng et al., 2018; Raykov and Marcoulides, 2011:): “FU8: My voice assistant provides entertainment, e.g., playing music or podcast, etc.” and “FU10: My voice assistant helps me connect with others, e.g., making phone calls or sending messages.” Another item of the relational use scale, “RU7: I would feel a sense of loss if I could no longer use my voice assistant,” was also removed because of its cross-loading (.45 and .26) on both factors.
With the items selected for each factor, we performed CFA on the data set from the second-wave survey (n2 = 354). The initial model fit was: χ²(134) = 866.75, root mean square error of approximation (RMSEA) = 0.124, comparative fit index (CFI) = 0.849, Tucker–Lewis Index (TLI) = 0.828. Upon inspecting the modification indices, we allowed correlated errors: FU3 and FU6, RU1 and RU2, RU4 and RU8, RU5 and RU6, RU8 and RU11, RU9 and RU10. Doing so significantly improved the fit of the measurement model: χ²(128) = 345.73, RMSEA = 0.069, CFI = 0.955 TLI = 0.946 (Table 1).
CFA factor loadings of the function use and relational use scales, nT2 = 354.
RMSEA: root mean squared error approximation; CFI: comparative fit index; TLI: Tucker–Lewis index. CFA was conducted using the data from the second wave of survey to cross validate the EFA results from the first wave. Based on the modification indices, we allowed six correlation paths between items: FU3 and FU6, RU1 and RU2, RU4 and RU8, RU5 and RU6, RU8 and RU11, RU9 and RU10. Model fit without correlation between items is: χ2 = 866.75 (134), RMSEA = 0.124, CFI = 0.849, TLI = 0.828.
Coefficients in predicting functional use, relational use, self-disclosure, and privacy concerns.
SE: standard error. i refers to each individual, t1 refers to data from the first wave survey, t2 refers to data from the second wave survey. Coefficients were unstandardized.
p < .05, **p < .01, ***p < .001.
Hypotheses testing results
A correlation matrix of variables is presented in Table 3.
Zero-order correlation matrix of variables.
t1 refers to data from the first wave survey, t2 refers to data from the second wave survey. Values of correlation coefficients > .10 are statistically significant, p < .05.
A cross-lagged panel model was fitted using the maximum likelihood estimator (MLE) to test the hypotheses and research questions. The model fit was moderate (χ²(7) = 77.32, p < .001; RMSEA = 0.169; CFI = 0.932; SRMR = 0.036). Based on the modification index, we added the correlated errors between relational use t2 and functional use t2, and between self-disclosure t2 and relational use t2. The fit statistics showed excellent fit of the second model to the observed data covariance matrix (χ²(5) = 6.21, p = .29; RMSEA = 0.026; CFI = 0.999; SRMR = 0.017). All the standardized coefficients in the hypothesized models are presented in Figure 1 and Table 2.

Results of the cross-lagged panel model.
H1 proposed the reinforcing pattern of functional use and relational use, respectively. Results supported H1. Functional use of AI voice assistants reinforced itself over time (b = .68, p < .001, 95% confidence interval [CI] = [.60, .76]), and a similar reinforcement pattern was observed for the relational use (b = .76, p < .001, 95% CI = [.66, .85]).
H2 proposed the relational use of voice assistants should reinforce subsequent functional use. Results supported H2 (b = .08, p < .05, CI = [.002, .16]).
H3 proposed the a positive relationship between functional use on subsequent relational use of AI voice assistants via self-disclosure. Results did not support H3. Specifically, functional interaction did not have a significant impact on subsequent self-disclosure (b = 1.83, p = .07, 95% CI = [−.004, .12]). Longitudinal mediation analysis results with 2000 bootstrapping simulations showed a negative, significant mediation effect of self-disclosure (z = −2.21, p = .03, 95% CI = [−.073, −.008]). Functional use had a significant, negative impact on self-disclosure (b = −.09, p = .02, 95% CI = [−.18, −.01]), but self-disclosure significantly and positively influenced relational use (b = .23, p < .001, 95% CI = [.13, .34]).
H4 suggested that the relational use of AI voice assistants should reinforce subsequent relational use via self-disclosure. Results supported H4. Specifically, Longitudinal mediation analysis results with 2000 bootstrapping simulations showed a significant, positive mediation effect of self-disclosure in reinforcing relational use (z = 3.64, p < .001, 95% CI = [.034, .106]). Relational use positively influenced self-disclosure (b = .17, p < .001, 95% CI = [.09, .26]), and self-disclosure also positively influenced relational use (b = .23, p < .001, 95% CI = [.13, .34]).
RQ1 asked whether self-disclosure to voice assistants would increase subsequent privacy concern, which would be associated with a decrease in self-disclosure. Results showed that self-disclosure to voice assistants indeed increased subsequent privacy concern (b = .33, p < .001, 95% CI = [.16, .50]), but privacy concern did not reduce self-disclosure at time 2; instead, it was positively associated with self-disclosure at time 2 (b = .07, p < .001, 95% CI = [.03, .11]). Longitudinal mediation analysis results with 2000 bootstrapping simulation showed a significant, positive mediation effect of privacy concern (z = 2.59, p = .01, 95% CI = [.008, .044])
Discussion
This study set out to integrate two lines of research: technologies as tools and technologies as social beings, under the theoretical framework of dynamic systems (Wang et al., 2011), to investigate the reciprocal relationships between functional use and relational use of AI voice assistants, as well as the mediation roles of self-disclosure and privacy concerns in this dynamic relationship. A two-wave longitudinal survey was conducted among 354 AI voice assistant users across 2 months. Factor analysis results supported our conceptualization and operationalization of functional use and relational use of AI devices as two distinct factors. Results from the cross-lagged panel model confirmed that functional use and relational use reinforced themselves over time. As predicted, relational use also increased subsequent functional use, and more relational use of AI voice assistants led to more self-disclosure, which in turn increased subsequent relational use. Surprisingly, functional use did not have a direct impact on subsequent relational use of AI voice assistants, and longitudinal mediation analysis found the lack of self-disclosure as the mechanism. Furthermore, while self-disclosure increased subsequent privacy concerns, privacy concerns did not reduce subsequent self-disclosure.
In this study, we distinguished the functional use and relational use of AI voice assistants, based on two lines of established research on human-technology interaction. Despite the abundance of research on human-chatbot interaction (Croes and Antheunis, 2021; Guzman and Lewis, 2020; Lee et al., 2020), few studies have distinguished the functional use and relational use of voice assistants or of other interactive technologies. The current study is probably among the first to put forth the conceptualization and operationalization for these two types of usage patterns, and the scales developed in this study lay out an important ontological classification.
Furthermore, we assume that these two usage patterns coexist, and mutually influence each other over time. Thus, we tested the dynamic reciprocity between functional use and relational use, building on the dynamic system perspective (Wang et al., 2011) and social penetration theory (Altman and Taylor, 1973). Results from this study endorsed the positive effects of relational use on subsequent functional use, and the reinforcement pattern of relational use of voice assistants via self-disclosure (Altman et al., 1981; Taylor and Altman, 1987), but did not support the proposition that more functional use would lead to greater relational use through self-disclosure (Cho et al., 2019; De Graaf et al., 2018). On the contrary, results indicated that functional use even reduced subsequent self-disclosure. This may be due to the nonsocial nature of functional use. During functional interactions, AI devices are merely used as tools and users are less likely to share information about themselves with AI. The lack of self-disclosure in this process decreased the likelihood of relationship development.
Now we put together two sides of the story. The significant direct influence of relational use on subsequent functional use, along with the insignificant direct path from functional use to relational use, reveals asymmetrical reciprocity between these two usage patterns. When users use voice assistants as social companions (relational use), they are more likely to use voice assistants as tools (functional use), but functional use did not necessarily evolve into relational use over time. Then what may be the antecedents of relational use of AI voice assistants? Future research should identify the boundary conditions of relational use, such as personality traits or individual differences that may result in the relational development with AI voice assistants.
Results from the current study also confirmed self-disclosure as a critical mediator in the dynamic transactional processes of relational usage. The importance of self-disclosure in relationship development is not only demonstrated in human-to-human interaction, but also in human–AI interaction, which resonates with findings from other studies on how people interact with chatbots (Ho et al., 2018; Skjuve et al., 2021). With more self-disclosure, users feel closer to and develop greater relational bonds with their voice assistants over time. In the current study, we measured users’ self-disclosure only but not voice assistants’ disclosure. As reciprocity of self-disclosure is the key catalyst for human-to-human relationship formation (Fox and Gambino, 2017), future studies should measure such disclosure as a dyadic process between human and voice assistants to provide a more nuanced understanding of self-disclosure in human–AI interaction.
Findings on the mediating role of privacy concern resonate with the accumulated evidence of the “privacy paradox” phenomenon (Kokolakis, 2017). That is, self-disclosure increased subsequent privacy concerns; privacy concerns, however, did not keep people from further self-disclosure. Such a paradox can be explained by cognitive biases and heuristics (Kokolakis, 2017). For example, people tend to believe that they are less at risk of experiencing a negative event, for example, a privacy breach, compared with others. Or, people may make decisions quickly based on their affective impression (Slovic et al., 2007), and positive emotions are associated with a higher level of intention for self-disclosure (Kokolakis, 2017). Therefore, users may underestimate the risks of information disclosure when confronted with their voice assistants that bring enjoyment and convenience to their everyday lives. Understanding the privacy paradox in human–technology interaction can provide useful perspectives on the legal and ethical framework of information privacy, as well as designing privacy awareness campaigns. Future research should continue investigating the psychological mechanisms of privacy concerns, and testing when and why people are still willing to continue disclosing to voice assistants despite privacy concerns (Chalhoub and Flechais, 2020; Liao et al., 2019).
Implications
This study offers important theoretical and practical implications. It offered a useful typology and operationalization of functional and relational uses of AI voice assistants and a longitudinal examination of the mechanisms underlying their dynamic relationship. In doing so, this study integrated theoretical perspectives and empirical findings in multiple lines of research including human–AI interaction, human–machine interaction, dynamic systems, and interpersonal communication. Overall, the findings contribute to existing research by illuminating the reciprocal impact of functional and relational use on themselves and on each other and the underlying relational processes through self-disclosure.
The findings add to the literature by showing that people use voice assistants as both tools and social beings. Scholars have studied the social settings related to the use of domestic technologies, for example, whether smart voice assistant is for personal use or communal use (Kraemer et al., 2019). This opens the question of whether and how the social settings of voice assistants may influence self-disclosure and privacy concerns in the dynamic relationships between functional and relational use. Future studies should include variables such as the frequency of personal use and the frequency of communal use of voice assistants, to provide a more nuanced understanding of the different usage patterns of and self-disclosure to AI voice assistants. Such understanding may inform the design of future AI voice assistants and other AI interactive technologies for better human–AI interactions.
The findings also raise further research questions regarding the dynamic relationship between voice assistant use and users’ psychological wellbeing. Previous studies showed that the relational use of voice assistants may generate psychological benefits (e.g. Ho et al., 2018). Would these psychological benefits further motivate people to use AI as companions? Are the psychological benefits from human–AI interaction comparable with the emotional support received from human–human interaction, and would the relational use of AI increase or decrease interactions with, and self-disclosure to, other humans? These are important questions to understand the socioemotional impact of human–AI interaction. Addressing those questions will provide practical guidance on how human–AI interactions can be utilized for improving users’ mental wellbeing and life satisfaction.
Limitations and future directions
This study has some limitations. First, this study examined the dynamics of human–AI interaction with a two-wave survey over 2 months. Future research may adopt a more intense data collection method, for example, experience sampling, to capture more nuances in people’s use of voice assistants at multiple times in daily life; or replicate this study with multiple-wave surveys over a longer period of time, to test the relationship identified in the current study. Second, the self-reported measures used in this study for the use of AI voice assistants may not fully capture how people use voice assistants in their daily lives. Future research will benefit from utilizing more objective measures: for example, naturally occurring data such as the actual conversations between users and AI devices to capture not only the frequency of use but also the content of the interactions. Third, only one item was used for measuring privacy concerns in this study. Although that measure was validated by previous studies, scholars have argued that privacy concerns are a multidimensional construct (Hong et al., 2021). A more comprehensive measure may be used in the future for a more detailed investigation of privacy concerns’ role. Moreover, we only measured privacy concerns once to avoid sensitizing participants to the privacy issues. Future studies can adopt a more implicit measure of privacy concerns and measure how privacy concerns change over time. Furthermore, our sample may not necessarily be representative of the population of voice assistant users in the United States, and future studies should use a representative sample of chatbot users to replicate this study.
Supplemental Material
sj-pdf-1-nms-10.1177_14614448221108112 – Supplemental material for A tool or a social being? A dynamic longitudinal investigation of functional use and relational use of AI voice assistants
Supplemental material, sj-pdf-1-nms-10.1177_14614448221108112 for A tool or a social being? A dynamic longitudinal investigation of functional use and relational use of AI voice assistants by Shan Xu and Wenbo Li in New Media & Society
Footnotes
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
Notes
Author biographies
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
