Abstract
It is critical to understand how algorithms structure the information people see and how those algorithms support or undermine society’s core values. We offer a normative framework for the assessment of the information curation algorithms that determine much of what people see on the internet. The framework presents two levels of assessment: one for individual-level effects and another for systemic effects. With regard to individual-level effects we discuss whether (a) the information is aligned with the user’s interests, (b) the information is accurate, and (c) the information is so appealing that it is difficult for a person’s self-regulatory resources to ignore (“agency hacking”). At the systemic level we discuss whether (a) there are adverse civic-level effects on a system-level variable, such as political polarization; (b) there are negative distributional or discriminatory effects; and (c) there are anticompetitive effects, with the information providing an advantage to the platform. The objective of this framework is both to inform the direction of future scholarship as well as to offer tools for intervention for policymakers.
Imagine a magic, wish-granting entity; however, unlike the magical avatars from folklore, this being grants your wish before you even make it. Is this entity enhancing or inhibiting your autonomy? Is autonomy about making the wish or having the wish come true? Now, let us modify this scenario a little so that the entity actually has its own agenda and can anticipate your wish and grant it but with some insinuation of its own preferences. At what point is this entity enabling or undermining your agency? In many ways, the algorithms of the modern internet that route information to people are like this magic entity, trying to anticipate your desires, in some cases before you have articulated any preferences. Feed-based platforms, such as Facebook and Twitter (now X), aim to anticipate who you want to hear from and about what topics. Google Search will attempt to answer your questions about the world on the basis of a seemingly inscrutable jumble of words. These modern miracles are driven by algorithms that are built on top of the massive streams of behavioral information that the internet produces. But in what ways do they help individuals and society achieve core values versus undermine them?
Our goal is to offer a classification system that encompasses the primary criteria for assessing internet information curation algorithms, where future work on developing more finely granular criteria is necessary. We split the assessment criteria of algorithms into two broad categories: the individual (user) level and the systemic level. At the individual level, we focus on the question of whether the results enhance or undermine agency: Do the results provide you with what you wanted or would have wanted? As can be seen from Table 1, we identify three types of criteria at the individual level: whether the information aligns with the person’s interests, whether the information is accurate, and whether the information is appealing in such a way that it undermines priorities set by the individual, what we term “agency hacking.” We also provide three criteria at the systemic level: whether there are generally adverse civic effects on some variable, such as political polarization; whether there are more specific negative distributional or discriminatory effects; and whether there are anticompetitive effects. We note that this classification system is not meant to be fully comprehensive; rather, it is induced from the existing literature that critically evaluates the role of algorithms in information curation on the internet.
Framework for Evaluating Internet Algorithms
We note the essential distinction between platform and algorithm. A platform is simply an internet-based service (such as Google or Facebook). All major platforms currently rely on algorithms—rapid, data-driven, automatic decision making—to curate information. Designers make choices when developing and deploying algorithms, and these choices have positive or negative effects. It is also possible that there are intrinsic effects to platform design that are independent of their algorithmic machinery. For example, research suggests adverse effects of Facebook on the mental health of college students, likely driven by processes of social comparison (Braghieri et al., 2022). These social comparisons may be intrinsic to the kinds of networked social processes of connections to peers and the resulting presentation of self in comparison to others. That is, this particular issue may be driven by Facebook’s nonalgorithmic affordances. We now turn to our six criteria for assessing information curation algorithms, starting with our three individual-level criteria and then shifting to our three system-level criteria.
Relevance
We define relevance as the alignment between what one likes or wants to see and what one is shown by an algorithm. Relevance is the core value proposition of algorithms that curate information. Finding a video on how to repair a dishwasher of a particular brand, identifying the location of the nearest decent pizzeria, or generating a list of reviews for a movie are such trivial exercises in 2023, that we lose sight of how extraordinary it is that we can find useful information so quickly. Finding content on the internet in the late 1990s relied on portals such as AltaVista, which were hierarchically organized and human-curated, very slow to update, and noncomprehensive. PageRank (Brin & Page, 1998) was the foundational algorithm of the web, using the hyperlink structure of web pages as a signal of the value of the page, and it offered a scalable approach to rapidly classify much of the content on the internet. The leap from AltaVista to the platform built on PageRank, Google, was epochal, both in terms of user experience and in putting information curation algorithms at the center of the internet more generally. The rise of Google also precipitated the rise of gaming the algorithm: a vast industry of search-engine optimization to raise the ranking of a page in Google Search (Wang et al., 2014).
This algorithm-driven model of connecting people to information of various kinds has many permutations, but the essential formula is always the same: automated, scalable, provision of information, and based on very complex statistical models deriving signals from observable behaviors. This is the principle behind a vast array of information curation algorithms, such as information retrieval (i.e., search), targeted advertising, recommender systems, and social feeds.
Any discussion of the value proposition of internet algorithms must begin with relevance, because that is the core value of the algorithms that have survived—they provide content that is relevant to what users want. As then-Google CEO Eric Schmidt stated in 2010, “Ultimately, we think we can understand things like what you really meant. . . . [w]hat is the problem you’re really trying to solve?” (McGee, 2010, para. 6). However, relevance is not the only value—either for individuals or society—at stake. We now turn to accuracy.
Accuracy
Accuracy—defined as the veracity of information—is also important to the individual and to society. If, for example, a user searches for and finds the closest pizzeria but the hours open are incorrectly reported, this information could adversely affect that user. This type of failure could result from an algorithmic failure (e.g., incorrect extraction of information) or from errors in the underlying posted information. However, the patterns of behavior from which algorithms are built on—most notably engagement—are often not related to accuracy. Consider, for example, the controversy that Google confronted in 2016, when its search engine, posed with the query “Did the Holocaust happen?” pointed to a Holocaust-denial site as its first result (Ortiz, 2016). The likely explanation for this is that the individuals who engaged with this question most were Holocaust deniers. More generally, “data voids”—topics on the internet with little high-quality information—are highly vulnerable to this kind of algorithmic failure (Golebiewski & boyd, 2019). This is particularly the case when elite actors intentionally and strategically populate data voids as a form of search engine optimization. For example, political elites can drive people to partisan content by deploying novel phrases that are used only by ideologically aligned online media sources, because those are the only places on the internet to find those exact sequences of words (Tripodi, 2022). Beyond data voids, there is the issue of a platform’s interest in providing accurate information. Relevance of content will (in part) drive engagement and thus platform profits. By contrast, accuracy only sometimes drives engagement. For example, if Google Search regularly provided inaccurate information regarding pizzeria locations and hours, people—upset by their inability to get pizza—would likely stop using Google Search for that purpose. However, this type of feedback does not always exist, and these are the corners of the internet where misinformation can potentially multiply.
E-commerce is one example in which misinformation can thrive, given that it is not in the interest of sellers to undermine their goods. For example, although e-commerce giants algorithmically curate the comments sections of goods to remove fraudulent comments, those comments sections are vulnerable to coordinated efforts to facilitate sales. If one looks up “apricot kernels cancer” in Amazon, it will show the user specific apricot kernel goods, even though apricot kernels are ineffective at curing cancer and may cause cyanide poisoning. In turn, Amazon presents shoppers with testimonials in the comments claiming how apricot kernels have cured people’s advanced cancer (Swire-Thompson & Lazer, 2020). Other facets of e-commerce platforms are also vulnerable to health misinformation: Juneja and Mitra (2021) found that products containing health misinformation were ranked highly in Amazon product search results and that engaging with these products caused Amazon to recommend more of them. Although false claims about the health benefits of products by sellers are regulated by statute, there is no regulation of false claims in the comments section or the algorithmic curation of comments or products in general. Amazon’s interest may be to make a sale rather than to provide accurate information about goods.
Agency Hacking
We define agency hacking to be when algorithms provide information that is so appealing to the individual that it becomes difficult for the individual’s self-regulatory resources to ignore. The metaphor of a horse and rider is often used to illustrate cognitive control, in which the rider (cognitive control) must constrain and regulate the horse (impulsive or automatic processes; Friese et al., 2011; Wiers et al., 2013). In other words, cognitive control is necessary to suppress impulses and execute goal-directed behavior (Menon & D’Esposito, 2022). Although this paradigm is often used in regard to addiction (Franken et al., 2017; Groman & Jentsch, 2012) or reward (Frömer et al., 2021; Otto & Vassena, 2021), it can also apply to behavior online. The information may be engaging and appealing to automatic processes but may be counter to top-down executive control goals or priorities set by the individual. People therefore require strong control processes to inhibit the desire to engage with the extremely appealing content algorithms provide. As such, individuals that are tired, distracted, or have limited inhibitory control would engage with the content presented (Snippe et al., 2019).
Current social media algorithms’ primary goal is to predict what the individual finds interesting and keep them engaged (Lewandowsky & Pomerantsev, 2022). Although television, books, and radio also have the goal of keeping their audience engaged, online algorithms can present uniquely curated content according to each person’s interests. Individuals may not have sufficient inhibitory control to resist clicking on personalized clickbait (Wegmann et al., 2020) when they wish to be engaging with other content online or reduce the time that they spend online altogether. In 2018, 45% of teens reported that they were online “almost constantly” (Anderson & Jiang, 2018). Furthermore, overuse of social media is generally associated with low work performance (Zivnuska et al., 2019), loneliness (Berryman et al., 2018), sleep problems (Koc & Gulyagci, 2013; Wolniczak et al., 2013), anxiety, and depression (Pantic, 2014), although causal direction is very difficult to parse.
There are clear examples of personalized content that is agency-enhancing (such as finding a local restaurant) and agency-undermining (such as presenting eating-disorder content to a recovering anorexic; Harriger et al., 2022). Indeed, people often do not want algorithmic personalization in certain domains. Kozyreva et al. (2021) surveyed people in the United States, Germany, and Great Britain about their attitudes toward algorithmic personalization. They found that people accepted personalized commercial services such as shopping and entertainment but objected to the personalization of political campaigning (in all countries) and news sources (in Germany and Great Britain).
Civic Effects
Although individual agency and autonomy are values of great importance, focusing solely on individuals may miss the forest for the trees. Algorithmic systems shape the experiences of billions and thus have the potential to create harmful effects at the “civic” or population level. An example of harmful civic effects are concerns that algorithmic systems may subvert public trust in institutions, such as democratic governments, public health agencies, mainstream journalism, and academia. Some have argued that public trust may be subverted because of personalization algorithms that create polarizing filter bubbles (Pariser, 2012) or recommender systems that radicalize individuals (Tufekci, 2018). Journalistic accounts also suggest that Facebook played a key role in mobilizing genocidal acts against the Rohingya in Myanmar, although the exact role that curation algorithms played compared with other affordances is unclear. At a minimum, the lack of algorithms to block hate speech was important; at most, it is plausible that prioritization on engagement increased the dissemination of hate speech (Mozur, 2018).
Current scholarly evidence for harmful civic effects of algorithms is mixed. Early research on political filter bubbles on Facebook, for example, suggested that there are significant but modest algorithm-driven filter bubble effects in newsfeeds (Bakshy et al., 2015). More generally, research on browsing highlighted that opposing partisans engage with fairly similar content online (Guess, 2021; see also Gentzkow & Shapiro, 2011). However, more recent research on the use of Facebook during the 2020 election (1) suggests a large segregation of news consumption on Facebook (González-Bailón et al., 2023), although disentangling algorithmic drivers from other affordances of Facebook is nearly impossible; but (2) that changing the algorithm for three months did not have significant effects on attitudes (Guess et al., 2023; Nyhan et al., 2023). Furthermore, although there was some preliminary evidence for algorithms sending users down the “rabbit hole” on YouTube (Alfano et al., 2021), other studies have been unable to replicate these effects (Bisbee et al., 2022; A. Y. Chen et al., 2021; Hosseinmardi et al., 2021; Ledwich & Zaitsev, 2020; Ribeiro et al., 2020). One alternative possibility is “filterless bubbles” in searches, in which partisans are shown similar options algorithmically but they themselves choose to engage with ideologically aligned content (Robertson et al., 2023). That is, ironically, providing diverse but relevant content may result in heavily partisan sorted news consumption, with resulting political polarization. Beyond partisanship, Yoon et al. (2022) found that once an individual watched one YouTube video that contained misinformation about a cancer remedy (fenbendazole, a deworming medication for pets), they were recommended other misinformation videos on the same topic. These results highlight that there is likely a fundamental tension between the individual interest in relevance and some of the civic effects that we have identified.
Distributional and Discriminatory Effects
A special case of systemic effects occurs when algorithms cause specific (often minoritized) groups to suffer harms. Examples of discriminatory effects of algorithms abound. Sweeney (2013) infamously demonstrated that Google Search would show innocuous advertisements when users searched for White-sounding names but racially charged ads for criminal background checks when searching for Black-sounding names, thus implying that the person in question was a criminal. Researchers found that Facebook’s advertising systems were discriminatory in two ways: Advertisers could explicitly choose targeting parameters that were discriminatory (Angwin et al., 2017; Angwin & Parris, 2016; Keegan, 2021; Sapiezynski et al., 2022; Venkatadri & Mislove, 2020; Waller & Lecher, 2022), or Facebook’s machine learning algorithms would “optimize” ad placements in a discriminatory manner even if advertisers did not select discriminatory targeting parameters (Ali et al., 2019, 2021; Kaplan et al., 2022). In both cases, individuals from marginalized (often legally protected) groups were prevented from seeing opportunity-related ads, for example, for housing, credit cards, and jobs.
Finally, several researchers have documented cases in which search algorithms produced results that perpetuate harmful stereotypes (L. Chen et al., 2018; Edelman & Luca, 2014; Hannak et al., 2017; Kay et al., 2015; Metaxa et al., 2021; Noble, 2018). The harms of these distributive effects are twofold. First, there are direct effects on affected individuals, such as loss of opportunities (Rajkumar et al., 2022). Second, there are indirect effects of propagating harmful stereotypes to the population as a whole. These two effects may be cumulative, reifying historical inequalities.
Anticompetitive Effects
One of the first algorithmic systems that had widespread impact on the general public was SABRE, an airline ticket reservation system developed by American Airlines in the 1960s that was quickly adopted by most of the major airlines of the day (Luo et al., 2015). Much later, it was discovered that American Airlines was manipulating the order of flights that appeared in search results to favor their own flights and thus reallocate revenue from their competitors to themselves (Cusumano et al., 2021). This tale perfectly exemplifies the potential for algorithms to have anticompetitive effects. The combination of systemically important platforms controlled by commercial interests, coupled with opacity surrounding these systems, creates fertile ground for systems to be designed in ways that favor their owners at the expense of competitors.
Amazon has been credibly accused of designing the algorithms on their platform to further their own business interests, for example, by self-preferencing their own products in search results and in the Buy Box (L. Chen et al., 2016; Jeffries & Yin, 2021). Despite this, merchants have little choice but to list their inventory on Amazon anyway to gain access to their enormous customer base. Google has also been accused of anticompetitive practices with regard to Google Search, such as self-preferencing their own vertical search services (e.g., for hotels, flights, and comparison shopping) over competitors in search results (Jeffries & Yin, 2020). Google argues that its services rise to the top of search results “organically” because users prefer them, but evidence suggests that this is not true (Subcommittee on Antitrust, Commercial and Administrative Law of the Committee on the Judiciary, 2020). This example highlights the potentially anticompetitive consequences of information dominance: Companies are heavily reliant on Google Search for referral traffic, yet those same data give Google a window into consumer preferences that it can use to identify competitors, develop competing services, and ultimately divert users toward those services (Competition & Markets Authority, 2020; Subcommittee on Antitrust, Commercial and Administrative Law of the Committee on the Judiciary, 2020).
Academics have been theorizing about the anticompetitive effects of algorithms (Khan, 2019; Srinivasan, 2018), and those concerns are starting to be heeded by lawmakers. In Europe, for example, the Digital Services Act prohibits a range of anticompetitive behaviors by dominant platforms, in line with the critical view that these platforms are akin to natural monopolies and should be regulated as such (boyd, 2010; Ghosh, 2019; Newman, 2011).
Conclusion
We have offered a six part normative framework for assessing a small but crucial slice of the algorithms that govern the 21st century. These algorithms are the essential “middleware” for democracy and markets that connect people to information on the internet. They determine what you hear about your friends, what you find out about politics and policy, what goods you see, and what prices you get. In many ways, they can be extraordinarily enabling of individuals—magical from a 20th-century point of view. In other ways, it is frightening how much influence these algorithms have on the shape of modern life relative to the opacity with which they operate and the conflicts inherent in their platforms’ position as profit-driven corporations.
We proposed the categories of relevance, accuracy, agency hacking, civic effects, distributional and discriminatory impacts, and anticompetitive effects. It is important to note that there will be trade-offs among the criteria of our framework. For example, accuracy, at times, may be in tension with relevance. People may not be always seeking the epistemic consensus of what is true. If you search for “Did the Holocaust happen?” you may really be looking for evidence that the Holocaust did not occur, perhaps for innocuous reasons such as a school project on Holocaust denialism. However, we argue that the platforms have often overprioritized relevance to the detriment of other values because relevance is most directly related to engagement and platform profits. Interventions that prioritize accuracy more than the status quo may often be desirable. Kington et al. (2021), for example, offered a framework for prioritizing accuracy as a value for health information; YouTube (n.d.) and Google (n.d.) have stated that they have adopted these principles into their algorithms. Such prioritization of authoritative sources may come at some cost of relevance (e.g., if someone is looking for antivaccine information) and engagement (as people may migrate to other platforms). However, in a domain such as health, there is a compelling argument that accuracy should be prioritized (see Swire-Thompson & Lazer, 2022).
This framework, and future elaborations, could be integrated into training for the professionals who are—consciously or not—embedding these values into platform design and supporting policies. Although professional ethics in the computer and data sciences are in their infancy, integrating ethics education into technical degree programs may yield an engineering workforce that is more engaged with the normative ramifications of their work. Frameworks such as “value sensitive design” that teach designers to identify stakeholders, surface their values, and build systems that embody these values (Friedman & Hendry, 2019) may be useful starting points for helping engineers identify and navigate normative trade-offs (Kopec et al., 2023).
When considering future solutions, regulation will have a crucial role to play, especially because the value prioritization of the internet giants often do not align with society’s interests. The least onerous interventions may be transparency requirements such as mandatory disclosures when algorithms are being used along with the broad contours of their design, the data they utilize, and safe harbors for data scraping and adversarial algorithm auditing. These requirements might help to foster a more informed public and enable platforms’ design choices to be critiqued. A step beyond general transparency would be mandatory pre- or postdeployment algorithmic impact assessments by independent auditors, or requirements for data sharing with researchers. Finally, outright prohibitions against certain types of algorithms (e.g., facial recognition) or platform designs (e.g., self-preferencing) may be needed in cases in which the values at play consistently favor powerful vested interests. The Digital Markets Act (which prevents large companies from abusing their market power) and the proposed Artificial Intelligence Act (which regulates artificial intelligence based on their risk to cause harm) in the European Union already utilize many of these regulatory levers (Larouche & de Streel, 2021; Veale & Zuiderveen Borgesius, 2021).
In each of the categories in our framework, it is remarkable how little we know. The literature contains some insights on each of these questions; however, our accumulated knowledge on the answers to each is sparse, particularly when the speed of changes of the platforms is considered. For example, what we once knew about Twitter algorithms circa early 2022 (which was not very much) is now obsolete and nonreplicable in 2023, after Elon Musk has reshaped the company as “X” and shut down its data access. Even absent the “Musk effect,” all of the major platforms are rapidly changing, in part because algorithms are changing and in part because how people use platforms is also changing. Algorithms are proprietary, and data regarding people’s experiences on platforms are scarce, especially around what people see (Lazer, 2020). Google provides some limited insights, for example, on what people search for (Google Trends) but none on what search results people are shown. Twitter’s application programming interfaces once allowed you to see what people shared but not what they saw. This points to a desperate need for greater transparency into the algorithms of the internet (Pasquetto et al., 2020). Most of the outputs of information algorithms are fleeting—a platform shows you something; you engage or move on. There are no long-standing data gathering efforts to capture the ephemera of the internet—although the internet Archive captures snapshots of static content, there is no archive to capture what Google Search from 2007 provided in response to a particular query. The curation algorithms of the internet continue to evolve. For example, what role will large language models have in shaping the information we see in the coming years? When making design choices for these models, there will be similar trade-offs in human values. We need to collectively make visible those platform decisions, and our effort here is to identify, in turn, the human values at stake in those choices.
