The ludic aesthetics of AR microgames in online self-communication

When one thinks about ARFs, the first thing that comes to mind is “beautification” digital effects: devices that, through image-processing techniques, can modify elements of the facial image, making the skin appear smoother, the nose smaller, the lips fuller, the eyes larger, etc. Many makeup brands have implemented this type of function within their virtual tryvertising software (which, in addition to beautifying the face, adds commercial makeup products). Many other applications, not necessarily connected to brands, such as FaceApp or Youmake-up, have followed the same trend, becoming indispensable tools for influencers in the second twenty-year period of the 2000s. However, as we will see, the use of AR for the beautification of selfies constitutes only one possible kind of ARFs. Several studies have shown that beautification is neither the only reason nor the prevailing one in the use of ARFs, and that aesthetic motivations are flanked by entertainment, “coolness,” curiosity, social interaction, silliness, fun, creativity, brand “fandoms,” and so on. Several research, for example, tacked with the ARF’s perceived curiosity and compatibility that influenced user satisfaction, contributing to detail the hedonic, utilitarian, social and personal reason that drove and derive when interacting with AR filters on social media (Ibáñez-Sánchez, Orús, and Flavián 2022).

It is possible to identify a plurality of visual rhetorics that go well beyond beautification. Among them:

rhetorics of distortion, in which filters deform the face or specific facial attributes (here, in its “purest” form, closest to the logic of the photographic filter) (Fig. 1A);

Fig. 1A - rhetorics of distortion

rhetorics of inscription, in which filters anchor tattoos, masks, or other objects onto the face (Fig. 1B);

Fig. 1B - rhetorics of inscription

rhetorics of modularization, in which body parts are duplicated and displaced (Fig.1C);

Fig. 1C - rhetorics of modularization

rhetorics of framing, in which the filter places the face within a frame or simulates a multi-stage recording process (Fig.1D);

Fig. 1D - rhetorics of framing

rhetorics of immersion, in which the filter inserts backgrounds, glitter, or concave three-dimensional objects, producing the effect of the face being immersed in a virtual environment (Fig.1E);

Fig. 1E - rhetorics of immersion

rhetorics of fandom, in which the filter inscribes attributes drawn from a narrative universe — whether real or fictional — onto the image (Fig.1F);

Fig. 1F - rhetorics of fandom

rhetorics of interrogation and quiz, in which the filter stages a random quiz whose outcome produces expressive consequences on the user’s face (Fig.1G);

Fig. 1G - rhetorics of interrogation and quiz

“gaming” rhetorics, in which the filter situates the face within a game-like environment where (often, though not necessarily) facial movements correspond to “moves” (Fig.1H);

Fig. 1H - gaming rhetorics

rhetorics of concealment (camouflage), in which the filter inserts elements that partially or fully occlude the face, making it difficult to interpret — precisely so that the filter itself can emerge as such (Fig.1I);

Fig. 1I - rhetorics of concealment

rhetorics of pareidolia, in which the filter isolates a limited set of features and transfers them onto its own texture, producing the visual effect commonly described as pareidolia (Fig.1J);

Fig. 1J - rhetorics of pareidolia

rhetorics of animalization, in which the filter confers animal-like traits on the face, ranging from minor embellishing attributes to fully covering masks (Fig.1K).

Fig. 1K - rhetorics of animalization

A more fine-grained approach, starting from this assumption, can account for how ARFs formally articulate self-representation and how this re-articulation radically differs from the experience of the mirror. ARFs activate a situation of real-time self-observation in which the user regulates their behaviour on the basis of the visual feedback provided by the front-facing screen; yet the rhetorics we have identified show that the ARF does not simply return a reflection, but reworks it, enriches it, diverts it, and inscribes it within a set of technical and discursive rules. Even if we consider only the specific case of gameful ARFs, for example, we can identify occurrences in which the mirror analogy no longer holds: it is true that the user’s face is always “represented” in the image, but it may be distorted in a way that thematically adapts it to the environment represented in the image; replaced by an avatar that remains “internal”; or externalized into a liminal position that is not so much the space of the game as that of the player’s action.

The starting postulate of our analysis, that already supported the taxonomical operation, is that these ARFs may constitute a recognizable form of visual self-communication within pop culture, that they could be approached as a genre, and that they were characterized by recurring motifs that can be traced, for instance, to the cinema of attractions. In Méliès’s films such as The Four Troublesome Heads (Fig. 2A), in fact, we encounter rhetorics of modularization of body parts, while in The Man with the Rubber Hand (Fig. 2B) we find a rhetoric of immersivity and framing.

Fig. 2A Fig. 2B

Importantly, and in a more direct connection with digital images, we can identify analogies with multimodal texts in the tradition of media/body art, such as Petra Cortright’s VVEBCAM (Fig. 2C), a YouTube video in which the artist stares intently into a webcam while cartoonish clip-art figures float around her face.

Fig. 2C

A key figure in media art to have experimented early on with these instruments is also Jeremy Bailey, whose works like The Future of Television (Fig. 2D) involved the creation of a face-tracking app that allowed to “broadcast” different TV programs onto his face: the performance was not only a mimicry of television’s logorrheic language, but also a rather accurate prediction: for Bailey, the future of television would be a matter of sharing — of wearing and displaying what one likes to other people.

Fig. 2D

However, ARFs have been used several times in digital art also, for example in Laurent Mignonneau and Christa Sommerer’s work Apis Humane (Fig. 2E) or Tamiko Thiel (Fig. 2F), who also made extensive use of face-tracking techniques to produce immersive scenes in underwater environments — often in a critical register focused on sea pollution.

Fig. 2E Fig. 2F







Regarding gameful ARFs, a first typology can be developed by drawing on the ludic genres: competitive games (agôn), characterized by a clearly formulated goal that competing players attempt to achieve within a regulated system (7A); games of chance (alea), which do not depend on the player’s abilities but on random processes (7B); games of simulation (mimicry), encompassing anything from children’s imitation games to theatrical performances (7C); and games of the pursuit of vertigo (ilinx), which “momentarily destroy the stability of perception” in order to “inflict a kind of voluptuous panic upon an otherwise lucid mind” (7D).

With regard to videogame aesthetics, we can identify multiple analogies with the visual forms and mechanics of “classic” minigames such as WarioWare (Fig. 3A) — the Nintendo series and Mario spinoff explicitly designed for “flash” sessions — structured as a barrage of microgames lasting only a few seconds (around five), immediately legible and meant to be solved on the fly. Examples include minigames like Dodge Balls, in which the player steers a small car in an arena and dodges bouncing boulders until the microgame ends, or Rock-Paper-Scissors. A related interaction mechanic was later implemented in titles such as Doodle Jump (Fig. 3B), where the player controls a small green creature, making it bounce on platforms in order to climb upward. This game exploited the accelerometer of early smartphones and, in doing so, proposed an interaction grammar different from WarioWare’s: it began to shift from discrete, button-like commands toward continuous analog input, such as the smooth movement of the device in the “real” world. Finally, a further videogame genre remediated not only by ARF but also by VR is the first-person endless runner in a rhythm-action format, well represented by Guitar Hero (Fig. 3C), where the player must interact and produce signals in synchrony with the music. Unlike WarioWare, what changes here is above all the scopic regime: we move from third person (or “god’s-eye”) perspectives to a first-person shot.

Fig. 4B
Fig. 4A
Fig. 4C










In order to analyse non verbal communication in embodied ARF, we draw on Ekman and Friesen’s taxonomy of facial and bodily behaviours that identify three main dimensions of non-verbal behaviour: the usage, the origin, and the coding. The usage refers to the circumstances surrounding a non-verbal act — for instance, its relation to verbal behaviour; the subject’s awareness and intentionality; and the type of information conveyed: this dimension sheds light on the decodability of embodied ARF for other users and, more broadly, on the interpersonal dimension of communication. Origin, by contrast, concerns how a given non-verbal behaviour became part of a person’s repertoire — i.e., the source of the action. This dimension offers insight into the virality of forms in ARF (for example, in challenges), as well as into influencers’ repertoires and culturally sedimented stereotypes. Finally, coding concerns the correspondence between an act and its meaning and therefore bears directly on the decoding of ARF’s multimodal languages. The authors primarily distinguish between arbitrarily coded acts and iconically coded acts. An example of the former is a raised hand signifying greeting or departure; an example of the latter is a gesture that imitates an action without being the action itself (for instance, when a person waves a fist menacingly). They also distinguish this kind of gesture — extrinsically iconic — from the act itself, which is likewise a non-verbal act but no longer merely iconic: it is intrinsically the very act that the gesture “imitates.” We can argue that embodied ARF involve both arbitrarily coded acts and iconically coded acts, but only in some cases are these behaviours configured as system-coded actions — that is, as signifiers that produce meaning also at the level of extended reality.

This first form of semiosis, for instance, may be configured through a pictorial relationship, where meaning is conveyed by “drawing” a picture of an event, object, or person (Video 5A); a spatial relationship, in which movement indicates distance between people, objects, or ideas (with distance functioning as a plastic category of space, Video 5B); a rhythmic relationship, in which movement traces the flow of an idea, accents a particular phrase, or describes the rate of an activity (Video 5C); a kinetic relationship, in which movement executes all or part of an action performance, where that performance either signifies — or partially constitutes — the meaning (Video 5D: in this case, is also pictorial); a pointing relationship, in which some part of the body — usually the fingers or the hand — points to a person, a body part, or an object/place (a very common format, notably, in videos where developers test or present new ARF) (Video 5E).

If this taxonomy is meant to describe the kind of biplanar relation between movement and meaning at the machinic level, and the syntax of the HCI itself, however, it does not allow us to grasp the meaning of other dimensions that are less codified yet still recognizable as signifying within a culture, for example, the meaning of a gesture that is used as an embodied arf, as well as the meanings ascribed to a specific entity during the performance. Yet, we can account for this dimension retrieving further categories proposed by Ekman and Friesen’s taxonomy such as emblems, illustrators, affect displays, regulators, and adaptors: these are designed to describe actions in terms of function/use, degree of intentionality, and relation to verbal language, rather than in terms of a “form of representation.”

Hence, we can recognize emblems in ARFs, that is, conventional signals with a kind of “dictionary translation” defined by a shared and decoded meaning and by conscious and intentional usage within a given group of individuals. Understanding embodied Arf as emblems allow to highlight those already codified gestures that are remediated in ARFs. For example, opening one’s mouth to eat computer graphics candies establishes a motivated, emblematic relation (Video 6A). By contrast, produce a sound with your voice to make an avatar jump is more arbitrary (Video 6B).

Then, we can recognize in ARF Illustrators, movements tied to speech that illustrate what is being said — for instance, pictographic movements that “draw” a picture of their referent; deictic movements, i.e., gestures that indicate an object present in the situation. In this sense, understanding ARFs as illustrators (and since verbal language is often absent) brings to the fore the strategies through which user package and convey meaning about specific entities — for instance, in the case of ARF designers who showcase a filter’s functionalities in a story. In the Video 7A, the users’ pictographic movements correspond to a bodily discourse made of turns of poses — and probably related, at another level, to a collective discourse, such as a challenge — whereas in the Video 7B, the ARF discourse enact a more theatrical, mimetic performance tied to the utterance “I am a boxer.”

Ultimately, ARF-related nonverbal behaviours can be understood as sites of affective expression: emotional cues — most often facial, but also vocal (prosody), postural, gestural, and linked to bodily tension — manifest expressive capacities and make perceptible affective states of varying kinds and intensities. Attending to these cues is crucial for grasping the thymic dimension of embodied affordance, that is, the effects of meaning generated by the assemblage of affective signals and affordances as well as to understand how visual and interactive components of the image are invested with affective meaning during the ARF embodiment.

Often, ARFs induce unnatural movements in order to make the user appear ridiculous, and this is , to play, or to amuse themselves are a typical example: in such cases, the experience is framed ludically. By contrast, this ludic framing seems absent from the 5B video, in which the underlying discourse appears to concern only the functioning of the filter itself. Yet, the impassive expression is by no means insignificant: Bruce Sterling, for instance, writing about Petra Cortright’s WEBCAM project, characterizes her expression as an absent gaze: solemn and contemplative, even as digital augmentations move frenetically around her. He reads this attitude as a gesture of resistance to the imperative of pleasure, one that sets Cortright apart from the “cam girls” of the same period, described instead as performing themselves in an anxious search for approval through viewers’ comments.