Unsupervised Discovery of Mythemes
Research Aim
This project investigates whether recurrent mythemes can be identified computationally across a corpus of myths and fairy tales without assigning predefined mytheme labels. The central hypothesis is that mythemes are not best understood as recurring characters, objects, actions, or motifs in isolation. A fox, forest, journey, theft, marriage, or death may perform very different functions in different narratives. What persists across stories is better understood as relational configurations: a structured sequence of tensions, transfers, inversions, mediations, and changes of state.
The project will therefore develop an unsupervised or self-supervised method for discovering recurrent relational structures across narratives. Rather than applying a model directly to raw text, it will build on an existing extraction method that identifies key narrative elements such as actor, action, patient, object, location, and outcome. These elements will be contextualised through their relations and transformations, enabling the system to compare episodes whose surface contents differ but whose structural organisation may be similar.
The intended outcome is not an automated declaration that a particular structure is definitively a mytheme. Instead, the system will generate interpretable clusters of recurrent relational configurations that can be evaluated as candidate mythemes through comparative structural analysis.
Conceptual Framework
The proposed approach distinguishes between three levels of narrative representation. The first is the level of elements: actors, objects, locations, and actions. The second is the level of relations: who acts upon whom, what is transferred, what boundary is crossed, and how roles are distributed. The third is the level of transformation: how an event changes possession, knowledge, status, identity, kinship, or categorical position.
An extracted narrative event can initially be represented as:
Ei = (ai, ri, pi, oi, li, si–, si+)
where ai is the actor, ri the action or relation, pi the patient or recipient, oi the relevant object, li the location, and si– and si+ represent the states before and after the event.
The transformational component is then:
ΔEi = si+ – si–
This difference should not be understood as a simple lexical vector subtraction. It represents a structured change, such as ignorance becoming knowledge, exclusion becoming inclusion, possession passing from one actor to another, or a human figure becoming animal. Two events may therefore be compared not because they contain similar entities, but because they enact comparable transformations.
The project will also examine oppositional and contrastive dimensions. However, semantic opposition will not be defined as a vector located exactly 180° from another vector, since standard embedding spaces do not reliably position antonyms as geometric opposites. Instead, relevant tensions will be represented as learned or induced axes, such as inside–outside, human–animal, hidden–revealed, kin–stranger, permitted–forbidden, possession–dispossession, and life–death.
For an opposition axis dk, derived from a set of contrastive examples, an event can be projected onto that axis:

The movement of an event or episode across an axis can then be represented as:
Δ πk = πk (si+) – πk (si–)
This allows the model to capture not simply whether a story contains terms associated with life and death, for example, but whether an episode moves from one pole towards another, suspends the opposition, or mediates between them.
Proposed Method
The corpus will consist of myths and fairy tales segmented into narratively meaningful episodes. Each episode will be converted into a relational representation containing extracted elements, role assignments, event relations, before-and-after states, and projections onto relevant contrastive axes. A single episode representation may include the following components:
| Component | Example |
|---|---|
| Actor and role | Fox as deceiver |
| Action relation | Deceives crow |
| Object relation | Food passes from crow to fox |
| Initial state | Crow possesses valued object |
| Final state | Fox possesses valued object |
| Tension | Appearance/reality; speech/action |
| Transformation | Inferior agent reverses possession through indirect action |
Narratives will then be represented as ordered sequences of events:
N = (E1, E2, …., En)
Short overlapping event windows will be formed:
Wi = (Ei, Ei+1, … Ei+m)
These windows function as the narrative equivalent of local image patches in a convolutional neural network. A one-dimensional CNN can scan across the event sequence and learn recurrent local configurations while remaining relatively insensitive to their absolute location in the story. A pattern such as prohibition, transgression, and altered status may therefore be detected whether it occurs near the beginning or middle of a narrative.
The CNN will not be trained as a supervised classifier. Instead, it will be incorporated into a self-supervised contrastive-learning architecture. Positive pairs will be created by generating structurally preserving variants of the same episode or event sequence. These transformations may alter actor identity, species, object, location, wording, and absolute narrative position while preserving role structure, event order, transfer relations, and before-and-after transformations. Negative pairs will consist of episodes whose relational organisation differs.
The training objective will encourage representations of structurally corresponding episodes to move closer together while separating unrelated configurations:

Here zi is the learned representation of a narrative window, zi+ is a structurally preserved variant, sim is cosine similarity, and τ is a temperature parameter. This approach is unsupervised with respect to mytheme labels, but theoretically guided in its definition of what should remain invariant.
The learned representations will then be clustered. Each cluster will be analysed to identify the features that remain stable across its members: role topology, event order, transformation, opposition profile, and direction of transfer. A cluster may, for example, contain stories with different characters and settings but share the pattern:
prohibition → boundary crossing → acquisition → irreversible change
Such a cluster would be treated as a candidate mytheme.
Model Architecture
The initial model will combine three complementary components. A relational event encoder will represent individual episodes. A one-dimensional CNN will detect recurrent local sequences of events. A clustering stage will group similar latent configurations. The CNN is appropriate for patterns whose event order and local sequence are significant.
A later extension may incorporate a graph neural network to model role topology and an attention mechanism to capture relations separated by long narrative distances. This is important because not all mythemic structures are local. A prophecy, prohibition, substitution, or kinship relation may be introduced early and resolved much later. The first phase should nevertheless remain deliberately constrained, testing whether recurrent local relational patterns can be identified before introducing a more complex hybrid architecture.
Evaluation
The project will be evaluated at three levels. First, structural invariance tests will assess whether the model recognises comparable patterns when actors, species, objects, locations, and wording are changed. Second, cluster coherence will be assessed by examining whether grouped episodes share interpretable relational and transformational structures rather than superficial vocabulary or common provenance. Third, expert interpretation will determine whether the resulting clusters constitute plausible candidate mythemes.
Evaluation should also test for unwanted shortcuts. The model must not cluster primarily by story collection, translator, genre label, recurring character name, or lexical similarity. Ablation experiments will therefore remove or mask actor names, species, locations, and source metadata in order to determine whether discovered patterns depend on relational structure.
Expected Contribution
The project proposes a computational account of the mytheme as a recurrent transformation of relational configurations. It moves beyond motif detection and semantic similarity by treating narrative meaning as distributed across roles, tensions, transfers, and changes of state.
The method is neither wholly theory-free nor conventionally supervised. It is best described as weakly guided unsupervised or self-supervised structural learning. Narrative events are represented according to an explicit relational ontology, but the mythemes themselves are not supplied in advance. The system instead proposes recurring structures whose interpretive and theoretical significance can subsequently be assessed.
The broader research question is whether structural invariants can be learned from variation. If successful, the project would demonstrate that computational models can identify patterns that persist not because the same figures recur, but because different figures occupy comparable relational positions and enact similar transformations. This would provide an empirical and methodological basis for investigating mythemes as latent, recurring structures across bodies of mythic and literary narrative.