Technical report
Life Data Stories
Generating and presenting biographical data stories
Abstract
Life Data Stories is an approach that transforms encyclopedic biographical prose into structured data stories. The stories are shown as sequences of slides that combine narrative text, timelines, maps, and social networks. A second story type, the meta story, traces a theme across several lives and presents it as one continuous scrolling document with similar visual encodings. The system separates two concerns. First, generation is performed offline by staged pipelines that combine large-language-model inference with deterministic transformation. Second, the presentation is rendered by an interactive web application without further use of artificial intelligence. This technical report documents the approach and its implementation and provides a basic scientific contextualization. The code is available at github.com/fabian-beck/life-ds and the live application at fabian-beck.github.io/life-ds.
1Introduction
Encyclopedic biography is a rich and well-sourced account of a life, written as continuous prose in long-form articles. Reading a life in that order is natural, but it hardly supports other forms of access, such as casual browsing. Such browsing would be supported by the biography’s temporal, geographic, and relational structure: the phases of a life, its places, and the people involved. However, the reader cannot access this structure directly.
Life Data Stories makes this structure explicit and explorable through specific data representations. It derives from the source prose a set of discrete events. Each event has a date and, where known or available, a place, the people involved, and illustrations. Our approach groups those events into the phases of a life and presents the result as narrative text, as a chronology, as a geography, and as a social network. Further, a visual identity is generated for each story. A similar derivation applied across several biographies yields a meta story as one continuous document that integrates a timeline with a lane per life, a map, and a social network merged from the individual biographies.
Segel and Heer [1], surveying how data and visualization are used to tell stories, place such a presentation on a spectrum between the author-driven, in which a fixed order carries the message, and the reader-driven, in which the audience decides what to look at. A story in Life Data Stories is author-driven at first glance and reader-driven in its depth. The sequence is fixed when the story is generated and can be followed to the end. Alternative exploration affordances are available throughout, in timelines, image galleries, maps, and social network views. A meta story is linear in the same way and links to the personal stories throughout.
BiographySampo [2] is one of the closest precedents with a digital humanities focus. It extracts events, places, and relationships from national biographical dictionaries with language technology and pattern rules, links them to external registers, and offers them through faceted search, maps, timelines, and ego networks. Its tools serve biographical research and show the unchanged biography text next to the data. VisKonnect [3] uses encodings similar to those of Life Data Stories and connects historical figures through the events they have in common. There, a reader’s prompt retrieves the matching events from an event knowledge graph, and an event timeline, an event map, and a relationship graph are shown beside a short answer. Textual answers are generated but refer to intersections of lives and stay disconnected from the visual representation. In contrast, Life Data Stories integrates text and visual representation and proposes a reading order while offering exploration options.
To achieve this, we rely on large language models (LLMs) for interpreting the data. A biography is at first unstructured text, and giving it structure implies deciding which episodes are important, which places they relate to, and which relationships matter. However, this needs to be done only once, before the stories are presented. The system is accordingly built in two parts that connect through the data (Figure 1). Generation runs offline as preprocessing and interleaves model calls with deterministic transformation, so steps that research or review material alternate with steps that resolve, cluster, merge, and validate the data. The processing produces a set of documents, data, and images describing one life or one theme. The interface loads those files and renders them in a single-page web application.
2Data model
The data that connects generation and interface is a semi-structured representation of a biography and the central abstraction of the system. For a single life, it holds several structured perspectives: the events of the life with their text, pictures, and geography. While these parts are relatively straightforward derivatives of the textual biographical source, two further perspectives are more interpretive: the social network of the subject, which the text rarely describes explicitly, and the story’s visual identity.
A meta story refers to the data of its subjects and adds further perspectives. It lists the subjects and groups them into subtopics. Its chapters divide the covered time span into periods, and each chapter references events from the subjects’ data, annotated with their relevance to the theme. The subjects’ ego networks are merged into one social network, and the locations of their events are clustered into places relevant to the theme. The text is written for the theme, and the pictures are selected from the events of the individual stories. Like a personal story, a meta story is presented in a generated visual identity.
3Generation
We split the generation of this data into two pipelines. The first, the personal story pipeline, reads an encyclopedia article and writes the data of an individual biography. The second, the meta story pipeline, reads the data of several individuals and writes a meta story connecting them. Both pipelines run offline as Python command-line scripts, invoked for one life or one theme at a time. A run stores the data described in the previous section as JSON documents. Each step reads the documents of earlier steps, writes its own, and is one of four kinds.
Both pipelines follow the same pattern: material is derived bottom-up and revised top-down. Both end in translation: each English document is translated into every other supported language (currently: German).
3.1Personal story pipeline
The personal story pipeline derives data for one biography, from an encyclopedia article to the data needed for an illustrated and individually styled story. The chart (Figure 2) groups its steps into seven stages. Once the sources are loaded, the events of the life are proposed from them, researched, geocoded, and matched with pictures. The story is then styled individually, and the generated pictures adopt that style. The social network is derived, and a review pass revises it together with the events. Background reports then add depth to selected events, and localization closes the run.
Sources. An acquisition step loads the subject’s Wikipedia article and linked article titles. An AI-based selection call chooses the most relevant titles before their full texts are fetched. The fetched articles are cached for reuse by later steps and reruns, together with Commons image metadata.
Events. The events and the main narrative are derived from this material. The first AI-based call proposes twelve to sixteen events with titles, dates, descriptions, an optional classification, and a weight. It also groups them into three to six chapters. An AI call per event then researches each event in detail against the relevant articles: the historic and modern name of its place, the people involved, a semantic icon, and image search queries. Geocoding looks up exact coordinates for each place by its modern name, if available.
Images. A planning step generates image queries for the person and the subjects of the events, which complement the per-event queries for specific objects such as a building or a document. The search queries are then executed against Wikimedia Commons, the life-wide ones also against Openverse, and the candidates are scored in code on resolution, bytes per pixel, filename, and a penalty for certain subjects (e.g., a plaque only remembering an event). An AI-based matching call then assigns pictures to events, writes captions, keeps each picture’s attribution, and picks a reference portrait. Because we consider the portrait the central visual element and the matching call judges by metadata only, a vision call checks whether it depicts the subject.
Presentation. A color and type system is generated, inspired by the subject’s own work. The step returns a color palette with a sentence justifying it, a tileable background pattern, a separator glyph, and a heading and a body font. The style is checked through rule-based heuristics regarding tile, font, and contrasts, and regenerated if necessary. Based on the style, especially its color palette, illustrations are generated. The reference portrait is style-transferred toward one master style shared by all subjects. Each chapter is turned into an abstract visual concept in one call over all chapters, and each concept is illustrated in the same style.
Network and review. The ego network is derived as typed, weighted, and described relationships. To increase coherence and avoid redundancies, a top-down review pass over the events and the network refines descriptions and applies high-confidence changes. Reading the slides in order, it potentially rewrites descriptions that rely on terms not yet introduced or repeat the previous slide.
Depth. A background report of a few paragraphs—explaining details of an event or the background of a key concept relevant to the event (e.g., an invention)—is written after the review for the most important events, roughly one event per chapter, and only where the sources provide enough material on the event. An image critic call compares the illustrations found for each report with its text and removes images that do not match its topic.
Localization. To support multiple languages, a lookup collects the target language’s names of the subject, the people, and the places from Wikipedia’s language links. A glossary call then fixes one translated name per person, so that each person is named the same way throughout the story. Translation calls finally translate all text fields, originally written in English, into each target language.
3.2Meta story pipeline
The meta story generation (Figure 3) concerns a theme that connects multiple lives (e.g., an area of science, art, or political development) and directly builds upon the output of the personal pipelines. The story is planned and its events collected and curated against the theme. Moreover, the social network and the map are derived. Then, a composition step joins the branches into one story and gives it a visual identity. Localization again comes last.
Planning. An AI planning call reads summaries of available personal stories and selects the subjects for the theme. It groups the selected people into two to four subtopics, proposes three to six chapters as named eras with estimated date ranges, and drafts a description and a conclusion.
Events. Batched curation calls judge all events of the selected people against the theme and keep the most relevant ones. The proposed chapters are then fitted to the filtered events. A context call adds up to two general historical events per chapter that directly affected the story’s people (e.g., a war).
Network. The network construction starts from the subjects’ ego networks. An initial merge joins them by normalized name into one graph of the subjects and the related people who appear in two or more of their ego networks. Since each ego network was generated for one person, a review call reads the merged graph against excerpts of the subjects’ articles and adds missing ties, rewords existing ones, and deletes indirect ones, adapting how much it changes to the density of the graph. Community detection by greedy modularity [4] yields the story’s circles (i.e., clusters of people).1 A narration call writes a title and a short description per circle.
Map. A rating call scores all filtered geo-located events by the importance of their place for the story. Clustering merges events into clusters whose events all lie close together and keeps the clusters with sufficient total event weight, ordered by the mean year of their events. A narration call writes a title and a short description per cluster and decides which clusters the map keeps as stops, discarding clusters whose events have nothing in common beyond their city and places of little relevance to the theme.
Composition. The composition call assembles the story from the generated material: it revises and connects the descriptions into a coherent story and decides the order of the sections and the final grouping of the circles. A style call derives from the story’s framing a color palette, background pattern, fonts, and ornamental elements.
Localization. The translation proceeds as for personal stories but reuses the subjects’ names and event titles already translated for their personal stories, so that both agree in wording.
3.3AI models and prompting
The generation uses OpenAI’s API models (gpt-5.6-luna, gpt-5.6-sol, gpt-5.6-terra, gpt-image-2.5-flare, gpt-image-2.5-sunburst). We configure the models per call site to only spend frontier intelligence and reasoning effort where needed and keep generation costs manageable (about 1 EUR per generated story). Composition, the one step that rewrites a whole assembled story, uses the largest model, and the remaining text steps use either a reasoning tier or a small model. A step uses the small model when its output is validated against existing entities, rewritten by a later step, or backed by a deterministic fallback. More advanced tasks use the reasoning tier, among them proposing the events of a life, selecting the people of a theme, reviewing another step’s output, the glossary, and the translation. For prompt design and context engineering, we tried to anticipate which materials a call needs and to limit the prompt to those, while still providing sufficient background. In the steps that make one call per event, the material shared by all calls is placed at the start of the prompt so that the API caches this prefix. The steps that write the story’s prose share one instruction block on writing style. It discourages typical patterns of AI-generated prose, for instance, contrasting a fact with an artificial alternative. Most text responses are requested as structured output against a Pydantic JSON schema.
4Interface
Life Data Stories’ frontend is a mobile-first Svelte application that loads and renders the generated data. The layout is designed for the phone but adapts responsively to larger screens; the screenshots in this report are deliberately captured as a mix of phone and desktop widths.
The landing page offers access to the two story types, which organize the data differently: a personal story follows one life in order, and a meta story follows a theme across several lives. On the landing page (Figure 4), the meta stories are displayed in a carousel at the top. Below them, the personal stories can be reached and filtered by role or search. Additionally, a map plots the place of every event across the corpus. Nearby events merge into circles that grow with their number of events and take the color of a person who contributes more than half of them, fading to gray as the share approaches half. Tapping a marker there shows the event, and another tap opens the event as part of the respective personal story.
In the following, we walk through both story types in detail. Throughout, we use the meta story Architecture as Living Form and Antoni Gaudí’s personal story as running examples.
4.1The meta story
A meta story is structured as an opening scene, up to three sections (a timeline, a map, and a social network), and a conclusion, followed by a grid linking to the personal stories of its subjects. The story is read from top to bottom by scrolling. Architecture as Living Form follows nine architects who explored organic shapes in architecture and worked across more than a century.
Timeline. After a short textual introduction, the vertical scrolling turns horizontal as the reader reaches the timeline visualization (Figure 5). In Architecture as Living Form, a timeline of nine lanes, one per architect, shows events as dots along each lane. A sequence of guiding story cards, marking chapters in the development, is pinned near the top and leads the reader through the timeline. A second layer above the lanes marks historical events affecting several lives at once, such as World War II. Selecting a dot, or stepping from one to the next with the arrow buttons, opens a brief description of the event. The surrounding card links to that architect’s personal story.
Map. The second section shifts the focus to the locations where the architects’ work took place. A non-interactive map (Figure 6) fills the background behind the scrollable story cards, flying to a new location as each one scrolls into view. It zooms to a point or fits a bounding box depending on how spread out the events are. Each event listed on a card links directly to the respective personal story.
Network. The third section is a graph (Figure 7) that displays the architects and the documented connections between them. Its force-directed layout is calculated in the background and finalized before it is displayed. A weak force pulls each subject toward a horizontal position given by their birth year, so that the graph tends to read chronologically from left to right. The precomputed data groups people into clusters. As the reader scrolls, one card per cluster moves over the graph and describes it, and the people and connections involved are highlighted, while the rest of the graph darkens.
The meta story ends with a brief summary of the nine lives, followed by a grid of tiles linking to the architects’ personal stories, which we use here to switch to Antoni Gaudí.
4.2Personal story
Gaudí’s personal story begins with the stylized portrait and a short biographical summary. The story is a horizontal sequence of slides, navigated by swiping sideways. Chapter slides, showing a headline and an illustration, structure the event slides (Figure 8) that provide the main information. Beside the details of the event, an image may be shown (and can be enlarged as part of an image gallery). For events with supporting material, the slide extends vertically: a chevron reveals a longer background passage about the context of the event. Special events such as birth, death, publication, or migration contain further structuring elements and specialized representations: for instance, a birth slide adds a box with the subject’s parents and birth name, and a migration draws an animated path from the old place to the new one on the map.
Timeline. As an explicit representation of time, a row along the bottom serves as both a progress marker and an outline. Each event is shown as an icon reflecting its type and each chapter as a plain marker. Tapping the chapter label above the row unfolds it into a full-screen outline of the life (Figure 9). Each entry is set in from the left in proportion to the subject’s age at that point, so the pace of the life stays visible where the list is dense. Tapping an entry moves the story to that slide.
Map. The map stays in place across slides and adds a marker at each event’s main place, so that by the end of the story it shows everywhere the subject worked and lived. In the example, the event’s main marker is in Barcelona, while faded smaller markers show the places of earlier events, in Paris and in Reus near Barcelona.
Network. When an event slide mentions another person, that person’s name appears beside the text with their role, as here for Eusebi Güell, tagged as patron. Tapping the tag opens a card describing the relationship. From this card, or the button at the top of the screen, readers can access Gaudí’s full network (Figure 10). It is organized by relationship category, family first and then the others in alphabetical order—academic, business, professional, and religious in Gaudí’s case. Each category shows a short introductory passage with the mentioned names and roles highlighted. Its people appear as chips grouped by role: the family as generations around the subject, joined by lines, and the other categories as chips labeled by role, gathered into boxes where a role repeats. Tapping a chip opens the description and strength of that relationship.
Readers swipe through the story to its end, where a conclusion is presented together with up to five related people whose personal stories are available. Relatedness is indirect: a person qualifies by sharing at least one primary role with the subject and ranks higher when also appearing in the subject’s network.
5Discussion and conclusion
This report has introduced the implementation of Life Data Stories, an approach for transforming unstructured biographical text into structured stories about individual lives and the themes that connect them. It defines the main structuring elements of these stories: life events, imagery, geography, and social network views, which together provide affordances for interactive exploration beside the linear reading of the narrative text. The stories share a consistent design, while their visual identity adapts to the specific person. The approach relies on AI-based extraction to derive the underlying structured data. This step runs as preprocessing, so the presentation layer stays independent of further AI calls. At the same time, the resulting stories remain highly interactive through the extracted data layers and the connections between them.
While this report provides implementation details and a basic scientific contextualization, a broader scientific discussion is beyond its scope. Such a discussion would examine related research and its empirical results in more detail. Likewise, an evaluation of the approach remains future work. As the approach targets a broad audience, feedback from this audience will be an important part of such an evaluation. Beyond usability and the perceived value of the approach, it should investigate to what extent users accept AI-based transformations of biographical sources and what expectations they have regarding such transformations.
6References
- E. Segel and J. Heer, “Narrative Visualization: Telling Stories with Data,” IEEE Transactions on Visualization and Computer Graphics, vol. 16, no. 6, pp. 1139–1148, 2010, doi: 10.1109/TVCG.2010.179.
- E. Hyvönen, P. Leskinen, M. Tamper, H. Rantala, E. Ikkala, J. Tuominen, and K. Keravuori, “BiographySampo – Publishing and Enriching Biographies on the Semantic Web for Digital Humanities Research,” in The Semantic Web (ESWC 2019), vol. 11503, pp. 574–589, doi: 10.1007/978-3-030-21348-0_37.
- S. Latif, S. Agarwal, S. Gottschalk, C. Chrosch, F. Feit, J. Jahn, T. Braun, Y. C. Tchenko, E. Demidova, and F. Beck, “Visually Connecting Historical Figures Through Event Knowledge Graphs,” in 2021 IEEE Visualization Conference (VIS), pp. 156–160, doi: 10.1109/VIS49827.2021.9623313.
- A. Clauset, M. E. J. Newman, and C. Moore, “Finding Community Structure in Very Large Networks,” Physical Review E, vol. 70, no. 6, Art. no. 066111, 2004, doi: 10.1103/PhysRevE.70.066111.