Narrative Landscape: Mapping Narrative Dispositions Across LLMs
A new study published on arXiv introduces a quantitative framework for profiling Large Language Model (LLM) dispositions, defined as stable, model-specific regularities in output under controlled conditions. The research evaluates six frontier models using a structured narrative constraint-selection task across three instruction types. The authors operationalize disposition through two key dimensions: consistency, measured by Jaccard similarity of cross-replication selection overlap, and diversity, assessed via the inverse Simpson index to gauge dispersion across options. Additionally, the paper presents the 'Narrative Landscape,' a PCA-based visualization tool that maps each model's selection profile into a shared space for direct comparison. The findings reveal a distinct rigidity-exploration spectrum among different model families. Crucially, the study demonstrates that varying instruction types can shift the geometry of selection spaces, even when scalar metrics remain similar. This indicates that comparable scores may mask qualitatively distinct selection topologies, highlighting the need for deeper structural analysis beyond simple metrics to understand LLM behavior and biases in narrative generation tasks.
Wire timeline
Narrative Landscape: Mapping Narrative Dispositions Across LLMs
A new study published on arXiv introduces a quantitative framework for profiling Large Language Model (LLM) dispositions, defined as stable, model-specific regularities in output under controlled conditions. The research evaluates six frontier models using a structured narrative constraint-selection task across three instruction types. The authors operationalize disposition through two key dimensions: consistency, measured by Jaccard similarity of cross-replication selection overlap, and diversity, assessed via the inverse Simpson index to gauge dispersion across options. Additionally, the paper presents the 'Narrative Landscape,' a PCA-based visualization tool that maps each model's selection profile into a shared space for direct comparison. The findings reveal a distinct rigidity-exploration spectrum among different model families. Crucially, the study demonstrates that varying instruction types can shift the geometry of selection spaces, even when scalar metrics remain similar. This indicates that comparable scores may mask qualitatively distinct selection topologies, highlighting the need for deeper structural analysis beyond simple metrics to understand LLM behavior and biases in narrative generation tasks.
cs.AI updates on arXiv.org