About

A research instrument for narrative data.

NarraGrid turns interviews, memories, open-ended survey responses and field notes into structured variables with language models, and keeps the method attached, so every result can be checked against human judgment and described in a methods section.

§ 1

What it is

A shared workspace for research groups. Each lab is an organization with its own projects, members and roles: admins manage the lab, assistants prepare data and run analyses, and raters score text by hand. Within a project, the source text, the prompt, the model outputs and the human ratings sit side by side.

§ 2

How a study moves through it

  1. 01

    Upload

    Bring a CSV or Excel file and choose the column that holds the text, or upload DOCX and PDF documents.

  2. 02

    Write the instrument

    A prompt sets what the model reads and the fields it must return: free text, numbers, yes or no, categories, multiple choice or a Likert scale. A codebook document can be turned into a first draft.

  3. 03

    Run models

    Run the prompt over all rows or a selection with a model from OpenAI or Anthropic. Each output is checked against the fields; rows that fail keep their error and can be retried.

  4. 04

    Rate by hand

    Rating tasks give human raters the same fields, so each model output has a human counterpart.

  5. 05

    Compare

    Agreement is computed per field: percent agreement and Cohen's κ (quadratic-weighted for Likert scales), Pearson r and mean absolute error for numbers, Jaccard overlap for multiple choice, and Krippendorff's α between raters.

  6. 06

    Report

    Each run drafts its own methods paragraph and checklist. Bring runs, ratings and benchmarks into a report that previews as a printable page and exports as Markdown.

§ 3

What every result keeps

Prompt
The exact version used, stored with the run, so later edits do not change past results.
Model
Provider, the model requested, the version the provider reports for each answer, and the settings it was sent with.
Data
The dataset and the rows the run covered.
Run
Who started it, when, and how many rows completed, failed or were skipped.
Agreement
The number of compared ratings, n, next to every coefficient.
Methods
A reporting checklist, a draft methods paragraph and a record of the exact instructions sent to the model.

§ 4

Where your text goes

Datasets, runs and ratings are stored in your organization's workspace. When you start a run, the text of the selected rows is sent with your prompt to the model provider you chose, through its API. Before a run starts, NarraGrid estimates what it will cost.

Organizations can add their own API keys. The keys are stored encrypted, and the provider bills runs made with them to the organization. Prompts are public by default and can be made private; benchmarks start private.

§ 5

Reporting standards

Every analysis run comes with a reporting checklist and a draft methods paragraph, built on two published sets of recommendations for research that codes text with language models:

  • Abdurahman, S., Salkhordeh Ziabari, A., Moore, A. K., Bartels, D. M., & Dehghani, M. (2025). A primer for evaluating large language models in social-science research. Advances in Methods and Practices in Psychological Science, 8(2), Article 25152459251325174. https://doi.org/10.1177/25152459251325174
  • Törnberg, P. (2024). Best practices for text annotation with large language models. Sociologica, 18(2), 67–85. https://doi.org/10.6092/issn.1971-8853/19461

NarraGrid records what it can check itself: the prompt version, the model version the provider reports for each answer, the settings each request was sent with, the dates, and every answer that failed. Where a step needs the researcher, it helps: a consistency check codes a random sample a second time and reports Krippendorff's α, and rating tasks compare the model with human coders. What it cannot check, such as whether errors cluster in one group, stays on the list as a question for the researcher.

§ 6

Access

NarraGrid is in early access, and accounts are by invitation. An invite code from a lab adds you to that lab; an access code lets you set up a lab of your own.