Skip to content

Navigation Menu

Sign in
Sign up

Repository files navigation

@peerbits/transcript-parser

Deterministic post-diarization transcript structuring — speaker normalization, utterance merging, and conversation metadata (not diarization, not AI)

Category: AI Clinical Documentation — Clinical Documentation Components · License: Apache-2.0 · Status: Stable

CI License npm version


1. What problem does this solve?

Raw diarized transcript output from ASR and speech-to-text engines (such as WebVTT, SRT, or provider-specific JSON) is fragmented, noisy, and inconsistently labeled. Single thoughts are frequently split across multiple millisecond timestamps, speaker labels vary arbitrarily ("Speaker 1", "spk_0", "Doctor"), and conversations lack mechanical structure.

@peerbits/transcript-parser normalizes and structures already-diarized transcripts into clean, merged, time-bounded conversation blocks and computes mechanical metrics (speaking times, turn counts, word counts) without touching audio and without calling AI/LLM models.

Scope Note: This library performs deterministic, rule-based structuring of ALREADY-DIARIZED transcripts. It does not perform acoustic speaker diarization, and it does not use AI or LLMs anywhere. See What This Is Not for full architectural boundaries.


2. Features

  • Multi-Format Adapters: Parse standard WebVTT, SubRip (SRT), and generic JSON transcript arrays into normalized raw utterances.
  • Speaker Normalization: Consolidate inconsistent labels ("Speaker 1", "SPEAKER_01", "spk_0") into stable internal identifiers ("speaker_1", "speaker_2").
  • Deterministic Utterance Merging: Merge consecutive same-speaker fragments using a configurable millisecond gap threshold (default 2000ms).
  • Explicit Role Hints: Assign roles (clinician, patient, caregiver) strictly from caller-provided mapping tables. Roles remain null if unmapped — zero content guessing.
  • Mechanical Conversation Metadata: Calculate total conversation duration, speaker turn counts, word counts, speaker changes, and per-speaker speaking time distributions.
  • Zero Dependencies: Pure TypeScript implementation with zero external runtime dependencies.

3. Installation

npm install @peerbits/transcript-parser

4. Demo and Quick Start

Peerbits HealthTech - Transcript Parser Demo

import { parseTranscript } from "@peerbits/transcript-parser";
const rawVtt = `WEBVTT

00:00:01.000 --> 00:00:03.200
<v Speaker 1>Good morning. How are you feeling today?</v>

00:00:03.400 --> 00:00:05.100
<v Speaker 1>I see you are here for a routine checkup.</v>

00:00:06.000 --> 00:00:09.500
<v Speaker 2>Good morning doctor. I have had a mild cough.</v>`;
// Parse and structure deterministically
const result = parseTranscript(rawVtt, {
 gapThresholdMs: 2000,
 roleMap: {
 speaker_1: "clinician",
 speaker_2: "patient",
 },
});
console.log(result.blocks);
// [
// {
// speakerId: "speaker_1",
// role: "clinician",
// text: "Good morning. How are you feeling today? I see you are here for a routine checkup.",
// startTime: 1000,
// endTime: 5100,
// wordCount: 17,
// fragmentCount: 2
// },
// {
// speakerId: "speaker_2",
// role: "patient",
// text: "Good morning doctor. I have had a mild cough.",
// startTime: 6000,
// endTime: 9500,
// wordCount: 9,
// fragmentCount: 1
// }
// ]
console.log(result.metadata.totalDurationFormatted); // "00:00:08.500"
console.log(result.metadata.speakerStats);

5. Architecture

src/
├── adapters/
│ ├── webvtt.ts # W3C WebVTT cue parser & voice tag extractor
│ ├── srt.ts # SubRip (SRT) parser
│ └── generic-json.ts # Flexible JSON array parser
├── normalize-speakers.ts # Label standardizer (spk_1 -> speaker_1)
├── merge-utterances.ts # Time-gap utterance merger & timestamp formatter
├── apply-role-hints.ts # Explicit role mapper (Strict: no content inference)
├── metadata.ts # Mechanical duration/turn/word metrics computation
├── validate.ts # Structural validation self-checks
├── types.ts # Core TypeScript interface definitions
└── index.ts # Public API entry point & pipeline runner

6. Examples

See the /docs/examples directory for end-to-end examples including WebVTT processing, SRT ingest, and generic JSON pipelines.


7. Roadmap

  • WebVTT, SRT, and Generic JSON parsing
  • Configurable time-gap utterance merging
  • Explicit speaker-to-role hint application
  • Mechanical conversation metrics computation
  • Export to formatted Markdown encounter transcript
  • Custom speaker label pattern plug-in interface

8. Contributing

Contributions are welcome! Please read CONTRIBUTING.md before submitting pull requests.


9. License

Apache License 2.0 — see LICENSE.


10. About Peerbits

transcript-parser is part of the Peerbits HealthTech Open Source initiative — reusable engineering components extracted from our healthcare technology work. This repository contains generalized, reusable logic only; it is not tied to any specific client engagement or commercial product.

About

Deterministic transcript structuring for speaker normalization, utterance merging, and conversation metadata.

Topics

Resources

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

AltStyle によって変換されたページ (->オリジナル) /