Indonesian corpus linguistics toolkit

SemanText

01

Compile the corpus

Import texts by keyword, URL, CSV export, or raw text

Enter search terms. Text URLs are collected from news search, then each page is harvested at the edge (up to 25 per round).

Optional date range. Leave both fields empty to harvest without a date filter. Texts without a usable date are always retained and counted.

One text URL per line. The first 25 are processed per round.

Corpora produced by SemanText: Datetime, Title, Text, URL, TextID, Publication.

Paste plain text or upload plain-text files (.txt). Each submission is treated as a single text and is limited to 100,000 tokens.

02

Corpus overview

All texts in the corpus, in one place

02No corpus yet. Compile one in step 01: import texts by keyword or URL, or add raw text.

03

Analysis tools

All computation belongs to your browser, not to a server

A

Word frequency

B

N-gram analysis

C

Concordance (KWIC)