Indonesian corpus linguistics toolkit

SemanText

01

Assemble the corpus

Import materials by keyword, URL, CSV export, or raw text

Enter any keywords. Article URLs are collected from news search, then scraped at the edge (up to 25 per round).

Optional date range. Leave both fields empty to scrape without a date filter. Rows without a usable date are always retained and counted.

One URL per line. The first 25 are fetched per round.

Corpora produced by SemanText: Datetime, Title, Text, URL, TextID, Publication.

Paste raw text or upload plain-text files (.txt). Each submission is treated as a single document and is limited to 100,000 tokens.

02

Corpus overview

All imported materials, in one place

02No corpus yet. Assemble one in step 01: import articles by keyword or URL, or add raw text.

03

Analysis tools

All computation belongs to your browser, not to a server

A

Word frequency

B

N-gram analysis

C

Concordance (KWIC)