Skip to main content

sbx-eng-tokenization-stanford

Standard reference Information

Manning, Christopher D., Mihai Surdeanu, John Bauer, Jenny Finkel, Steven J. Bethard, and David McClosky. 2014. The Stanford CoreNLP Natural Language Processing Toolkit In Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics: System Demonstrations, pp. 55-60.

Analysis citation Information

Språkbanken (2026). sbx-eng-tokenization-stanford (updated: 2026-05-12). [Analysis]. Enriched and distributed by Språkbanken. https://doi.org/10.23695/nzf0-jm46
BibTeX Additional ways to cite the dataset.
English tokenization with Stanford CoreNLP

Example

This analysis is used with Sparv. Check out Sparv's quick start guide to get started!

To use this analysis, add the following line under export.annotations in the Sparv corpus configuration file:

- stanford.token  # Token segments

In order to use this annotation you need to add the following settings to your Sparv corpus configuration file:

metadata:
  language: eng

classes:
  token: stanza.token

For more info on how to use Sparv, check out the Sparv documentation.

Example output:

<token>Language</token>
<token>is</token>
<token>the</token>
<token>human</token>
<token>ability</token>
<token>to</token>
<token>acquire</token>
<token>and</token>
<token>use</token>
<token>complex</token>
<token>systems</token>
<token>of</token>
<token>communication</token>
<token>.</token>

Other references

Type

  • Analysis

Task

  • tokenization

Unit

  • token

Licence

MIT

Dependencies

Keyword

  • stanford-parser

Created

2020-04-24

Updated

2026-05-12

Contact

sb-info@svenska.gu.se