BERTeus
BERT language model for Basque
BERT hitzkuntza eredua euskararako
Descripción (en):
We have trained a BERT (Devlin et al., 2019) model for Basque Language using the BMC corpus (Basque Media Corpus). The training corpus contains 224.6 million tokens, of which 35 million come from the Wikipedia.
Descripción:
BERT (Devlin et al., 2019) hizkuntza eredua entrenatu dugu euskararako BMC corpusa (Basque Media Corpus) erabiliz. Entrenamendurako corpusak 224,6 milioi token ditu, eta horietatik 35 milioi Wikipediatik jaso dira.
Descarga
Enlace para acceder online o descargar:
Tipo:
Grammars and language models
Persona de contacto:
Ander Barrena
Email persona de contacto:
ander.barrena@ehu.eus
Grupo de investigación:
IXA-UPV/EHU
Euskara
Displaying 1 - 20 of 20
Tools and services
Averell
Averell is a Python library and command line interface to download and to standardize corpora from ten multi-lingual poetry repositories |
Jollyjumper
Jollyjumper is our enjambment detection Python library for Spanish |
Rantanplan
Rantanplan is a Python library for the automated scansion of Spanish poetry |
PoetryLab app
PoetryLab: An Open Source Toolkit for the Analysis of Spanish Poetry Corpora |
PDMapping
Tool for documenting and analyzing speakers' judgments about spatial and sociocultural linguistic variation. |
Ferramenta On-Line de ExpeRimentación PerceptivA (FOLErPa)
FOLERPA is an online tool for carrying out perceptual experiments. |
Cartografía dos apelidos de Galicia
Research tool for the study of the geographical distribution of surnames in Galicia. |
Vocabulary analyzer Web Service
This web service calculates different lexicometric measures and displays them graphically (tokens, types, hapaxes & type/token ratio). |
Ngram Statistics de Pedersen
Pedersen's Ngram Statistics Package |
UPF Freeling-based part-of-speech tagger.
This is the UPF Freeling-based part-of-speech tagger. |
Análisis de relaciones de dependencias
This WS performs dependency parsing using Bohnet's graph-based Parser. The input is text in plain text or CoNLL format. The languages supported are English and Spanish. |
Freeling Named Entity Recognition - NER
Freeling-based Named Entity Recognition - NER |
WSD-IXA
Word-Sense Disambiguation |
Ixa pipes
Multilingual NLP tools |
ixaKat
A modular chain of Natural Language Processing tools for Basque |
Maltixa
Statistical Syntactic analyzer for Basque |
Eustagger
Morphosyntactic tagger for Basque |
Xuxen
Spelling and grammar checker for Basque |
BASYQUE
A web application to analyse syntactic variation of Basque dialects |
Analhitza
Category analyzer |