arxiv:2409.04269

Open Language Data Initiative: Advancing Low-Resource Machine Translation for Karakalpak

Published on Sep 6

· Submitted by

IAMJB on Sep 10

Authors:

Mukhammadsaid Mamasaidov ,

Abror Shopulatov

Abstract

This study presents several contributions for the Karakalpak language: a FLORES+ devtest dataset translated to Karakalpak, parallel corpora for Uzbek-Karakalpak, Russian-Karakalpak and English-Karakalpak of 100,000 pairs each and open-sourced fine-tuned neural models for translation across these languages. Our experiments compare different model variants and training approaches, demonstrating improvements over existing baselines. This work, conducted as part of the Open Language Data Initiative (OLDI) shared task, aims to advance machine translation capabilities for Karakalpak and contribute to expanding linguistic diversity in NLP technologies.

View arXiv page View PDF Add to collection

Community

AdinaY

Sep 10

Wow this is super cool 🔥 Thanks for your work!!

AdinaY

Sep 10

Dataset :https://huggingface.co/datasets/tahrirchi/dilmash
Models: https://huggingface.co/collections/tahrirchi/dilmash-66c35b12d7a9770138c837fd

Sep 11

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Expanding FLORES+ Benchmark for more Low-Resource Settings: Portuguese-Emakhuwa Machine Translation Evaluation (2024)

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.

Tap or paste here to upload images

· Sign up or log in to comment

Models citing this paper 3

Datasets citing this paper 1

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2409.04269 in a Space README.md to link it from this page.

Collections including this paper 1