EXTRACTING PARALLEL PHRASES FROM ENGLISH-PUNJABI CORPORA

Printed Book
SR 367
Inclusive of VAT
Sold as: EACH
SR22Per Month/24 months
Author:Lehal, Manpreet Singh
Date of Publication: 2024
Book classification:Engineering,English Books
No. of pages:204 Pages
Format:Paperback

This book is printed on demand and is non-refundable after purchase

Available Formats :

Printed Book

It will be sent to your address

SR367
Incl. VAT

Choose your delivery preference

Or

About this Product

This study presents a novel approach to extract parallel data from a comparable English-Punjabi corpus, addressing the scarcity of parallel corpora for this language pair. Unlike previous research, this approach focuses on creating high-precision parallel data using minimal resources. The data is sourced from diverse domains, including Wikipedia articles, TDILs noisy parallel sentences, and Gyan Nidhi reports. The methodology consists of three phases: extracting and aligning documents, translating Punjabi texts into English using OpenNMT-py, and calculating content similarity through three measures-Euclidean Distance, Cosine, and Jaccard. These algorithms are run individually, and then their results are integrated to improve accuracy. By combining the scores of all three measures, the system achieves a precision of 93% and an accuracy of 86%. This integrated approach significantly enhances parallel data extraction for English-Punjabi corpora and holds potential for improving Statistical Machine Translation (SMT) models.
Show more

Specifications

SKU9786208225414
Manufacturer Number9786208225414
year published2024
Show more

Report an issue with this product.

Customer Reviews