HomeGuides

How to Convert Old Hindi Documents to Unicode

By Anjali Sharma · Hindi Localization Specialist · Updated May 29, 2026

Short answer: To convert an old Hindi document to Unicode, first identify the legacy font it uses (often KrutiDev), extract the text, run it through the matching legacy-to-Unicode converter, and proofread the result. Editable files like Word convert most easily; PDFs need their text extracted first; and scanned images require Hindi OCR before any conversion is possible.

Start by identifying the format

Before converting anything, work out what you actually have, because the route differs sharply. An editable document (such as a Word file) that shows Hindi only when a legacy font is installed is the easy case. A PDF may contain real text or may be an image of text. A scanned document is just a picture and contains no text data at all until it is recognised. Knowing which of these you are holding determines every step that follows.

A quick check for legacy text is to change its font to a plain Roman font: if the Hindi turns into English letters and symbols, it is a legacy font like KrutiDev and needs conversion, not just a font change. If it stays as readable Hindi, it may already be Unicode.

Converting editable documents

Editable files are the most straightforward. Copy the Hindi text out of the document, paste it into the appropriate legacy-to-Unicode converter (a KrutiDev-to-Unicode converter for KrutiDev text), run the conversion, and paste the Unicode result back into a new document. Do this in sections for a long file so you can proofread as you go and keep track of your place.

Remember that converters work on plain text, so formatting such as headings, bold, and layout is not carried across. Convert the text first, then reapply the formatting in the new Unicode document. Keep the original file untouched until you have confirmed the converted version is correct.

Source typeExtra step neededThen
Editable doc (legacy font)NoneCopy → convert → proofread
PDF with real textExtract/copy the textConvert → proofread
PDF that is an imageHindi OCR firstConvert if still legacy → proofread
Scanned documentHindi OCR firstConvert if needed → proofread

Converting PDFs

PDFs are two problems wearing one file extension. If the PDF contains selectable text, you can copy that text out and convert it like any editable document — though be aware that copying from PDFs built on legacy fonts can sometimes scramble characters, so check the extracted text before converting. If the PDF is really an image of a page, there is no text to copy, and you must first run Hindi OCR (optical character recognition) to turn the picture into text.

Whether OCR gives you legacy or Unicode text depends on the tool. Good modern Hindi OCR often outputs Unicode directly, in which case no further conversion is needed. If it produces legacy text, convert that to Unicode as a second step. Either way, PDFs almost always require more checking than clean editable files.

Converting scanned documents

A scan is purely an image, so it holds no text until recognition is done. The process is: run the scan through Hindi OCR to produce text, confirm whether that text is Unicode or a legacy encoding, convert if necessary, and then proofread thoroughly. OCR on Hindi has improved a great deal but is not perfect, especially with old print, faint scans, or unusual fonts, so expect to correct more errors than with a born-digital file.

For important scanned records, budget time for careful proofreading against the original image. Names, numbers, and dates deserve particular attention, because an OCR slip there is both easy to miss and costly to get wrong.

Converting many documents at once

If you are facing an archive rather than a single file, plan the work rather than converting ad hoc. Group the documents by source type — editable files, text PDFs, and scans — because each group follows a different route and it is efficient to handle them in batches. Start with the editable files, which convert fastest and cleanest, then the text PDFs, and leave the scans for last since they need OCR and the most proofreading.

Keep a consistent naming and folder scheme so the Unicode versions map clearly to their originals, and record which documents have been checked. For a large or important archive, it is worth converting and proofreading a small sample first to gauge how clean the source is before committing to the whole set. That sample tells you whether the job is a quick copy-and-convert or a careful, proofread-heavy effort.

Proofread, then preserve

Whatever the source, the final step is always the same: read the Unicode result against the original and fix any characters that did not translate cleanly. Complex conjuncts, nukta marks, and unusual symbols are the usual suspects. Once the Unicode version is confirmed accurate, it becomes your portable, future-proof copy — searchable, editable, and readable on any device.

Keep the original legacy file or scan as a reference archive rather than deleting it. Conversion does not harm the source, and retaining it means you can always recheck a questionable word later. With the Unicode version in hand, your old Hindi content is finally free of the font dependency that kept it locked to a single program or machine.

Frequently asked questions

How do I convert an old Hindi Word document to Unicode?

Copy the Hindi text, paste it into the matching legacy-to-Unicode converter (such as KrutiDev-to-Unicode), run it, and paste the Unicode result into a new document, then proofread.

Can I convert a Hindi PDF to Unicode?

If the PDF has selectable text, copy and convert it. If it is an image, run Hindi OCR first to extract text, then convert to Unicode if the result is still a legacy encoding.

How do I convert a scanned Hindi document?

A scan is an image, so run it through Hindi OCR to produce text, check whether that text is Unicode or legacy, convert if needed, and proofread carefully against the original.

Will converting keep my document formatting?

No. Converters handle plain text, so headings and styling are not carried over. Convert the text, then reapply formatting in the new Unicode document.

Which characters should I check after converting?

Complex conjuncts, nukta marks, and unusual symbols convert least reliably. Proofread these along with names, numbers and dates against the original.

Should I delete the original legacy file after converting?

No. Keep it as a reference archive. Conversion does not harm it, and retaining it lets you recheck any questionable word against the source later.

Anjali Sharma — Hindi Localization Specialist
Anjali works on Devanagari typography and document localization, helping offices and publishers move between legacy KrutiDev files and modern Unicode. She writes plain-English guides for people who just need their Hindi text to work everywhere.