Deep Learning: The Answer for Challenging Character Recognition Projects


Deep Learning OCR simplifies challenging character recognition projects in packaging, shipping and logistics applications.

DPM code on an automotive part being read by a ZebraFS42 scanner

Introduction

E-commerce and global trade growth significantly boost the volume of goods transported. That’s one reason why automation in packaging, shipping, and logistics is becoming increasingly important. Automation also helps companies package products faster and more accurately, improves package sorting and handling, and enhances warehouse management and transportation, thereby enabling businesses to operate more efficiently, meet customer expectations, and stay competitive in today’s fast-paced and rapidly evolving market.

Optical character recognition (OCR), a technology that allows machines to recognise and extract text from images, plays a crucial role in automated systems that scan, sort, and label packages. By quickly converting printed or handwritten text from shipping labels or documents into digital data, OCR eliminates the need for manual data entry, reducing errors and saving time.

Though OCR continues to advance and can address a wide range of packaging and labelling formats, stylised fonts, distorted or hidden characters, reflective surfaces, and intricate backgrounds remain a challenge that can only be addressed by deep learning (DL).

As production, packaging, and sorting line rates increase to meet greater demand, so do labelling requirements and regulations. Packages and shipments need to comply with specific labelling standards, including 1D and 2D barcodes, product identification numbers, as well as allergen and country of origin labelling requirements.

OCR technology allows for the automatic recognition and interpretation of these labels, ensuring compliance and enabling seamless traceability throughout the supply chain, while deep learning helps enhance the accuracy and reliability of the recognition process. This helps companies meet regulatory requirements, enhance inventory management, and improve overall operational efficiency.

OCR Benefits

OCR helps improve the accuracy and traceability of packages and shipments. By automatically reading and interpreting shipping labels, OCR systems can ensure that packages are accurately labelled, sorted, and tracked throughout the supply chain. This reduces the chances of misrouted or lost packages, leading to greater customer satisfaction and improved profit margins.

OCR automatically recognises and records products and serial numbers—along with 1D and 2D barcodes—or QR codes on packaging to enable efficient inventory management. By automating the tracking and updating of stock levels, OCR technology helps reduce manual effort and the risk of human error. Accurate inventory data allows businesses to optimise their supply chain operations and avoid stockouts or overstocking.

By automating the extraction and processing of textual information, OCR contributes to overall operational efficiency. It enables faster document processing, reduces manual intervention, and accelerates decision-making processes. This efficiency gain translates into faster order processing, improved shipment accuracy, and enhanced productivity across various packaging, shipping, and logistics applications.

Benefits
  • Ensure that packaging and labels contain accurate/legible text
  • Inspect and read date and lot codes
  • Improve track-and-trace operations
  • Enable maximum throughout for packaging applications

Traditional OCR Challenges

OCR has evolved to address the challenges presented by diverse packaging and labelling formats, but stylised fonts, distorted or obscured characters, reflective surfaces, and complex backgrounds can pose difficulties for traditional OCR techniques, which are user-taught and therefore typically require industrial imaging professionals for setup, training, and deployment (Figure 1).

Training a traditional OCR system involves several steps. First, the input images are pre-processed to enhance their quality and prepare them for character recognition. This may involve tasks such as noise reduction, image binarization (converting to black and white), and deskewing (adjusting the image to correct for rotation).

Segmentation is next. The pre-processed image is divided into individual characters or text lines. This step separates the characters or lines from each other, making them easier to recognise and analyse separately. Features are then extracted from each segmented character. These features are distinctive characteristics that help differentiate one character from another. Traditional OCR systems use various techniques to extract features, such as analysing the contours, strokes, and other geometric properties of the characters. 

A set of training data is created by labelling the extracted features with their corresponding characters. To train the OCR system effectively, a diverse set of training images is necessary to cover different fonts, sizes, styles, orientations, and noise conditions that the system might encounter in real-world scenarios. Human operators manually annotate each character in the training images to create a dataset that pairs character features with their correct labels.

Then comes classifier training, where a classification algorithm is trained using the labelled training data. This algorithm learns to recognise patterns in the extracted features and associate them with the corresponding characters. Common algorithms used in traditional OCR systems include support vector machines, hidden Markov models, and k-nearest neighbours.

Next, the trained classifier is evaluated using a separate set of test data. This evaluation measures the accuracy and performance of the OCR system. If the performance is not satisfactory, the training process may be iterated by adjusting parameters, improving pre-processing techniques, or expanding the training dataset.

Once the OCR system achieves the desired level of accuracy, it can be deployed to recognise characters in new, unseen images. The pre-processed image is segmented, and the trained classifier is used to recognise and convert the segmented characters into digital text. It’s important to note that traditional OCR systems rely heavily on handcrafted features and specific algorithms, making them less adaptable to different font styles, sizes and languages compared to deep-learning-based OCR systems.

5 examples of printed products which are difficult to scan  using Traditional OCR methods due to reflective surfaces, printing on curved surfaces, and characters that are skewed, distorted, or obscured

Figure 1: Traditional OCR methods may encounter challenges when dealing with reflective surfaces, printing on curved surfaces, and characters that are skewed, distorted, or obscured.

Figure 2: Traditional OCR systems are user-taught, which involves several steps, including pre-processing, segmentation, labelling, classification, and algorithm evaluation.

Examples of challenging character recognition including complex backgrounds, blurry or damaged characters, distortion, or reflective surfaces that make traditional OCR techniques ineffective.

Figure 3: The OCR capabilities provided by the DL-OCR tool offer a solution for challenging character recognition projects. These projects involve complex backgrounds, blurry or damaged characters, distortion, or reflective surfaces that make traditional OCR techniques ineffective. The tool includes a pre-trained neural network, ready for use with thousands of diverse image samples. It achieves approximately 97% accuracy right from the start, even in very difficult scenarios.

Enter Deep Learning

Artificial intelligence and machine learning, and specifically deep learning, present a promising solution to the challenges of OCR solutions. With the use of deep learning algorithms, it becomes possible to detect irregularities in patterns—even when alphanumeric characters are difficult to define using rigid rules—and advancements are underway to address these challenges and improve the OCR process. 

Modern approaches such as deep-learning-based OCR leverage convolutional neural networks (CNNs) and recurrent neural networks (RNNs) to automatically learn and extract features from characters, reducing the reliance on explicitly engineered features. These models can handle a wider variety of fonts and adapt faster and more effectively to new or unfamiliar fonts without requiring extensive manual adjustments.

The process of gathering and annotating large datasets necessary for training deep learning models has, however, emerged as a hindrance to widespread implementation. The number of images required for a dataset varies by the complexity of the code-reading application, and it is often determined through experimentation and iterative refinement to achieve the desired performance.

Updating the OCR process to handle font changes more efficiently is an ongoing area of research and development, aiming to reduce the burden of manual adjustments and enhance the system’s adaptability to new fonts and text variations (Figure 3). To this effect, transfer learning techniques are also being employed to leverage pre-trained models on large datasets, allowing for better generalisation and reducing the need for excessive training data for each specific font.

Zebra’s Deep Learning OCR Tool

Advancements in OCR have made it possible to overcome all these challenges. For example, the industrial-quality Deep Learning OCR (DL-OCR) tool is an add-on to Zebra Aurora Focus™ software that makes reading text quick and easy. DL-OCR comes with a ready-touse neural network that is pre-trained using thousands of different image samples. So it delivers high accuracy straight out of the box, even when dealing with very difficult cases. Users can create robust OCR applications in just a few simple steps—all without the need for machine vision expertise (Figure 4). The intuitive Zebra Aurora Focus software interface makes setup easy, and the system can be deployed in just minutes.

Unlike traditional OCR applications, Zebra has taken a zero-learning approach with its DL-OCR tool, which delivers accurate reads without having to train different texts or fonts up front. The tool also leverages state-of the-art techniques that allow novices to easily set up fast, accurate text recognition and challenging reading applications. Another key advantage of the DL-OCR tool is its ability to more effectively deal with text that is inconsistent or has degraded contrast.

With the new tool, an Aurora Focus DL-OCR deployment can be streamlined to a few quick steps: placing the region of interest (ROI) around the text to be read, fixturing the ROI to a locating tool that facilitates reading when the part moves or rotates within the camera’s field of view, and adjusting a few thresholds that help differentiate between different character strings, including mathematical, alphanumeric size, or spacing as needed.

Optical character recognition OCR scan examples of difficult product labels using Deep Learning OCR

Figure 4: Zebra DL-OCR users can easily create a robust OCR application with just a few simple steps, without having extensive machine vision expertise. 

Optical character recognition OCR scan examples of difficult product labels using Deep Learning OCR

Figure 5: Key features of Zebra's DL-OCR tool include its readiness for use, the ability to handle difficult OCR cases that are not achievable with traditional methods, high accuracy out of the box, user-friendliness, and compatibility with both an NVIDIA GPU and CPU.

Complete Scalability and Optimisation

As part of Zebra’s Aurora software, the DL-OCR tool can be deployed on a range of different products, including any FS fixed industrial scanner or VS smart camera, delivering on-camera deployment that provides a compact, self-contained inspection system. Embedded smart camera platforms present a faster and simpler approach to setting up and capturing images for various OCR applications when compared to traditional industrial PC-based machine vision systems.

However, end users are not locked into deploying the tool just on Zebra devices, as Aurora software can be reliably deployed on third-party industrial PCs and vision controllers available on the market today. With customer needs constantly expanding and evolving, end users need tools that are not only effective but also easy to use. The new DL-OCR tool was designed with this in mind, as it requires no training and can be deployed on edge or distributed systems, delivering users the flexibility needed to deploy a system that meets their exact need and can be scaled and optimised as required.

Regardless of the hardware platform selected, deep learning tools offer several benefits in OCR applications compared to traditional OCR methods. DL-OCR systems, like the one from Zebra Technologies, can read fonts directly out of the box and will continue to learn even more as the Zebra team trains the algorithm on new images and samples for subsequent software releases. This end-to-end learning approach eliminates the need for explicit feature extraction, making the system more adaptable to various fonts, languages and styles (Figure 5). Traditional OCR systems require handcrafted features, which can be less flexible to implement and time-consuming to maintain.

4SightXV6

Zebra VS40 machine vision smart sensor

VS40 Machine Vision Smart Camera

Conclusion


Deep learning models have demonstrated superior performance in character recognition tasks. They can automatically learn and identify complex patterns, making them highly effective in handling variations in characters, noise, and distortion. Deep learning models often achieve higher accuracy rates compared to traditional OCR methods, saving manufactures and integrators time and money while enabling OCR that can more effectively deal with text that might be inconsistent or have degraded contrast.

 

Deploying deep learning in OCR applications doesn’t have to be difficult. Zebra DL-OCR brings significant advantages to the automation of packaging, shipping, and logistics processes, offering an easy-to-deploy OCR solution that can be set up within minutes and effectively address challenges such as automated package scanning, sorting, and labelling tasks while improving accuracy, traceability, and compliance with labelling standards.

 

For more information on Zebra's DL-OCR tool and machine vision software, click below.