<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE root>
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="1.1d1" xml:lang="en"><front><journal-meta><journal-id journal-id-type="publisher">Kazakhstan journal for oil &amp; gas industry</journal-id><journal-title-group><journal-title>Kazakhstan journal for oil &amp; gas industry</journal-title></journal-title-group><issn publication-format="print">2707-4226</issn><issn publication-format="electronic">2957-806X</issn><publisher><publisher-name>KMG Engineering</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="publisher-id">109018</article-id><article-id pub-id-type="doi">10.54859/kjogi109018</article-id><article-categories><subj-group subj-group-type="heading"><subject></subject></subj-group></article-categories><title-group><article-title>Digital monitoring of the use of personal protective equipment at industrial facilities using neural network architectures</article-title></title-group><contrib-group><contrib contrib-type="author"><name name-style="western"><surname>Abdimanap</surname><given-names>Galymzhan S.</given-names></name><email>g.abdimanap@kmge.kz</email><uri content-type="orcid">https://orcid.org/0000-0003-1676-4075</uri><xref ref-type="aff" rid="aff-1"/><xref ref-type="aff" rid="aff-2"/></contrib><contrib contrib-type="author"><name name-style="western"><surname>Alimova</surname><given-names>Anel N.</given-names></name><bio>&lt;p&gt;PhD&lt;/p&gt;</bio><email>a.alimova@kmge.kz</email><uri content-type="orcid">https://orcid.org/0000-0002-5155-2417</uri><xref ref-type="aff" rid="aff-1"/></contrib><contrib contrib-type="author"><name name-style="western"><surname>Bostanbekov</surname><given-names>Kairat A.</given-names></name><bio>&lt;p&gt;PhD&lt;/p&gt;</bio><email>k.bostanbekov@kmge.kz</email><uri content-type="orcid">https://orcid.org/0000-0003-2869-772X</uri><xref ref-type="aff" rid="aff-1"/></contrib><contrib contrib-type="author"><name name-style="western"><surname>Abdrakhmanov</surname><given-names>Renat E</given-names></name><email>r.abdrakhmanov@kmge.kz</email><uri content-type="orcid">https://orcid.org/0009-0004-1997-9818</uri><xref ref-type="aff" rid="aff-1"/></contrib><contrib contrib-type="author"><name name-style="western"><surname>Orenkyzy</surname><given-names>Dana</given-names></name><email>d.orenkyzy@kmge.kz</email><uri content-type="orcid">https://orcid.org/0009-0000-2265-2091</uri><xref ref-type="aff" rid="aff-1"/></contrib><contrib contrib-type="author"><name name-style="western"><surname>Nurseitov</surname><given-names>Daniyar B.</given-names></name><bio>&lt;p&gt;Cand. Sc. (Physics and Mathematics),&amp;nbsp;professor (associate)&lt;/p&gt;</bio><email>d.nurseitov@kmge.kz</email><uri content-type="orcid">https://orcid.org/0000-0003-1073-4254</uri><xref ref-type="aff" rid="aff-1"/><xref ref-type="aff" rid="aff-2"/></contrib></contrib-group><aff id="aff-1">KMG Engineering</aff><aff id="aff-2">Satbayev University</aff><volume>8</volume><issue>3</issue><history><pub-date date-type="received" iso-8601-date="2026-06-22"><day>22</day><month>06</month><year>2026</year></pub-date><pub-date date-type="accepted" iso-8601-date="2026-07-30"><day>30</day><month>07</month><year>2026</year></pub-date></history><permissions><copyright-statement>Copyright © , Abdimanap G.S., Alimova A.N., Bostanbekov K.A., Abdrakhmanov R.E., Orenkyzy D., Nurseitov D.B.</copyright-statement></permissions><abstract>&lt;p&gt;&lt;strong&gt;ABSTRACT&lt;/strong&gt;&lt;/p&gt;&#13;
&lt;p&gt;&lt;strong&gt;Relevance. &lt;/strong&gt;Ensuring employees occupational safety at the oil and gas industry facilities remains a relevant objective, as far as traditional monitoring methods for the use of Personal Protective Equipment (PPE) are based on manual visual inspections and are susceptible to human factor. Most existing computer vision systems are limited to detecting only a small number of 2–6 classes of PPE categories and to verifying the anatomical consistency between detected protective equipment and the corresponding body parts of employees.&lt;/p&gt;&#13;
&lt;p&gt;&lt;strong&gt;Study purpose. &lt;/strong&gt;Developing and validating a digital monitoring method for compliance with PPE use requirements based on neural network architectures, integrating algorithms for object detection, human pose estimation, and anatomical matching of PPE elements in real time mode.&lt;/p&gt;&#13;
&lt;p&gt;&lt;strong&gt;Materials and Methods. &lt;/strong&gt;A unique dataset of 16 352 images (after augmentation) containing 13 object classes, including 6 types of PPE and 6 negative classes was created for training the model. The YOLOv8 model was used for object detection, and HRNet for human pose estimation. A two-tier video stream processing architecture was implemented, combining object tracking (BoT-SORT), spatial and anatomical matching and TensorRT quantization to improve system performance.&lt;/p&gt;&#13;
&lt;p&gt;&lt;strong&gt;Results. &lt;/strong&gt;During the training phase, the YOLOv8 model achieved Precision = 0.98, Recall = 0.97, and F1 = 0.94. When testing the developed system on 16 video files obtained from industrial sites, the system achieved a precision of 96.58%, a recall of 68.48%, and an F1 score of 0.8014. Detecting small objects (gloves) in cropped images improves detection efficiency by 2-3 times compared to full-frame processing.&lt;/p&gt;&#13;
&lt;p&gt;&lt;strong&gt;Conclusion. &lt;/strong&gt;The developed approach, combining the YOLOv8 and HRNet models and an anatomical matching algorithm, provides effective digital monitoring of compliance with PPE requirements in real-world production conditions. The obtained results confirm the potential of its application in the development of intelligent industrial safety monitoring systems at industrial facilities.&lt;/p&gt;</abstract><kwd-group xml:lang="en"><kwd>industrial safety, personal protective equipment (PPE), computer vision, YOLOv8 neural network, HRNet, human pose estimation, object detection.</kwd></kwd-group><kwd-group xml:lang="kk"><kwd>өндірістік қауіпсіздік, жеке қорғаныс құралдары (ЖҚҚ), компьютерлік көру, YOLOv8 нейрондық желісі, HRNet, адамның позасын бағалау, объектілерді детекциялау.</kwd></kwd-group><kwd-group xml:lang="ru"><kwd>производственная безопасность, средства индивидуальной защиты (СИЗ), компьютерное зрение, нейросеть YOLOv8, HRNet, оценка позы человека, детекция объектов</kwd></kwd-group></article-meta></front><body>&lt;p&gt;&lt;strong&gt;Introduction&lt;/strong&gt;&lt;/p&gt;&#13;
&lt;p&gt;In modern industry employee safety is a key element of effective process organization. Failure to comply with personal protective equipment (PPE) usage regulations poses a direct threat to the health and lives of employees, leading to the implementation of innovative technologies to automate the PPE use monitoring. Traditional approaches based on manual visual inspections and periodic checks are susceptible to human error and delays in identifying  safety rules violation [1, 2].&lt;/p&gt;&#13;
&lt;p&gt;Today, there are two main approaches to PPE monitoring such as sensor-based methods based on radio frequency identification (RFID) and the Internet of Things (IoT) [1, 2] and machine vision methods using deep learning algorithms [3–15]. The second group of methods has attracted particular attention due to their ability to analyze visual data without the involvement of additional means. However, existing solutions are primarily limited to detecting a few classes of objects and are not always capable of verifying the correct wearing of PPE.&lt;/p&gt;&#13;
&lt;p&gt;The purpose of the study is to develop a digital monitoring method for compliance with PPE use standards at industrial sites, combining employee detection (YOLOv8), human pose estimation (HRNet), detection of 13 classes of PPE objects, and anatomical matching to verify the correct wearing of PPE in real time mode.&lt;/p&gt;&#13;
&lt;p&gt;&lt;strong&gt;&lt;/strong&gt;&lt;/p&gt;&#13;
&lt;p&gt;&lt;strong&gt;Literature overview&lt;/strong&gt;&lt;/p&gt;&#13;
&lt;p&gt;This section reviews key research in the field of automated PPE detection using deep learning methods. The reviewed works reflect the main development directions in this field, including the application of deep learning algorithms, transfer learning methods, and various approaches to improving PPE detection accuracy.&lt;/p&gt;&#13;
&lt;p&gt;Xiong R. and Tang P. [4] proposed a method for detecting compliance with PPE use at construction yards based on employee pose analysis using pose estimation technology, instead of traditional object detection. In the proposed approach, the employee's body is represented as a skeletal model, and individual body parts are used as spatial landmarks to define part attention regions where the presence of relevant PPE elements (e.g., a hard hat on the head, a vest on the upper body) are expected to be worn. To localize these regions, the authors developed a set of rules based on human anatomical features, after which they used convolutional neural networks to classify the extracted image fragments.&lt;/p&gt;&#13;
&lt;p&gt;Experimental results demonstrated the high efficiency of the proposed approach: the F1-score was 96.95% for detecting hard hats and 94.61% for detecting vests. Furthermore, the developed method demonstrated higher accuracy compared to traditional object detection models, including Faster R-CNN and YOLO. These results demonstrate the potential of using human pose estimation methods for automated compliance control of industrial safety requirements and monitoring the use of personal protective equipment (PPE) at construction yards.&lt;/p&gt;&#13;
&lt;p&gt;A number of studies have been devoted to the application of traditional object detection methods for monitoring the use of PPE. In the work of Chen S. and Demachi K. [5], an approach to monitoring compliance with requirements for the use of PPE at nuclear power plant facilities was proposed. To train the model, an annotated dataset was formed, including 3808 images obtained both from webcams installed at an industrial facility and from open internet sources. A separate manually collected dataset was used for testing. In the framework of the study, two classes of objects were considered: a protective helmet and a full-face protective mask. To solve the detection problem, a one-stage YOLOv3 model with the basic Darknet-53 architecture was used. Training was performed in two stages: the first stage involved training with frozen convolutional layers of the basic network, and the second stage involved tweaking the model with unfreezing of all layers. To optimize the training process, the Adam optimizer algorithm with a batch size of 8 and different learning rates at individual stages was used. Based on experimental results, the developed model demonstrated an accuracy of 97.64% and a recall of 93.11%, processing a video stream in real time at 7.96 frames per second (FPS). However, the study is limited to monitoring only two types of PPE and does not consider a wider range of PPE. The authors note the need to expand the dataset and implement the developed model in real-world production conditions as future directions for development.&lt;/p&gt;&#13;
&lt;p&gt;In a study by Delhi V.S.K. et al. [6], the problem of using PPE on construction yards in real time mode was examined using deep learning and computer vision methods to improve industrial safety. To train the model, the authors manually compiled a dataset consisting of approximately 2 500 images collected from open internet sources. The images were divided into four classes: "SAFE," "UNSAFE," "no clothing" and "no jacket."&lt;/p&gt;&#13;
&lt;p&gt;The YOLOv3 architecture was used as the object detection model. To improve the robustness and generalization ability of the model, image augmentation methods were used during the data preparation stage, including horizontal flipping and rotation by up to 30°. The dataset was divided into training, validation, and test sets at a ratio of 90%, 8%, and 2%, respectively. Experimental results demonstrated the high efficiency of the proposed model then the mAP, Recall, and F1-score values reach 97%.&lt;/p&gt;&#13;
&lt;p&gt;Wang Z. et al. [7] investigated the problem of employee compliance with occupational safety requirements using deep learning neural network methods adapted for real-time mode object detection. The authors developed and compared detection models based on the YOLOv3, YOLOv4, and YOLOv5 architectures for recognizing six classes of PPE objects. The CHV dataset [8] containing 1 330 high-quality images was used for training and testing the models. The dataset included six classes of objects: protective helmets of four colors, employee face images and vests. A comparative analysis revealed that the YOLOv5x model achieved the highest detection accuracy, reaching 86.55%, while the YOLOv5s model demonstrated the highest processing speed reaching up to 52 frames per second. The obtained results indicate a tradeoff between detection accuracy and computational efficiency.&lt;/p&gt;&#13;
&lt;p&gt;Ma L. et al. [9] proposed a combined PPE detection algorithm using a portable version of the YOLOv4 model. A dataset containing approximately 25 000 images obtained from CCTV cameras installed at construction yards was generated to perform experiments. The images were unevenly distributed across six object classes and divided into training and test sets. Two object detection architectures were considered such as YOLOv4 and YOLOv4-Tiny. To improve the accuracy of the models, fine-tuning and optimization were performed using the original dataset. The best results were obtained using the CLSlim YOLOv4 model, based on the Channel and Layer Slimming (CLSlim) method. This approach contributed to reduce the number of model parameters by 98.2%, and to increase the output speed by 1.8 times, and, in the same time, to reduce the mAP value by only 2.1% compared to the baseline model. The results obtained demonstrate the high efficiency of the CLSlim method in reducing computational power and increasing the speed of object detection.&lt;/p&gt;&#13;
&lt;p&gt;In the study by Lo J.-H., Lin L.-K., and Hung C.-C. [10], the problem of developing deep learning algorithms for detecting the use of PPE in real time mode was considered. A dataset containing over 11 000 images was generated for training and evaluating the models. The capabilities of the YOLOv3, YOLOv4, and YOLOv7 architectures were investigated. To improve training quality and reduce the risk of overfitting, data preprocessing methods were used, including image flipping, cropping, noise addition, and color space transformation. A comparative analysis showed that the YOLOv7 model demonstrates the best results, achieving a mAP value of 97.29% with a processing speed (FPS) of 25.02 frames per second. The implemented system achieved a PPE detection accuracy of 97.5%, confirming the effectiveness of modern deep learning models for monitoring safety compliance at industrial facilities in real time mode. The authors note the need to further expand the dataset by including more PPE use scenarios, images with overlaying objects, and negative examples, as well as by improving data collection methods. It should be noted that most images in the dataset were captured under good lighting and favorable weather conditions, which may limit the generalization ability of the model and its effectiveness in more complex industrial conditions.&lt;/p&gt;&#13;
&lt;p&gt;In the study by Lee Y.-R et al. [11], a computer vision-based platform was proposed to monitor the proper use of PPE at construction yards. To compile the dataset, images with a resolution of 1280×720 pixels obtained from various sources, including Google Images, CCTV cameras and smartphones used at construction yards. The dataset included three classes of objects: “Person”, “Hard Hat”, and “Safety Vest”. A total of 1 031 images were used for the model training, and 257 images were used for testing. The dataset was used to train a contemporary pixel-based PPE detection model of the YOLACT architecture. The DeepSORT algorithm was used to track objects between frames. The results of the study show that the YOLACT model achieved an mAP50 score of 66.4%, and the DeepSORT algorithm provided an accuracy of 91.3% in detecting the PPE use status. As a future research direction, the authors consider expanding the platform's functionality by identifying potentially hazardous situations on construction sites through analyzing the co-relation between employees, the PPE they wear, and the surrounding work environment.&lt;/p&gt;&#13;
&lt;p&gt;Nugraha K.O.P.P. and Rifai A.P. [12] proposed an approach to monitoring compliance with PPE use requirements in industrial laboratories based on convolutional neural networks (CNN). The authors created their own dataset of images of laboratory employees containing various PPE wearing scenarios. A trained model classified the images with the aim of determination of the presence or absence of mandatory PPE items, including lab coats, safety glasses, and gloves. Experimental results showed that the developed approach provides acceptable accuracy in classifying elements under controlled operating conditions. However, the authors also found a significant dependence of classification quality on shooting conditions. In particular, changes in illumination and background characteristics leads to a noticeable decrease in recognition accuracy compared to the conditions presented in the training set. The obtained results demonstrate the existence of a domain shift problem inherent in PPE usage monitoring systems and highlight the need for more variable datasets covering a wider range of operating conditions, including different illumination levels, viewing angles, and environmental features.&lt;/p&gt;&#13;
&lt;p&gt;Zhao M. and Barati M. [13] proposed an approach to detecting PPE in electrical substations based on graph neural networks (GNNs). Unlike classical object detection methods, which consider objects independently, the proposed model takes into account spatial and semantic relationships between objects, including employees, process equipment, and PPE, representing them as a graph structure. This approach improves the robustness and speed of PPE detection in a specific industrial environment characterized by a high density of objects, a complex background, and the presence of visually similar elements.&lt;/p&gt;&#13;
&lt;p&gt;In the meantime, the authors note that the effectiveness of the proposed method largely depends on the quality and representativeness of the training data. Furthermore, constructing a graph structure requires additional labeling of the connections between scene objects, which increases the labour intensity of data preparation and complicates the implementation of such solutions in real-world industrial settings.&lt;/p&gt;&#13;
&lt;p&gt;Ngoc-Thoan N. et al. [14] developed an improved model based on the YOLOv5 architecture for preventing emergency situations and monitoring the use of PPE at construction yards. The authors modified the network architecture and training procedure to improve the accuracy of detecting PPE elements and potentially hazardous industrial situations. Conducted experiments demonstrate improved performance compared to the base version of YOLOv5 model. Despite these results, the authors note that the developed model remains sensitive to changes in environmental conditions, including variations in illumination, weather conditions, and scene background characteristics. These factors may limit its reliability during around-the-clock operation in real industrial conditions.&lt;/p&gt;&#13;
&lt;p&gt;Contemporary research is aimed at improving object detection architectures by eliminating anchoring mechanisms and improving the quality of data preparation. For example, Wang H. [15] investigated the applicability of the YOLOv8 model to PPE detection objectives. Based on the experimental results, the model demonstrates precision, recall, and F1-score values of 0.95, 0.84, and 0.74, respectively, outperforming models such as YOLOv5s, Faster R-CNN, SSD, and RetinaNet in detection quality.&lt;/p&gt;&#13;
&lt;p&gt;YOLOv8's high performance is due to its use of an anchor-free detection mechanism and an expanded dataset using augmentation methods. Eliminating fixed anchors leads to the reduction of the model's dependence on predefined object parameters and to increasing its robustness to changes in scale, illumination, and scene dynamics. The results obtained demonstrate the high potential of the YOLOv8 architecture for building automated monitoring systems for PPE use in industrial conditions.&lt;/p&gt;&#13;
&lt;p&gt;&lt;strong&gt;&lt;/strong&gt;&lt;/p&gt;&#13;
&lt;p&gt;&lt;strong&gt;Materials and methods&lt;/strong&gt;&lt;/p&gt;&#13;
&lt;p&gt;&lt;em&gt;Data collection and processing.&lt;/em&gt; A unique covering various scenarios and situations at production sites image dataset was formed for the study. For its formation, video materials obtained directly from the production sites were used. The initial sample included 300 video files, each of which was analyzed manually. Based on the results of the analysis, 91 of the highest quality video recordings were selected that met the requirements for training data. The selection took into account the absence of a black screen, sufficient lighting, the correct camera angle, the presence of activity in the frame and the minimum level of noise in the frame.&lt;/p&gt;&#13;
&lt;p&gt;Based on the selected video materials, a dataset was created, including 2 877 images containing 13 object classes: person, helmet, glasses, jacket, gloves, pants, boots, as well as negative classes for each element of PPE (helmet-negative, glasses-negative, jacket-negative, gloves-negative, pants-negative, boots-negative). The use of negative classes made it possible to reduce the number of false positives in cases where elements of everyday clothing are visually similar to personal protective equipment.&lt;/p&gt;&#13;
&lt;p&gt;&lt;em&gt;Data Augmentation.&lt;/em&gt; To increase the volume and diversity of the dataset, augmentation methods were used, which is an important stage in data preparation in computer vision tasks. The use of data augmentation is relevant in the context of PPE detection, as it improves the model's ability to handle real-world scenarios.&lt;/p&gt;&#13;
&lt;p&gt;During the augmentation process, geometric transformations were used, including rotation, scaling, and image translation, simulating changes in the position and orientation of objects typical of industrial settings. Horizontal image flipping was also applied, allowing the model to better generalize to PPE examples viewed from different angles. To simulate a variety of operating conditions, color corrections were performed on the images, including brightness and contrast adjustments. Blurring and noise addition methods were also applied to simulate real-world industrial conditions, where PPE may be partially hidden or exposed to external influences.&lt;/p&gt;&#13;
&lt;p&gt;As a result of the augmentation, the dataset size increased from 2 877 up to 16 352 images. This allows for significant increase the variability of the training set and improvement of the model's robustness to changes in the external environment and shooting conditions.&lt;/p&gt;&#13;
&lt;p&gt;&lt;em&gt;Object Detection.&lt;/em&gt; Various computer vision architectures were considered for object detection and instance segmentation tasks, including Faster R-CNN [16], SSD [17], and models from the YOLO family [18]. Among them, special attention was given to the YOLOv8 architecture, developed by Ultralytics, which is an advanced single-stage detector optimized for real-time operation. The YOLOv8 architecture uses a fully convolutional approach with the CSPDarknet backbone, ensuring efficient gradient propagation and increasing computational efficiency. Unlike previous versions of the YOLO family, the model implements an anchor-free detection mechanism, which allows to increase the detector's flexibility and to reduce computational complexity without significantly reducing accuracy.&lt;/p&gt;&#13;
&lt;p&gt;For segmentation, YOLOv8 predicts masks simultaneously with bounding rectangles, using deformable convolutions to improve spatial perception of objects. By virtue of pre-training on large-scale datasets such as COCO, the model demonstrates high generalization ability and performance when solving a wide range of object detection tasks.&lt;/p&gt;&#13;
&lt;p&gt;In this study, the YOLOv8 architecture was used to train and subsequently evaluate object detection and segmentation performance on the created dataset.&lt;/p&gt;&#13;
&lt;p&gt;&lt;em&gt;Human Pose Detection.&lt;/em&gt; Various deep learning architectures can be used to solve the problem of 2D human pose estimation, including ResNet [19], networks of Stacked Hourglass type [20], OpenPose [21], as well as methods based on Part Affinity Fields and Convolutional Pose Machines (CPMs) [22].&lt;/p&gt;&#13;
&lt;p&gt;In this study, the model proposed by Sun K. et al. [23], based on the High-Resolution Network (HRNet) architecture, was used to detect human body keypoints. HRNet uses a high-resolution subnet and repeatedly integrates representations formed at different resolutions, achieving high accuracy in pose estimation objective.&lt;/p&gt;&#13;
&lt;p&gt;To predict keypoint coordinates in images, the HRNet model generates high-resolution heatmaps of objects, modeling the true position of each point using a two-dimensional Gaussian distribution. On the COCO and MPII datasets, the HRNet model demonstrates high accuracy in human pose estimation. This study used a pre-trained version of HRNet trained on the COCO dataset.&lt;/p&gt;&#13;
&lt;p&gt;&lt;strong&gt;&lt;/strong&gt;&lt;/p&gt;&#13;
&lt;p&gt;&lt;strong&gt;Experiments&lt;/strong&gt;&lt;/p&gt;&#13;
&lt;p&gt;&lt;strong&gt;Model Training and Evaluation&lt;/strong&gt;&lt;/p&gt;&#13;
&lt;p&gt;&lt;em&gt;Training the Object Segmentation Model. &lt;/em&gt;To train the object segmentation model, the resulting dataset was divided into subsets to effectively utilize its potential. The total dataset size was 16 352 images, of which 13 081 were used for training the model, and the remaining 3 271 images were used for testing and validation. The distribution of objects by class is shown in Figure 1, allowing for a clear understanding of the dataset structure and its correspondence to real-world data.&lt;/p&gt;&#13;
&lt;p&gt;The model was trained over 300 epochs, providing the model with a sufficient number of iterations to train on diverse data and adaptation to a variety of conditions.&lt;/p&gt;&#13;
&lt;p&gt;&lt;em&gt;Model Evaluation.&lt;/em&gt; To evaluate the performance of the YOLOv8 model for solving the PPE detection problem, standard object detection quality metrics were used: Precision, Recall, and Mean Average Precision (mAP). These metrics allow us to evaluate the model's ability to correctly detect and localize PPE in images.&lt;/p&gt;&#13;
&lt;p&gt;Precision measures the accuracy of positive predictions made by the model. The formula for Precision is as follows:&lt;/p&gt;&#13;
&lt;p&gt;,&lt;/p&gt;&#13;
&lt;p&gt;where&lt;/p&gt;&#13;
&lt;p&gt;TP (True Positives) is the number of correct PPE detections. FP (False Positives) is the number of instances where the model incorrectly predicted the presence of PPE despite its absence.&lt;/p&gt;&#13;
&lt;p&gt;Recall, also known as sensitivity or true positive rate, measures the model's ability to detect all relevant PPE items. The formula for Recall is as follows:&lt;/p&gt;&#13;
&lt;p&gt;,&lt;/p&gt;&#13;
&lt;p&gt;where&lt;/p&gt;&#13;
&lt;p&gt;FN (False Negatives) is the number of actual PPE items that the model failed to recognize.&lt;/p&gt;&#13;
&lt;p&gt;&lt;/p&gt;&#13;
&lt;p&gt;&lt;strong&gt;Figure 1 – Class Distribution in the Dataset&lt;/strong&gt;&lt;/p&gt;&#13;
&lt;p&gt;&lt;/p&gt;&#13;
&lt;p&gt;Mean Average Precision (mAP) is one of the main metrics used to evaluate the quality of object detection models. This metric calculates the average accuracy at different confidence thresholds and is often visualized as a Precision-Recall Curve. Calculating mAP involves several steps:&lt;/p&gt;&#13;
&lt;ul&gt;&#13;
&lt;li&gt;To calculate Precision and Recall for different confidence thresholds.&lt;/li&gt;&#13;
&lt;li&gt;To calculate the area under the Precision-Recall curve.&lt;/li&gt;&#13;
&lt;li&gt;To average the areas under the curve for different classes, which yields the mAP value.&lt;/li&gt;&#13;
&lt;/ul&gt;&#13;
&lt;p&gt;The mAP provides a comprehensive picture of model performance at different confidence levels in detecting PPE objects. A higher mAP value indicates a more robust and accurate model.&lt;/p&gt;&#13;
&lt;p&gt;The analysis of these metrics data allows to identify the strengths and weaknesses of the model, enabling adjustments to improve detection accuracy. To improve PPE detection accuracy, several key hyperparameters must be considered as follows:&lt;/p&gt;&#13;
&lt;p&gt;Learning rate is a critical parameter that affects model convergence. The optimal value is selected experimentally, starting with a low value and gradually increasing it if necessary.&lt;/p&gt;&#13;
&lt;p&gt;Batch size affects stability and training speed. Larger batches speed up training but require more memory, so it is important to find a balance between efficiency and available resources.&lt;/p&gt;&#13;
&lt;p&gt;Data augmentation involves the methods such as rotation, shear, brightness and contrast adjustments. improve the model's ability to generalize and adapt to real-world conditions.&lt;/p&gt;&#13;
&lt;p&gt;Dataset balance ensures an even distribution of images from different classes (PPE and negative examples) to prevent the model from being biased toward a particular category.&lt;/p&gt;&#13;
&lt;p&gt;A validation dataset was used to monitor training quality, playing a key role in assessing the model's generalization ability and identifying potential overfitting. The validation set was designed to cover various production scenarios, camera angles, and lighting conditions, ensuring its representativeness of real-world operating conditions. During validation, hyperparameters such as learning rate, batch size, and augmentation parameters were adjusted to determine the optimal configuration for the best object detection performance.&lt;/p&gt;&#13;
&lt;p&gt;After completing the training and validation phases, the model was tested on a test dataset not previously used in the training process. The testing module evaluates the performance of the YOLOv8 model and verifies its accuracy in detecting PPE. For accurate evaluation, the test set must be diverse and cover potential operational scenarios in which the model will be used. The key testing metrics are Precision, Recall, and F1 score, which reflect the quality of predictions and the confidence of the model.&lt;/p&gt;&#13;
&lt;p&gt;Therefore, when creating a YOLOv8 model, training parameters such as learning rate, batch size, and data augmentation methods must be considered. Balancing the dataset and ensuring diversity in the validation and testing datasets are crucial for accurate model evaluation. These recommendations should be adjusted to the specific requirements for improving PPE detection accuracy, ensuring the effective performance of the model.&lt;/p&gt;&#13;
&lt;p&gt;For the further evaluation of the model's performance, the Precision-Confidence Curve was analyzed. This curve allows for evaluating the range of object detection accuracy changing with different model confidence thresholds. As shown in Figure 2, the model achieves an accuracy of approximately 0.98 for all classes, demonstrating high prediction reliability with a minimal number of false positives. It is of great importance in industrial conditions where strict compliance with safety standards is required.&lt;/p&gt;&#13;
&lt;p&gt;The Recall-Confidence Curve (refer Figure 3) shows the relationship between Recall and model confidence. High Recall values, reaching 0.97, indicate the model's ability to effectively detect virtually all types of PPE, reducing the probability of missed events. This confirms the reliability of YOLOv8 for PPE detection tasks in real-world conditions.&lt;/p&gt;&#13;
&lt;p&gt;&lt;/p&gt;&#13;
&lt;p&gt;&lt;/p&gt;&#13;
&lt;p&gt;&lt;strong&gt;Figure 2 – Precision-Confidence Curve based on bounding rectangles&lt;/strong&gt;&lt;/p&gt;&#13;
&lt;p&gt;&lt;/p&gt;&#13;
&lt;p&gt;&lt;/p&gt;&#13;
&lt;p&gt;&lt;strong&gt;Figure 3 – Recall-Confidence Curve based on bounding rectangles&lt;/strong&gt;&lt;/p&gt;&#13;
&lt;p&gt;&lt;/p&gt;&#13;
&lt;p&gt;An analysis of the Recall curves for individual classes reveals heterogeneity in the model's performance. Of particular interest is the "person" class, for which recall begins to decline sharply at a confidence threshold of approximately 0.4, while for most PPE classes, a steady decline is observed only at values above 0.7. This behavior is explained by the high intra-class variability of objects in the "person" class, due to the variety of poses, viewing angles, object scales, and the presence of partial overlapping of employees in the frame. As a result, the model assigns lower confidence values to correct human detections.&lt;/p&gt;&#13;
&lt;p&gt;For the comprehensive assessment of the quality of PPE detection, the Precision-Recall Curve was analyzed. Figure 4 illustrates the tradeoff between Precision and Recall, allowing for a detailed assessment of the model's performance. YOLOv8 demonstrates an average accuracy of mAP@0.5 equal to 0.964 across all object classes, demonstrating the model's high performance and reducing the risk of missing PPE items, contributing to improved safety in industrial conditions.&lt;/p&gt;&#13;
&lt;p&gt;&lt;/p&gt;&#13;
&lt;p&gt;&lt;/p&gt;&#13;
&lt;p&gt;&lt;strong&gt;Figure 4 – Precision-Recall curve for bounding rectangles&lt;/strong&gt;&lt;/p&gt;&#13;
&lt;p&gt;&lt;/p&gt;&#13;
&lt;p&gt;An analysis of the class-specific mean accuracy (mAP@0.5) values reveals significant variation: from 0.990 for the "person" class and 0.992 for "safety glasses" to 0.867 for "non-safety glasses" and 0.942 for "boots." The lowest value for the negative "non-safety glasses" class is conditional upon the visual proximity of regular eyeglasses and other facial features to safety glasses, making it difficult to distinguish between the positive and negative classes. The relatively low value for the "boots" class is conditional upon to the small size of the object and the frequent overlap of shoes at the bottom of the frame when viewed from a wide-angle camera angle. These observations are consistent with the imbalanced distribution of classes in the dataset (Figure 1) and confirm the need for targeted expansion of the training set for poorly represented and visually ambiguous categories.&lt;/p&gt;&#13;
&lt;p&gt;For the further evaluation of the model's quality, the F1-score dependance on confidence level curve (F1 Confidence Curve) was analyzed, allowing us to assess the balance between Precision and Recall of YOLOv8. Figure 5 demonstrates the influence of the confidence threshold on the F1-score value of the object detection quality metric. YOLOv8 achieves an F1-score of approximately 0.94 at the optimal confidence threshold for all PPE classes.&lt;/p&gt;&#13;
&lt;p&gt;&lt;/p&gt;&#13;
&lt;p&gt;&lt;/p&gt;&#13;
&lt;p&gt;&lt;strong&gt;Figure 5 – F1-score curve, confidence for bounding rectangles &lt;/strong&gt;&lt;/p&gt;&#13;
&lt;p&gt;&lt;/p&gt;&#13;
&lt;p&gt;The F1-score curve clearly shows that the optimal balance between Precision and Recall for all classes is achieved at a confidence threshold of approximately 0.425, resulting in an F1 value of 0.94. However, the "human" class demonstrates lower F1-score values (a plateau at approximately 0.8) across the entire threshold range compared to the PPE classes, which is consistent with the recall curve analysis.&lt;/p&gt;&#13;
&lt;p&gt;The model's results are presented in Figure 6.&lt;/p&gt;&#13;
&lt;p&gt;&lt;/p&gt;&#13;
&lt;p&gt;&lt;/p&gt;&#13;
&lt;p&gt;&lt;strong&gt;Figure 6 – Model result on test data&lt;/strong&gt;&lt;/p&gt;&#13;
&lt;p&gt;&lt;strong&gt;&lt;/strong&gt;&lt;/p&gt;&#13;
&lt;p&gt;&lt;strong&gt;Development of a PPE Monitoring System&lt;/strong&gt;&lt;/p&gt;&#13;
&lt;p&gt;To ensure continuous analysis of the video stream in real time, a dual-threaded data processing architecture is implemented in the system.&lt;/p&gt;&#13;
&lt;p&gt;The first thread (Capture Thread) asynchronously reads frames from the video stream and stores them in a buffer for subsequent processing. This approach minimizes delays in data receipt and ensures the transmission of the most relevant frames.&lt;/p&gt;&#13;
&lt;p&gt;The second thread (Processing Thread) is responsible for object detection and subsequent semantic analysis. It extracts frames from the first thread and processes them independently of the capture process. Splitting the functions into two threads reduces processing delays and ensures the relevance of the analyzed frames even when the computation time exceeds the interval between new images received.&lt;/p&gt;&#13;
&lt;p&gt;This approach prevents data transmission delays and efficiently utilizes computing resources (CPU/GPU).&lt;/p&gt;&#13;
&lt;p&gt;At the first stage of video stream processing a people detection model based on YOLOv8 and the BOTSORT tracker is used. These models effectively identify people in the frame and track their movements. Next, a pose estimation is performed for each detected person using the HRNet model. This model uses frame images and the coordinates of the person's bounding box (bbox) as input. The model can identify 17 key points of the human skeleton with high accuracy, which is especially important when multiple people are present in the frame.&lt;/p&gt;&#13;
&lt;p&gt;Following the pose estimation, the person's bounding box is expanded by N pixels, after which the corresponding region of the image is cropped, and the cropped region is being uploaded to the object detection model. This approach allows for focused analysis on the employee location and eliminates the influence of the surrounding background. The model is trained to recognize 13 object classes: one class corresponds to the person, and the remaining 12 classes correspond to PPE items and their negative classes (helmets, goggles, gloves, etc.). Experiments have shown that using cropped images significantly improves the detection of small objects. For example, for the "gloves" class, 58 detections were obtained in cropped images, compared to only 25 in the full frame. These results indicate that image cropping enhances the accuracy of small object detection, improving the overall performance of the system.&lt;/p&gt;&#13;
&lt;p&gt;To reduce the number of false positives, a two-stage mechanism for matching detected objects to a specific employee was implemented. The first stage uses spatial matching, which matches the coordinates of detected PPE objects with the bounding rectangle of the person based on the intersection of their regions. The second stage uses anatomical matching, which verifies that the PPE matches the pose keypoints. For example, a helmet should be positioned near the head, and gloves near the hands.&lt;/p&gt;&#13;
&lt;p&gt;The analysis results are stored in a "person" class object, which contains the following information: frames, a list of detected PPEs with confidence levels, a pose keypoint vector, and a tracking ID.&lt;/p&gt;&#13;
&lt;p&gt;To improve decision stability, the system accumulates data for each person over M-consecutive frames (M = 24 at 8 FPS, which corresponds to a three-second video fragment). PPE presence classification is performed using the majority rule: if more than 50% of the frames in the interval contain a negative class (absence of an object), an event is generated. To reduce the impact of noise and random detection errors, a model confidence threshold exceeding 0.85 is additionally used for negative classes.&lt;/p&gt;&#13;
&lt;p&gt;To prevent repeated registrations of the same violation, a time filter has been implemented. This filter uses a hash table of events with a set time-to-live (TTL) of 5 minutes. Each event is encoded with a combination of parameters (track ID, PPE type), preventing repeated registrations of the same violation within a specified time interval.&lt;/p&gt;&#13;
&lt;p&gt;When a violation is detected in asynchronous mode, a GIF animation consisting of 24 consecutive frames is generated and then saved to the database via a separate record thread. This approach eliminates delays in the main thread associated with animation generation, maintaining high system performance.&lt;/p&gt;&#13;
&lt;p&gt;Experiments have shown that generating GIF files is a resource-intensive operation and can reduce video stream processing speed. Therefore, the system implements a dynamic buffer clearing mechanism after object lifetimes, preventing the accumulation of out-dated data and optimizing RAM usage.&lt;/p&gt;&#13;
&lt;p&gt;To improve system performance, TensorRT model quantization methods are used to accelerate GPU inference and an asynchronous frame preloading algorithm. RAM consumption is controlled using an object manager, which removes excess instances of the Person class if their number exceeds a specified threshold (5 instances). This approach enables efficient resource management and ensures uninterrupted system operation even during lengthy video stream analysis.&lt;/p&gt;&#13;
&lt;p&gt;&lt;/p&gt;&#13;
&lt;p&gt;&lt;strong&gt;Results and Discussion&lt;/strong&gt;&lt;/p&gt;&#13;
&lt;p&gt;This section presents the results of testing the PPE compliance monitoring system using 16 video files obtained in real-world production conditions. The system uses two models: a PPE segmentation and detection model based on YOLOv8 and a human pose estimation model using HRNet, which significantly improves the accuracy of determining whether PPE is being worn correctly.&lt;/p&gt;&#13;
&lt;p&gt;The following metrics were used to evaluate the model's performance: True Positives (TP) is equal to 113, True Negatives (TN) is equal to 0, False Positives (FP) is equal to 4, and False Negatives (FN) is equal to 52.&lt;/p&gt;&#13;
&lt;p&gt;Based on these values, key metrics were calculated:&lt;/p&gt;&#13;
&lt;ul&gt;&#13;
&lt;li&gt;Recall comes up to 0.6848. This metric characterizes the system's ability to detect actual violations, amounting to 68.48%. This indicates that the model is not always able to identify all violations.&lt;/li&gt;&#13;
&lt;li&gt;Precision comes up to 0.9658. A high value for this indicator indicates that 96.58% of the detected violations were correct.&lt;/li&gt;&#13;
&lt;li&gt;F1-score comes up to 0.8014. This metric is an integral indicator of the system's quality, as it takes into account both precision and recall.&lt;/li&gt;&#13;
&lt;li&gt;Accuracy comes up to 0.6686. This value indicates that 66.86% of the analyzed cases were correctly classified. However, this metric has limited informativeness for the task at hand, as the test sample did not contain any true negatives (TN = 0). This is due to the specifics of the dataset's formation: the "no violation" class was not included in the video tagging, as it has no practical significance in the context of the task. The model focuses exclusively on identifying violations, so only cases where they were present were considered during the tagging process. Consequently, no true negatives were recorded in the test videos.&lt;/li&gt;&#13;
&lt;/ul&gt;&#13;
&lt;p&gt;The system demonstrated an F1 score of 0.8014, reflecting a balanced balance between high Precision (96.58%) and Recall (68.48%), as evidenced by a low false positive rate (FP = 4). This indicates a good ability of the model to detect violations. However, Recall (0.6848) indicates the presence of missed events (FN = 52). This suggests the need to further improve the system's sensitivity to reduce the number of undetected events.&lt;/p&gt;&#13;
&lt;p&gt;To improve the completeness of violation detection in real-world conditions, it is necessary to optimize frame processing algorithms, improve models for more accurate violation detection, and expand the training set.&lt;/p&gt;&#13;
&lt;p&gt;&lt;em&gt;Training vs. Real-World Gap.&lt;/em&gt; It is worth emphasizing the difference between the conditions under which the model is evaluated during training and during real-world use. Model training metrics (Precision = 0.98, Recall = 0.97) were obtained on a static test dataset containing 3,271 images with annotated objects, where detection quality was assessed exclusively under controlled conditions. The system training metrics (Precision = 0.9658, Recall = 0.6848) reflect the performance of the full monitoring system, which includes not only the detection model but also additional data processing stages: object tracking (BOTSORT), temporal aggregation (24 frames in 3 seconds), threshold filtering (confidence  0.85), anatomical matching, and a temporal filter for repeated events (TTL = 5 minutes). Each of these components can be a source of missed violations. Thus, the recall of 68.48% characterizes not an isolated detection model, but the full chain from video stream to event generation.&lt;/p&gt;&#13;
&lt;p&gt;&lt;em&gt;Comparative analysis with existing approaches.&lt;/em&gt; The obtained results should be considered in the context of previously published studies.&lt;/p&gt;&#13;
&lt;p&gt;&lt;/p&gt;&#13;
&lt;p&gt;Table 1. Comparative analysis with existing approaches&lt;/p&gt;&#13;
&lt;table&gt;&#13;
&lt;tbody&gt;&#13;
&lt;tr&gt;&#13;
&lt;td width="226"&gt;&#13;
&lt;p&gt;Research&lt;/p&gt;&#13;
&lt;/td&gt;&#13;
&lt;td width="132"&gt;&#13;
&lt;p&gt;Model&lt;/p&gt;&#13;
&lt;/td&gt;&#13;
&lt;td width="123"&gt;&#13;
&lt;p&gt;Number of classes&lt;/p&gt;&#13;
&lt;/td&gt;&#13;
&lt;td width="160"&gt;&#13;
&lt;p&gt;Main results&lt;/p&gt;&#13;
&lt;/td&gt;&#13;
&lt;/tr&gt;&#13;
&lt;tr&gt;&#13;
&lt;td width="226"&gt;&#13;
&lt;p&gt;DelhiV.S.K., SankarlalR. и ThomasA. [6]&lt;/p&gt;&#13;
&lt;/td&gt;&#13;
&lt;td width="132"&gt;&#13;
&lt;p&gt;YOLOv3&lt;/p&gt;&#13;
&lt;/td&gt;&#13;
&lt;td width="123"&gt;&#13;
&lt;p&gt;4&lt;/p&gt;&#13;
&lt;/td&gt;&#13;
&lt;td width="160"&gt;&#13;
&lt;p&gt;mAP = 97 %,&lt;/p&gt;&#13;
&lt;p&gt;Recall = 97 %,&lt;/p&gt;&#13;
&lt;p&gt;F1-score = 97 %&lt;/p&gt;&#13;
&lt;/td&gt;&#13;
&lt;/tr&gt;&#13;
&lt;tr&gt;&#13;
&lt;td width="226"&gt;&#13;
&lt;p&gt;Wang Z. et al. [7]&lt;/p&gt;&#13;
&lt;/td&gt;&#13;
&lt;td width="132"&gt;&#13;
&lt;p&gt;YOLOv5x&lt;/p&gt;&#13;
&lt;/td&gt;&#13;
&lt;td width="123"&gt;&#13;
&lt;p&gt;6&lt;/p&gt;&#13;
&lt;/td&gt;&#13;
&lt;td width="160"&gt;&#13;
&lt;p&gt;Accuracy = 86,55 %&lt;/p&gt;&#13;
&lt;/td&gt;&#13;
&lt;/tr&gt;&#13;
&lt;tr&gt;&#13;
&lt;td width="226"&gt;&#13;
&lt;p&gt;Lo J.-H., Lin L.-K., and Hung C.-C. [10]&lt;/p&gt;&#13;
&lt;/td&gt;&#13;
&lt;td width="132"&gt;&#13;
&lt;p&gt;YOLOv7&lt;/p&gt;&#13;
&lt;/td&gt;&#13;
&lt;td width="123"&gt;&#13;
&lt;p&gt;not specified&lt;/p&gt;&#13;
&lt;/td&gt;&#13;
&lt;td width="160"&gt;&#13;
&lt;p&gt;mAP = 97,29 %&lt;/p&gt;&#13;
&lt;/td&gt;&#13;
&lt;/tr&gt;&#13;
&lt;tr&gt;&#13;
&lt;td width="226"&gt;&#13;
&lt;p&gt;Wang H. [15]&lt;/p&gt;&#13;
&lt;/td&gt;&#13;
&lt;td width="132"&gt;&#13;
&lt;p&gt;YOLOv8&lt;/p&gt;&#13;
&lt;/td&gt;&#13;
&lt;td width="123"&gt;&#13;
&lt;p&gt;not specified&lt;/p&gt;&#13;
&lt;/td&gt;&#13;
&lt;td width="160"&gt;&#13;
&lt;p&gt;Precision = 0,95,&lt;/p&gt;&#13;
&lt;p&gt;Recall = 0,84,&lt;/p&gt;&#13;
&lt;p&gt;F1-score = 0,74&lt;/p&gt;&#13;
&lt;/td&gt;&#13;
&lt;/tr&gt;&#13;
&lt;tr&gt;&#13;
&lt;td width="226"&gt;&#13;
&lt;p&gt;Model training (this study)&lt;/p&gt;&#13;
&lt;/td&gt;&#13;
&lt;td width="132"&gt;&#13;
&lt;p&gt;YOLOv8&lt;/p&gt;&#13;
&lt;/td&gt;&#13;
&lt;td width="123"&gt;&#13;
&lt;p&gt;13 (including 6 negative classes)&lt;/p&gt;&#13;
&lt;/td&gt;&#13;
&lt;td width="160"&gt;&#13;
&lt;p&gt;Precision = 0,98,&lt;/p&gt;&#13;
&lt;p&gt;Recall = 0,97,&lt;/p&gt;&#13;
&lt;p&gt;F1-score = 0,94&lt;/p&gt;&#13;
&lt;/td&gt;&#13;
&lt;/tr&gt;&#13;
&lt;tr&gt;&#13;
&lt;td width="226"&gt;&#13;
&lt;p&gt;Digital monitoring of PPE (this study)&lt;/p&gt;&#13;
&lt;/td&gt;&#13;
&lt;td width="132"&gt;&#13;
&lt;p&gt;YOLOv8 + HRNet + BoT-SORT&lt;/p&gt;&#13;
&lt;/td&gt;&#13;
&lt;td width="123"&gt;&#13;
&lt;p&gt;13 (including 6 negative classes)&lt;/p&gt;&#13;
&lt;/td&gt;&#13;
&lt;td width="160"&gt;&#13;
&lt;p&gt;Precision = 0,9658,&lt;/p&gt;&#13;
&lt;p&gt;Recall = 0,6848,&lt;/p&gt;&#13;
&lt;p&gt;F1-score = 0,8014&lt;/p&gt;&#13;
&lt;/td&gt;&#13;
&lt;/tr&gt;&#13;
&lt;/tbody&gt;&#13;
&lt;/table&gt;&#13;
&lt;p&gt;&lt;/p&gt;&#13;
&lt;p&gt;As the table shows, during the training phase, the proposed YOLOv8 model demonstrates results comparable to the best solutions presented in the literature. However, when testing the monitoring system on real video recordings, Recall dropped to 68.48% while maintaining a high accuracy of 96.58%.&lt;/p&gt;&#13;
&lt;p&gt;This discrepancy is explained by several significant factors. Firstly, the proposed system detects 13 object classes, including six negative classes, while most of the reviewed studies are limited to recognizing 2-6 object categories. Secondly, the system was evaluated on a continuous video stream obtained in real-world oil and gas field operational conditions, which presents additional challenges such as motion blur, changing illumination, partial occlusion of workers, and the small size of individual PPE elements in the frame.&lt;/p&gt;&#13;
&lt;p&gt;&lt;em&gt;Contribution of Anatomical Matching &lt;/em&gt;The key difference between the proposed approach and works [5–7, 9–15] is the use of a two-stage matching mechanism, including spatial and anatomical matching. Most existing systems associate PPE with a worker based solely on the spatial intersection of bounding boxes (IoU), which can lead to false positives. For example, the system may mistakenly interpret a hard hat held by a worker as being correctly worn.&lt;/p&gt;&#13;
&lt;p&gt;The closest approach to the proposed approach is the study by Xiong R. and Tang P. [4], which also uses human pose estimation to determine attention zones. However, this approach is focused only on two types of PPE (hard hat and vest). The proposed system controls six types of PPE, including gloves and boots, which requires more precise determination of the positions of the hands and feet.&lt;/p&gt;&#13;
&lt;p&gt;Using the HRNet architecture, it is possible to identify 17 key points of the human skeleton with high accuracy, ensuring correct anatomical matching of PPE elements with corresponding body parts even when multiple employees are present in the frame.&lt;/p&gt;&#13;
&lt;p&gt;&lt;em&gt;Practical Relevance and Industrial Implementation.&lt;/em&gt; Despite the moderate Recall value, the system's high accuracy (96.58%) ensures its practical significance for industrial applications. During the training phase, the YOLOv8 model achieved an F1 value of 0.94, significantly exceeding the result reported by Wang H. [15] for a similar YOLOv8s model (F1 = 0.74). Furthermore, the developed monitoring system, when processing a real video stream, demonstrated an F1 value of 0.8014, which is also higher than the results of the aforementioned study, despite significantly more complex operating conditions (13 classes, video stream, anatomical matching).&lt;/p&gt;&#13;
&lt;p&gt;From a practical standpoint, high detection accuracy is of particular importance for industrial safety systems. Frequent false positives (unfounded warnings) lead to reduction of user confidence in the system and to ignoring the warnings by operators. Therefore, the achieved result, characterized by only four false positives (FP = 4) out of 117 positive predictions, demonstrates the system's high reliability and the accuracy of the generated violation notifications.&lt;/p&gt;&#13;
&lt;p&gt;An additional advantage of the developed solution is its two-tier video stream processing architecture, based on asynchronous frame reading and processing. The implemented approach ensures a performance of 8 frames per second, which is sufficient for monitoring production processes characterized by relatively low employee movement speeds.&lt;/p&gt;&#13;
&lt;p&gt;To improve computational efficiency, model quantization using TensorRT was used, which allows for reducing inference time and data processing delay. This solution is an important factor for the subsequent integration of the system into the existing video surveillance infrastructure of oil and gas fields and ensuring its operation in near-real time.&lt;/p&gt;&#13;
&lt;p&gt;&lt;em&gt;Limitations of the Research.&lt;/em&gt; Despite the obtained results, this study has several limitations. The system was tested on 16 video files obtained from a single oil field, which limits the generalizability of the results to other industrial facilities. The absence of a True Negatives class (TN = 0) is conditional upon to the nature of the dataset, which is designed solely for safety rules violation detection. Consequently, several quality indicators, such as Specificity and negative predictive value (NPV), cannot be calculated correctly. Furthermore, the dataset predominantly contains daylight scenes with a small number of workers in the frame (up to 5 people), which may limit the system's effectiveness at nighttime shooting, complex lighting, and dense employee occurrence.&lt;/p&gt;&#13;
&lt;p&gt;&lt;em&gt;Directions for Further Research.&lt;/em&gt; To improve the comprehensiveness of safety rules violation detection, the following is planned:&lt;/p&gt;&#13;
&lt;ol&gt;&#13;
&lt;li&gt;To expand the dataset to more than 30 000 images, including scenes with night lighting, various weather conditions, and data from multiple industrial sites.&lt;/li&gt;&#13;
&lt;li&gt;To explore adaptive mechanisms for selecting confidence thresholds instead of a fixed value (0.85), including the use of reinforcement learning methods for dynamic threshold adjustment.&lt;/li&gt;&#13;
&lt;li&gt;To conduct a comparative evaluation of modern object detection models, including YOLOv11, YOLOv26, and RT-DETR, to improve the performance of small PPE elements detection.&lt;/li&gt;&#13;
&lt;li&gt;To implement a feedback mechanism for further training the model based on accumulated false negatives (FNs).&lt;/li&gt;&#13;
&lt;li&gt;To conduct multi-center validation of the system on data from various industrial sites to confirm the generalizability of the approach.&lt;/li&gt;&#13;
&lt;/ol&gt;&#13;
&lt;p&gt;&lt;em&gt;&lt;/em&gt;&lt;/p&gt;&#13;
&lt;p&gt;&lt;strong&gt;Conclusions&lt;/strong&gt;&lt;/p&gt;&#13;
&lt;p&gt;This paper presents the system for monitoring compliance with PPE wearing standards in real time. The developed approach is based on a two-tier video stream processing architecture and combines the detection of 13 object classes using the YOLOv8 model, human pose estimation using HRNet, and a two-stage spatial and anatomical matching mechanism to verify proper PPE wearing.&lt;/p&gt;&#13;
&lt;p&gt;Testing results on 16 video files obtained from a real oil field yielded an F1 score of 0.8014, with a Precision of 96.58% and a Recall of 68.48%. The system's high accuracy with a low false positives rate (FP = 4) confirms its practical applicability for industrial safety applications, where false positives reduce employee confidence in the system and ignoring warnings, creating greater operational damage than individual system gaps.&lt;/p&gt;&#13;
&lt;p&gt;The key scientific and practical result of the study is the integration of anatomical matching based on human pose estimation, allowing the correlation of PPE elements with corresponding body parts. Unlike traditional approaches based solely on the spatial intersection of bounding rectangles, the proposed solution reduces the number of false positives and increases the reliability of safety compliance monitoring.&lt;/p&gt;&#13;
&lt;p&gt;The results obtained confirm the potential of combining deep learning, object tracking, and human pose analysis for automated monitoring of PPE compliance.&lt;/p&gt;</body><back><ref-list><ref id="B1"><label>1.</label><mixed-citation>1. Kelm A., Laußat L., Meins-Becker A., Platz D., et al., Mobile passive Radio Frequency Identification (RFID) portal for automated and rapid control of Personal Protective Equipment (PPE) on construction sites // Automation in Construction. 2013. Vol. 36, P. 38–52. doi: 10.1016/j.autcon.2013.08.009</mixed-citation></ref><ref id="B2"><label>2.</label><mixed-citation>2. Zhang H., Yan X., Li H., Jin R. and Fu H. Real-Time Alarming, Monitoring, and Locating for Non-Hard-Hat Use in Construction // Journal of Construction Engineering and Management. 2019. Vol. 145, P.1–13. doi: https://doi.org/10.1061/(ASCE)CO.1943-7862.000162</mixed-citation></ref><ref id="B3"><label>3.</label><mixed-citation>3. Dubey S. Dixit M. A comprehensive survey on human pose estimation approaches // Multimedia Systems. 2022. Vol. 29, P. 167–195. doi: 10.1007/s00530-022-00980-0</mixed-citation></ref><ref id="B4"><label>4.</label><mixed-citation>4. Xiong R., Tang P. Pose guided anchoring for detecting proper use of personal protective equipment // Automation in Construction. 2021. Vol.130, Article 103828. 17 p. doi: https://doi.org/10.1016/j.autcon.2021.103828.</mixed-citation></ref><ref id="B5"><label>5.</label><mixed-citation>5. Chen S. and Demachi K. A vision-based approach for ensuring proper use of personal protective equipment (PPE) in decommissioning of Fukushima Daiichi nuclear power station // Applied Sciences. 2020. Vol.10, Article 5129. 14 p. doi: https://doi.org/10.3390/app10155129.</mixed-citation></ref><ref id="B6"><label>6.</label><mixed-citation>6. Delhi V.S.K., Sankarlal R. and Thomas A. Detection of Personal Protective Equipment (PPE) Compliance on Construction Site Using Computer Vision Based Deep Learning Techniques // Frontiers in Built Environment. 2020. Vol. 6, Article 136. 10 p. doi: https://doi.org/10.3389/fbuil.2020.00136.</mixed-citation></ref><ref id="B7"><label>7.</label><mixed-citation>7. Wang Z., Wu Y., Yang L., Thirunavukarasu A., Evison C. and Zhao Y. Fast Personal Protective Equipment Detection for Real Construction Sites Using Deep Learning Approaches // Sensors. 2021. Vol. 21, Article 3478. 22 p. doi: https://doi.org/10.3390/s21103478.</mixed-citation></ref><ref id="B8"><label>8.</label><mixed-citation>8. GitHub [Internet]. ZijianWang Z. PPE_detection [cited 2026 Мау 5]. Available from https://github.com/ZijianWang-ZW/PPE_detection.</mixed-citation></ref><ref id="B9"><label>9.</label><mixed-citation>9. Ma L., Li X., Dai X., Guan Z., Lu Y. A Combined Detection Algorithm for Personal Protective Equipment Based on Lightweight YOLOv4 Model // Wireless Communications and Mobile Computing. 2022. Vol. 2022. Article ID 3574588. 11 p. doi: https://doi.org/10.1155/2022/3574588.</mixed-citation></ref><ref id="B10"><label>10.</label><mixed-citation>10. Lo J.-H., Lin L.-K., and Hung C.-C. Real-Time Personal Protective Equipment Compliance Detection Based on Deep Learning Algorithm // Sustainability. 2023. Vol. 15, Article 391. 15 p. doi: https://doi.org/10.3390/su15010391.</mixed-citation></ref><ref id="B11"><label>11.</label><mixed-citation>11. Lee Y.-R., Jung S.-H., Kang K.-S., Ryu H.-C., Ryu H.-G. Deep Learning-Based Framework for Monitoring Wearing Personal Protective Equipment on Construction Sites // Journal of Computational Design and Engineering. 2023. Vol. 10, P. 905–917. doi: https://doi.org/10.1093/jcde/qwad040.</mixed-citation></ref><ref id="B12"><label>12.</label><mixed-citation>12. Nugraha K.O.P.P., Rifai A.P. Convolutional Neural Network for Identification of Personal Protective Equipment Usage Compliance in Manufacturing Laboratory // Jurnal Ilmiah Teknik Industri. 2023. Vol. 22, P. 11–24. doi: 10.23917/jiti.v22i1.21826.</mixed-citation></ref><ref id="B13"><label>13.</label><mixed-citation>13. Zhao M., Barati M. Substation Safety Awareness Intelligent Model: Fast Personal Protective Equipment Detection Using GNN Approach // IEEE Transactions on Industry Applications. 2023. Vol. 59, P. 3142 - 3150. doi: 10.1109/TIA.2023.3234515</mixed-citation></ref><ref id="B14"><label>14.</label><mixed-citation>14. Ngoc-Thoan N., Bui D.-Q. T., Tran C.N.N., and Tran D.-H. Improved Detection Network Model Based on YOLOv5 for Warning Safety in Construction Sites // International Journal of Construction Management. 2023. Vol. 24, P. 1007-1017. doi: https://doi.org/10.1080/15623599.2023.2171836.</mixed-citation></ref><ref id="B15"><label>15.</label><mixed-citation>15. Wang H. Detection of Personal Protective Equipment (PPE) Using an Anchor Free-Convolutional Neural Network // International Journal of Advanced Computer Science and Applications. 2024. Vol. 15, 9 p. doi: 10.14569/IJACSA.2024.0150239</mixed-citation></ref><ref id="B16"><label>16.</label><mixed-citation>16. Ren S., He K., Girshick R., Sun J. Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks // IEEE Transactions on Pattern Analysis and Machine Intelligence. 2017. Vol. 39, P. 1137–1149. doi: 10.1109/TPAMI.2016.2577031.</mixed-citation></ref><ref id="B17"><label>17.</label><mixed-citation>17. Liu W., Anguelov D., Erhan D., Szegedy C., Reed S., Fu C.-Y., Berg A.C. SSD: Single Shot MultiBox Detector // Computer Vision. 2016. Lecture Notes in Computer Science. Vol. 9905, Springer. P. 21–37. doi: 10.1007/978-3-319-46448-0_2.</mixed-citation></ref><ref id="B18"><label>18.</label><mixed-citation>18. Redmon J., Divvala S., Girshick R., Farhadi A. You Only Look Once: Unified, Real-Time Object Detection // Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2016. P. 779–788. doi: 10.1109/CVPR.2016.91.</mixed-citation></ref><ref id="B19"><label>19.</label><mixed-citation>19. He K., Zhang X., Ren S., Sun J. Deep Residual Learning for Image Recognition // Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2016. P. 770–778. doi: https://doi.org/10.1109/CVPR.2016.90.</mixed-citation></ref><ref id="B20"><label>20.</label><mixed-citation>20. Newell A., Yang K., Deng J. Stacked Hourglass Networks for Human Pose Estimation // Computer Vision. 2016. Lecture Notes in Computer Science. Vol. 9912, Springer. P. 483–499. doi: https://doi.org/10.1007/978-3-319-46484-8_29.</mixed-citation></ref><ref id="B21"><label>21.</label><mixed-citation>21. GitHub [Internet]. CMU-Perceptual-Computing-Lab/openpose [cited 2026 Мау 5]. Available from https://github.com/CMU-Perceptual-Computing-Lab/openpose.</mixed-citation></ref><ref id="B22"><label>22.</label><mixed-citation>22. Wei S.-E., Ramakrishna V., Kanade T., Sheikh Y. Convolutional Pose Machines // IEEE Conference on Computer Vision and Pattern Recognition. 2016. P. 4724–4732. doi: 10.1109/CVPR.2016.511.</mixed-citation></ref><ref id="B23"><label>23.</label><mixed-citation>23. Sun K., Xiao B., Liu D., Wang J. Deep High-Resolution Representation Learning for Human Pose Estimation // IEEE Conference on Computer Vision and Pattern Recognition. 2019. P. 5686–5696. doi: 10.1109/CVPR.2019.00584.</mixed-citation></ref></ref-list></back></article>
