RESEARCH PROPOSAL Research title Tăng cường theo dõi người qua nhiều camera trực tuyến bằng phân (Vietnamese) cụm tự hiệu chỉnh (SCRC) cùng biểu diễn đặc trưng sử dụng Attention đa đầu tích chập (MDHA) và tích chập đồ thị động (DGC) trên không gian toàn cục. Research title Enhancing Online Multi-Camera People Tracking via Self(English) Correcting Re-Clustering (SCRC) and Feature Embedding with Multi-DConv-Head Attention (MDHA) and Dynamic Graph Convolution (DGC) on a Global Spatial Topology. Sub-committee Information Technology Group name BDs.M Authors CE190201 – Huynh Thanh Duy – Leader CE190355 – Tien Quoc Bao CE191210 – Le Ngoc Minh Supervisors Lam Le Huan & HoangTN Abstract Online Multi-Camera People Tracking (MCPT) remains a challenging problem due to inconsistent appearance across views, occlusions, and ID-switches. To address these issues, we propose an integrated framework that combines a novel clustering mechanism and an enhanced feature embedding architecture. Our contributions include: (1) a real-time detector based on YOLOv12 for accurate object localization; (2) a lightweight but expressive OSNet backbone augmented with Transformer blocks, comprising Multi-DConv-Head Attention (MDHA) and Dynamic Graph Convolution (DGC), strategically inserted after conv3 and conv4 layers. These blocks model both local-global attention and spatial topology-aware feature relations; (3) keypoint-based descriptors from RTMPose to increase robustness under partial occlusion; (4) a spatial topology-aware feature association mechanism guided by homographybased graph construction; and (5) a Self-Correcting Re-Clustering (SCRC) scheme that periodically refines global identity assignments. Our MDHA module expands lightweight attention with depthwise convolutions to enable scalable attention modeling over fine-grained spatial cues, while DGC incorporates a dynamic graph reasoning layer that constructs k-NN topologies directly from intermediate feature maps. With both being integrated into OSNet, our network becomes more discriminable without sacrificing inference speed. Experimental evaluation shows significant improvements over recent online tracking baselines such as BotSORT in both accuracy (HOTA, AssA, DetA) and real-time applicability. The system is well-suited for smart city surveillance in real-world scenarios like supermarkets, public transport, and indoor monitoring. Key words: Multi-Camera People Tracking (MCPT), Online Re-Identification (Re-ID), Self-Correcting Re-Clustering (SCRC), Enhanced OSNet, MDHA-augmented OSNet, Lightweight Transformer Attention, Multi-DConv-Head Attention (MDHA), Pose-aware Feature Embedding, Global Spatial Reasoning, Dynamic Graph Convolution (DGC), Graph Neural Network(GNN), Global Spatial Topology, YOLOv12, RTMPose, Smart Surveillance Systems. 1. Introduction 1.1 Literature review Multi-Camera People Tracking (MCPT) is a significant and challenging computer vision problem with wide practical application in tasks such as public surveillance, human behavior analysis, and advanced security networks. To address this, we introduce a dual-branch backbone combining MDHA and Dynamic Graph Convolution (DGC) within an OSNet structure to enhance spatiotemporal representation and improve robustness against view variations. A primary difficulty in MCPT is the accurate and consistent tracking of individuals across unrelated, overlapping camera views, particularly under complex conditions involving partial occlusions, appearance confusion, and frequent trajectory overlaps – an active area of research. The latest advances in the field have been towards enhancing tracking procedures to be more robust and precise by utilizing new object detection techniques and deep neural networks, specifically in the area of person re-identification (Re-ID). Figure 1. Example of Multi-Camera People Tracking: The center shows a 2D map with colored regions for each camera's view. Black dots mark people’s locations, red lines connect each person in the images to their respective red dot on the map. MCPT aims to track each person with a consistent ID across cameras. Object detection is the core of any MCPT system, and YOLO models are extremely popular following their effectiveness and speed in real-time processing. Earlier versions of YOLO, such as YOLOv10 and YOLOv11 [1], demonstrated a strong balance between detection accuracy and processing speed, making them suitable for real-time applications across multiple camera streams. These models paved the way for more advanced architectures by introducing concepts like anchor boxes and multi-scale predictions. The sequential evolution in the YOLO series proves the continued pursuit of faster and more accurate object detection in diverse computer vision applications. MCPT's success also relies on extracting quality and discriminative features for successful person re-identification. Traditional appearance-based methods are also less effective in crowded or dynamically changing environments. To avoid these limitations, recent research has increasingly sought to leverage pose keypoint extraction. High-resolution human pose information from models like RTMPose [2] greatly enriches the representative features for ReID. Utilizing pose estimation gives one encouraging avenue to obtain more discriminative features, particularly in instances of occlusion or high visual similarity among individuals where appearance cues alone may not prove effective [3]. Person Re-ID backbones such as OSNet [4] have demonstrated strong performance in learning omni-scale features. However, their reliance on purely convolutional streams can limit their ability to capture long-range dependencies and global context effectively. To address this, recent research has explored augmenting convolutional architectures with attention mechanisms and graph-based modeling. For instance, the integration of Transformer blocks [12, 13], which have shown success in capturing long-range dependencies [5, 6], has been investigated in various Re-ID models. Similarly, graph convolutional networks [14, 15] [20] have been explored to model relationships between different parts of the feature map, potentially enhancing robustness to occlusion and clutter by leveraging contextual information [7, 8]. Inspired by these advancements, this research investigates the potential benefits of incorporating both attention and graph-based mechanisms into the OSNet architecture to improve feature discriminativeness and robustness for online multi-camera people tracking. In addition to robust feature extraction for Re-ID, tracking also plays a crucial role in maintaining identity consistency. Techniques like BotSORT [5] have shown great online multiobject tracking capability by harnessing Kalman filtering-based motion prediction and appearance feature-based association matching. BotSORT has since become a popular and competitive baseline for evaluating novel Re-ID features and association techniques in multicamera tracking. Despite important advances in the field, maintaining coherent identity labels across overlapping camera views in online MCPT systems remains an open issue, particularly under frequent trajectory intersections and inherent visual ambiguities. Although clustering-based methods for tracklet merging and temporal identity linking exist, they are often unsuitable for highly dynamic, real-time environments and may not fully leverage the global spatial topology for continuous self-refinement. To address such limitations, subsequent studies have advanced more sophisticated methods of camera-based identity correlation. For example, Kim et al. [6] conceptualized a Cluster Self-Refinement module that allows for the enhancement of ID assignment and feature memory using homography transformation for spatial reconstruction between overlapping views and pose features via RTMPose. These developments form a strong foundation for the design of more effective MCPT systems, yet major identity association challenges remain unsolved in online settings. 1.2 The necessity of the research While huge progress has been achieved in online MCPT, strong and stable identity tracking across numerous overlapping cameras over intricate real-world scenes remains a challenge. Most existing online systems have difficulty with visual ambiguities, intersection of multiple trajectories, and occlusion [7, 8, 9], leading to identity drift and fragmentation when relying solely on appearance features [10]. Further, existing methods employing spatial [7, 11] or pose [3] information may not leverage the Global Spatial Topology as efficiently as possible for continuous, self-supervised identity association updates. In order to address these existing problems, we propose a fivefold improvement over traditional online MCPT systems: (1) We employ YOLOv12 as our detection module to achieve accurate, low-latency person detection in multiple views. (2) Instead of applying ResNet50IBN [19] as in FastReID and previous works like Kim et al., we use a better OSNet backbone. This one includes Transformer Blocks each being composed of MDHA and DGC added after conv3 and conv4 for the effective extraction of local and global feature dependencies. (3) RTMPose is used to extract high-precision pose features and keypoints, which contribute greatly to Re-ID under visual similarity or occlusion. (4) We integrate a homography-based spatial topology to guide inter-camera graph construction, allowing topology-aware identity association. (5) Lastly, we introduce a Self-Correcting Re-Clustering (SCRC) module that periodically re-clusters cluster-based identity assignments by online pruning of incorrect or redundant feature representations. By synergistically integrating these advancements, our new approach is anticipated to significantly improve the accuracy, robustness, and long-term consistency of real-world online multi-camera people tracking by addressing the drawbacks of current methods in detection, feature representation, identity association, and online refinement. 1.3 Feasibility of research Figure 2. Overview of our proposed framework The feasibility of this research is strongly augmented by the reality that there are advanced deep learning techniques and software that are available and waiting to be utilized. Firstly, the remarkable progress in robust object detection, represented flawlessly by YOLOv12, gives a solid platform for efficiently detecting individuals in different camera streams. Secondly, the maturity of pose estimation methods like RTMPose allows us to pull out distinctive and informative pose-based features. These are crucial for tackling changes in how people look and when they're partially hidden, ultimately boosting person Re-ID performance. Notably, this research beefs up the OSNet backbone by inserting Transformer Blocks – each packing MultiDConv-Head Attention (MDHA) and Dynamic Graph Convolution (DGC) as you can see in Figure 2 – right after the conv3 and conv4 stages in OSNet achitecture. These Transformer Blocks combine local and global visual information using depthwise-conv attention and graphbased information sharing, achieving improved and more robust embeddings for future Re-ID and tracking tasks. This facilitates the more efficient and accurate construction of tracklets with feature-rich embeddings from the outset. Thirdly, the idea of leveraging how cameras spatially relate to each other, possibly through the module that combines appearance-based similarities (using Cosine and Euclidean Distance) with spatially adjusted distances (via Homography) as shown in the framework (Figure 2), has proven its worth in similar challenges and offers a promising way to improve identity linking across different cameras. Lastly, the step-by-step design and improvement of a Self-Correcting Re-Clustering (SCRC) mechanism, along with integrating insightful Global Spatial Topology for feature embedding and clustering, is totally achievable through well-thought-out algorithm design and extensive experiments on relevant datasets. We can objectively compare our proposed system against today's best online tracking algorithms, like BotSORT, using standard metrics such as HOTA, AssA, and DetA. All these points, both individually and together, highlight the methodological and technical soundness of this research and its potential to significantly contribute to the field of online multi-camera people tracking. Figure 3. Euclidean Distance with Homography: the positions of three cameras (colored squares) and their fields of view (FOVs, transparent areas) in a 2D space. The labeled dots ("ID") represent the locations of the tracked objects. 2. Research Objectives ● Propose and outline an innovative Self-Correcting Re-Clustering method for dynamically improving the consistency and accuracy of identity mappings across online Multi-Camera People Tracking (MCPT) systems between overlapping camera views. This method supports continuous self-adaptation of identity mappings with time, especially in realistic occlusion and intersection scenarios. ● Discover and utilize the Global Spatial Topology of the camera network to provide informative contextual information, enhancing the discriminative capability of feature embeddings for person re-identification and enabling more robust cross-camera identity linking. ● Design and deploy an online MCPT system that most effectively combines the SelfCorrecting Re-Clustering mechanism with the incorporation of Global Spatial Topology into a customized OSNet model enriched with Transformer blocks, Multi-DConv-Head Attention (MDHA), and Dynamic Graph Convolution (DGC). The architecture builds knearest-neighbor graphs over the feature space to diffuse contextual relationships with EdgeConv, significantly improving the robustness of feature representations on multiviewpoints and dense scenes. ● Test the system performance extensively on applicable benchmark datasets with standard metrics such as AssA (Association Accuracy) and DetA (Detection Accuracy). Interestingly, the HOTA (Higher Order Tracking Accuracy) metric will be the key metric for assessing the system's ability to have consistent identity tracking in difficult multicamera environments. ● Describe and contrast the private and collaborative contributions of the Self-Correcting ReClustering mechanism, the Global Spatial Topology, and the enhanced OSNet architecture (Transformer + MDHA + DGC) to the end-to-end performance improvement of the online MCPT system with state-of-the-art techniques such as BotSORT. 3. Research Scope This research aims to create and evaluate a real-time Multi-Camera People Tracking (MCPT) system, and one of the main thrusts is in improving person re-identification (Re-ID) performance with a more efficient OSNet model by introducing Transformer blocks based on Multi-DConv-Head Attention (MDHA) and Dynamic Graph Convolution (DGC). The system will also weave in a Self-Correcting Re-Clustering mechanism and tap into the Global Spatial Topology of the camera network to achieve more accurate and consistent tracking of individuals across overlapping camera views. The scope of the study encompasses the following: • Training the person detector: We'll be using the well-established COCO dataset to train the YOLOv12 model, enabling it to effectively spot people in each camera's view. • Building our base Re-ID engine: This includes vigorous research, customization, and rigorous training of our OSNet model with Transformer blocks with Multi-DConv-Head Attention (MDHA) and Dynamic Graph Convolution (DGC), on the Market-1501 dataset. Here, attention lies in achieving extremely discriminative features towards recognizing individuals drastically from other cameras. Concretely, after each OSBlock, we construct a k-nearest-neighbor graph over the C-dimensional feature embeddings (using cosine similarity), and then apply a graph-convolution layer (EdgeConv) to aggregate information from neighboring nodes, and this result will be fed into the selfattention mechanism of the Transformer block. • Designing the Self-Correcting Re-Clustering (SCRC) procedure: We will implement and employ an intelligent algorithm to refine and update identity groupings (clusters) independently based on observed segments (tracklets) and how the scene is observed by multiple cameras. This method supports the continuous self-adaptation of identity mappings over time, especially in realistic occlusion and intersection scenarios. • Soliciting the Global Spatial Topology: We shall investigate and draw on methods for learning and exploiting the spatial setup of the cameras to create informative contextual information, enhancing the discriminative capability of feature embeddings for person re-identification and enabling more robust cross-camera identity linking. • System integration: This step involves integrating the YOLOv12 model (which we trained on COCO), our enhanced OSNet model with Transformer, MDHA, and DGC (which we trained on Market-1501), RTMPose (if we use it), and the spatial reasoning modules to produce a full online MCPT system. • Testing the system: We will test our system extensively on standard MCPT benchmarks (e.g., MOT17 and DukeMTMC) and observe how our whole system performs, with special focus on how our enhanced OSNet model with Transformer, MDHA, and DGC impacts key tracking metrics (HOTA, AssA, DetA) over a strong baseline like BotSORT. • Analyzing contributions: We'll break down the individual and collaborative contributions of the Self-Correcting Re-Clustering mechanism, the Global Spatial Topology, and the enhanced OSNet architecture (Transformer + MDHA + DGC) to the end-to-end performance improvement of the online MCPT system compared to state-ofthe-art techniques such as BotSORT. This research will primarily deal with calibrated video feeds from fixed cameras that don't overlap in their view. Factors like sudden changes in lighting, bad weather, or very crowded scenes could affect the results, and we'll discuss these potential limitations in our conclusions. 4. Approach and Method The research approach underpinning this research is designed methodically in a series of essential stages, with each stage directed towards assisting the development and evaluation of an effective online Multi-Camera People Tracking (MCPT) system. The stage-by-stage approach offers an in-depth investigation and the seamless integration of novel approaches to successfully overcome the inherent challenges in tracking individuals' identity across overlapping camera views. 4.1 Data Preparation and Feature Extraction We initially pre-process multi-camera observation video streams and discriminate downstream identity tracking. Person detection is performed with YOLOv12 that balances CSPDarknet and deformable convolutions to possess robust precision and low-latency detection over varying viewings. To calculate feature embeddings, we enhance OSNet's backbone by uniting Transformer blocks after conv3 and conv4. Figure 4. Overview of the enhanced OSNet backbone. Transformer blocks are inserted after conv3 and conv4 stages, each comprising Multi-DConv-Head Attention (MDHA) for localglobal channel interaction and Dynamic Graph Convolution (DGC) for adaptive topological reasoning. This design strengthens the model's representational power under complex MCPT scenarios. For every Transformer block: • Multi-DConv-Head Attention (MDHA): Taking cues from Primer [16], MDHA replaces standard self-attention with a convolutionally more economical variant. It employs a 1×1 qkv projection to get the query, key, and value maps, which are then subjected to depthwise 3×3 convolution individually. These are split into multi-head representations, normalized, and employed as input in a scaled dot-product attention across channels. This enables the model to learn fine-grained intra-channel dependences as well as localglobal context at the same time with minimal computational expense. • Dynamic Graph Convolution (DGC): − Inspired by EdgeConv [14] and , we treat each spatial location in the feature map as a graph node and dynamically construct a k-nearest-neighbor (k-NN) graph over the flattened feature space using cosine similarity. These neighborhoods are not based on physical proximity in image space but rather on semantic similarity in feature space, enabling the model to form connections between regions that are visually or contextually alike, regardless of their spatial distance. Alternative dynamic graph structures have also been explored in general vision tasks, such as in Manessi et al. [8], and could inspire future variants of our DGC formulation. − Each node (representing a pixel or patch) is connected to its k semantic neighbors, forming a semantic neighborhood graph. The edge features between node xix_ixi and neighbor 𝑥𝑖 are computed as ℎ(𝑥𝑖 , 𝑥𝑗 ) = 𝑀𝐿𝑃(𝑥𝑗 − 𝑥𝑖 ∥ 𝑥𝑖 ), combining both local geometry (through difference vectors 𝑥𝑗 − 𝑥𝑖 ) and semantic identity (via 𝑥𝑖 ). This hybrid formulation allows the model to encode both local structure and global context, in contrast to fixed-grid convolution which is purely spatial. − These edge features are then aggregated using max-pooling to produce updated node embeddings: 𝑥′𝑖 = max 𝑅𝑒𝐿𝑈(𝜃(𝑥𝑗 − 𝑥𝑖 ) + 𝜙(𝑥𝑖 )) j∈N(i) where 𝜃, 𝜙 are shared MLP weights, and 𝑁(𝑖) denotes the dynamically computed neighborhood. − This EdgeConv-style operation is repeated after both conv3 and conv4 layers, and the resulting topology-aware embeddings are merged with the output of the MDHA block via residual connections. The integration of DGC enhances the network's capacity to learn from both local geometric patches (i.e., neighborhood-based edges) and non-local semantic relationships (i.e., similarity-based graph links), ultimately improving robustness under occlusion, motion blur, and cross-view variance. Let X∈RB×C×H×W be the input feature maps. We reshape it into X′∈RB×N×C, compute k-NN graphs 𝐺 = (𝑉, 𝐸), and perform aggregation via: ℎ′𝑖 = ∑𝑗∈𝑁(𝑖) 𝑀𝐿𝑃(ℎ𝑖 ∥ ℎ𝑗 − ℎ𝑖 ) where 𝑁(𝑖) denotes neighbors of node 𝑖. This output is reshaped and fused with the MDHA block via residual connection. MDHA and DGC work together to introduce semantic coherence and spatial consistency into the feature embeddings crucial for maintaining person identity under occlusion and various viewpoints. This graph‑based aggregation propagates contextual cues across spatial locations, further sharpening our feature descriptors for re‑ID under multi‑camera conditions. Furthermore, we'll explore leveraging RTMPose [2], pre-trained for human pose estimation, to capture highly distinctive pose-based features. This is particularly aimed at tackling the complexities arising from people being partially hidden (occlusion) and looking very similar, issues we touched upon in the introduction. These appearance-based and potentially pose-based features will then serve as the bedrock for our subsequent identity association efforts. 4.2 Single-Camera People Tracking (SCPT) For each individual camera stream, we perform intra-camera tracking using the extracted Transformer-enhanced feature embeddings. Each detected person is encoded via OSNet + MDHA + DGC to generate a compact descriptor. These descriptors are matched over consecutive frames using appearance similarity (cosine distance), optionally integrated with motion cues via Kalman filtering. The outputs of this stage are short but stable tracklets temporally linked bounding boxes with consistent intra-camera identities. These form the basis for higher-level cross-camera association and global identity refinement. 4.3 Cross-Camera Identity Association (CCIA) and Clustering To resolve identities across overlapping views, we construct cross-camera tracklet graphs. For each pair of tracklets, we pool the feature vectors from their MDHA+DGC-enhanced embeddings, and compute pairwise similarities using cosine and Euclidean distance metrics. We incorporate homography-based spatial constraints by transforming each detection's pixel coordinates into a shared world coordinate frame. These 2D projections enable measurement of inter-camera proximity, which is fused with appearance similarity to guide association. During this stage, we reapply the DGC module on the pooled features, now within a graph constructed over both feature similarity and spatial proximity. The topology-aware graph structure 𝐺𝑐𝑐𝑖𝑎 allows the system to model both appearance continuity and physical adjacency across camera views: ℎ′ 𝑖 = ∑𝑗∈𝑁(𝑖) 𝑀𝐿𝑃(ℎ𝑗 − ℎ𝑖 ) where 𝑁𝐺𝑆𝑇 (𝑖) incorporates both visual and spatial proximity under Global Spatial Topology (GST). We then apply agglomerative clustering on this fused representation space to group tracklets belonging to the same individual. This dual-constraint clustering ensures that only tracklets that are both visually similar and spatially feasible are linked, substantially reducing false merges across unrelated views. 4.4 Self-Correcting Re-Clustering Internet-based MCPT systems typically suffer from ID drift and fragmentation due to noisy detections, occlusions, or visual ambiguity. As a countermeasure to this, we introduce here a Self-Correcting Re-Clustering (SCRC) mechanism that re-inspects and refines global identity assignments from time to time. SCRC consists of two phases: (1)Appearance-Based Purification: We use unsupervised clustering on buffer features for a given global ID. Outlier features that belong to inconsistent clusters (e.g., spurious records from ID-switches) are pruned, tightening up identity representation. (2) Overlap-Based Merging: With the aid of homography-projected spatial topology, we analyze candidate global ID pairs for overlap in space and appearance. When their trajectory projections and visual embeddings are close enough, we merge them, preserving the more temporally coherent ID. By continuous application of these two corrections throughout streaming operation, SCRC compensates for long-term drift and resolves redundant identities. This feedback mechanism maintains the system stable under continuous occlusion, re-entry, or viewpoint change. 4.5 Output Multi-Camera People Tracking In the final stage of the pipeline, our system combines all of the filtered tracklets into consistent global identities. Each one is assigned a unique global ID that is the same across all of the cameras, even in cases of temporary occlusion or exit-reentry disappearance. The output is a real-time stream of tracked identities across a multi-camera topology, where each detection is semantically linked via both appearance embeddings and spatial topology. The system delivers smooth transitions of identities across views while ensuring robustness against identity drift, occlusion, and trajectory overlap. 4.6 Validation and Evaluation Systematic evaluation with vast general sets of datasets, for example, Market 1501 and DukeMTMC, is offered in our work outline. We will utilize conventional metrics for multiobject tracking, for example, HOTA (Higher Order Tracking Accuracy), AssA (Association Accuracy), and DetA (Detection Accuracy). Our best measure will be HOTA as our primary metric for assessing overall tracking performance, but with particular interest in how uniformly and reliably our identity links hold across cameras – a primary research objective. We will compare the performance of our system against leading cutting-edge online MCPT methods like BotSORT to demonstrate how powerful our Self-Correcting Re-Clustering mechanism and leveraging Global Spatial Topology along with the enhanced OSNet+MDHA feature embedding, including MDHA layers after conv3 and conv4, truly are, as we highlighted in the abstract. While our system emphasizes global feature consistency and re-clustering, recent occlusion-aware tracking models such as Nasseri et al. [12] also offer potential enhancements under dense or partially observable scenes. 4.7 Analysis of Contributions We conduct an ablation study to assess the individual and synergistic contributions of the three major components of our system: a) Self-Correcting Re-Clustering (SCRC): • Without SCRC, identity fragmentation occurs more frequently due to cumulative feature drift in online settings. • When activated, SCRC reduces ID-switches and consolidates split trajectories using cluster-level refinement. b) Global Spatial Topology (GST): • Enables the system to embed physical spatial relations between camera views. • Improves the accuracy of cross-camera identity associations, especially in dense scenes with overlapping FOVs. c) Enhanced OSNet Backbone (w/ MDHA + DGC): • Provides stronger local-global feature representations through MDHA. • Reinforces topological structure using DGC, yielding better generalization under occlusion and appearance ambiguity. Through controlled removal of each module, we observe dramatic drops in HOTA and AssA, confirming their critical roles. The full system, with all three modules together, outperforms strong baselines like BotSORT by huge margins on all major tracking metrics. 4.8 Implementation and Real-World Testing Our proposed MCPT system is designed for deployment in real-world surveillance use cases such as smart city CCTV networks, malls, and transportation hubs. We evaluate our implementation under practical settings, including: High density of crowds, Heavy occlusions, Disjoint or overlapping camera locations, Limited computing resources (e.g., edge devices) Figure 5. Preliminary results of the proposed system in a real-world warehouse scenario. The model successfully detects and tracks multiple individuals with consistent IDs, even under occlusion and varying viewpoints, leveraging YOLOv12 + RTMPose + enhanced OSNet with Transformer blocks. Each module (YOLOv12 detector, OSNet+Transformer encoder, RTMPose, SCRC) is inference-performance- and modularity-optimized. The system runs in real time (>30 FPS) on a single GPU setup and horizontally scale to numerous streams. Field trials demonstrate our approach achieves robust identity continuity and high association accuracy even in challenging environments. Homography-aware graph construction and dynamic clustering correction ensure reliable performance with long-term deployments. The next generations will incorporate trajectory prediction and behavior understanding modules to further increase identity coherence and expand the scope to group activity monitoring. 5. Research plan No. Date Task Output Person in charge 1 25/04/2025 – Dataset Tune the existing YOLOv12 COCO-trained Minh, Bảo 05/05/2025 Preparation & person detector. Normalize, sanitize, and 2 06/05/2025 – 18/05/2025 3 19/05/2025 – 30/05/2025 expand datasets like Market-1501 and DukeMTMC for training Re-ID. Reduce model complexity, tweak parameters, Minh, Duy and improve cross-view recognition accuracy of identities. Improve the existing Re-Clustering mechanism Duy, Bảo in order to realize faster and more dynamic clustering of identities, especially during occlusion and intersection. Improve modeling of spatial topology between Bảo, Minh cameras to improve contextual awareness and cross-camera identity association. 5 13/06/2025 – Integrate all modules of a comprehensive Minh, Duy 25/06/2025 online MCPT system. Finish communication among components and debug the system for easy deployment. 6 26/06/2025 – Deploy and test the system on various realistic Bảo, Duy 15/07/2025 datasets and self-recorded video streams. Optimize model speed, decrease latency, and maintain edge device accuracy. 7 16/07/2025 – Contribution Replace SCRC individual and combined effort Bảo, Duy, 27/07/2025 Analysis & of Spatial Topology and enhanced OSNet Minh SOTA model. Benchmark end-to-end system Comparison performance against state-of-the-art (SOTA) methods like BotSORT with major measures like HOTA, AssA, and DetA. 6. Expected results This effort aims to produce an extremely reliable online Multi-Camera People Tracking (MCPT) system for optimal use in real-world smart city public space and CCTV deployments. Our enhanced OSNet model (by employing MDHA, DGC, and Global Spatial Topology) should yield robust real-time tracking on moving surveillance. Secondly, our integration of our Self-Correcting Re-Clustering with OSNet+MDHA+DGC's feature enhancement on the MOT17 and DukeMTMC benchmarks will be capable of significantly improving cross-camera person identity accuracy. We expect this will significantly enhance the BotSORT tracking system's HOTA score, as well as AssA and DetA, reducing spurious identity assignments. The robust appearance of components of the system from OSNet+MDHA+DGC should be endowed with robust performance in realistic occlusions and heavy scene situations. Smarter city-like simulation will test its workability and scale-up practicability for real-world large-scale deployment. Although preliminary results appear in the appendix, the actual paper will entail rigorous experimentation over more datasets for verifying generalizability, other optimized tunings for real-time feasibility, and exploration for hybridizing trajectory prediction or behavior analysis for enhanced identity coherence. Lastly, we aim to show the practicality of this MCPT system in the real world and its value to society through its applications in public safety and security enhancement, with the effect of each new feature contribution clearly demarcated. 4 31/05/2025 – 12/06/2025 Detector Tuning Re-ID Model Optimization (OSNet enhancement with Transformer) Refinement of SelfCorrecting Re-Clustering (SCRC) Global Spatial Topology Utilization System Integration & Full Pipeline Debugging Real-world Deployment and Testing References [1] Zijian He, Kang Wang, Tian Fang, Lei Su, Rui Chen, and Xihong Fei. (2024). Comprehensive Performance Evaluation of YOLOv11, YOLOv10, YOLOv9, YOLOv8, and YOLOv5 on Object Detection of Power Equipment. arXiv preprint arXiv:2411.18871. [Online] [2] Tao Jiang, Peng Lu, Li Zhang, Ningsheng Ma, Rui Han, Chengqi Lyu, Yining Li, and Kai Chen. (2023). RTMPose: Real-Time Multi-Person Pose Estimation based on MMPose. arXiv preprint arXiv:2303.07399. [Online] [3] Jeongho Kim, Wooksu Shin, Hancheol Park, and Jongwon Baek. (2023). “Addressing the Occlusion Problem in Multi-Camera People Tracking with Human Pose Estimation,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops (CVPRW), Jun. 2023, pp. 5462–5468. [Online] [4] Kaiyang Zhou, Yongxin Yang, Andrea Cavallaro, and Tao Xiang. (2019). Omni-Scale Feature Learning for Person Re-Identification. arXiv preprint arXiv:1905.00953. [Online] [5] Syed Waqas Zamir, Aditya Arora, Salman Khan, Munawar Hayat, Fahad Shahbaz Khan, and MingHsuan Yang. (2021). Restormer: Efficient Transformer for High-Resolution Image Restoration. arXiv preprint arXiv:2111.09881v2. [Online] [6] Yalei Zhou, Peng Liu, Yue Cui, Chunguang Liu, and Wenli Duan. (2022). Integration of Multi-Head Self-Attention and Convolution for Person Re-Identification. Sensors, 22(16): 6293. [Online] [7] Yue Wang, Yongbin Sun, Ziwei Liu, Sanjay E. Sarma, Michael M. Bronstein, and Justin M. Solomon. (2019). Dynamic Graph CNN for Learning on Point Clouds. arXiv preprint arXiv:1801.07829v2. [Online] [8] Franco Manessi, Alessandro Rozza, and Mario Manzo. (2017). Dynamic Graph Convolutional Networks. arXiv preprint arXiv:1704.06199v1. [Online] [9] Nir Aharon, Roy Orfaig, and Ben-Zion Bobrovsky. (2022). BoT-SORT: Robust Associations MultiPedestrian Tracking. arXiv preprint arXiv:2206.14651. [Online] [10] Jeongho Kim, Wooksu Shin, Hancheol Park, and Donghyuk Choi. (2024). Cluster Self-Refinement for Enhanced Online Multi-Camera People Tracking. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops, 2024, pp. 7190–7197. [Online] [11] Pierre Baqué, François Fleuret, and Pascal Fua. (2017). Deep Occlusion Reasoning for MultiCamera Multi-Target Detection. arXiv preprint arXiv:1704.05775v2. [Online] [12] Mohammad Hossein Nasseri, Hadi Moradi, Reshad Hosseini, and Mohammadreza Babaee. (2021). Simple Online and Real-Time Tracking with Occlusion Handling. arXiv preprint arXiv:2103.04147v1. [Online] [13] Ran Eshel and Yael Moses. (2010). Tracking in a Dense Crowd Using Multiple Cameras. International Journal of Computer Vision, 88(1): 129–143. [Online] [14] Peng Zhang, Siqi Wang, Wei Zhang, Weimin Lei, Xinlei Zhao, Qingyang Jing, and Mingxin Liu. (2023). Cross-Camera Tracking Model and Method Based on Multi-Feature Fusion. Symmetry, 15(12): 2145. [Online] [15] Ran Eshel and Yael Moses. (2008). Homography-Based Multiple Camera Detection and Tracking of People in a Dense Crowd. In Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR), 2008, pp. 1–8. [Online] [16] David R. So, Wojciech Mańke, Hanxiao Liu, Zihang Dai, Noam Shazeer, and Quoc V. Le. (2022). Primer: Searching for Efficient Transformers for Language Modeling. arXiv preprint arXiv:2109.08668v2. [Online] [17] Kai Han, Yunhe Wang, Jianyuan Guo, Yehui Tang, and Enhua Wu. (2022). Vision GNN: An Image is Worth Graph of Nodes. arXiv preprint arXiv:2206.00272. [Online] [18] Lingxiao He, Yutian Lin, Zhihao Chen, Sheng Chen, Xiaodong Yang, and Jianchao Yang. (2020). FastReID: A PyTorch Toolbox for General Image Re-identification. arXiv preprint arXiv:2006.02631. [Online] [19] Xingang Pan, Ping Luo, Jianping Shi, and Xiaoou Tang. (2018). Two at Once: Enhancing Learning and Generalization Capacities via IBN-Net. In Proceedings of the European Conference on Computer Vision (ECCV), pp. 464–479. [Online] [20] Alexandre Alahi, Raphael Ortiz, and Silvio Savarese. (2016). Tracking the Untrackable: Learning to Track Multiple Cues with Long-Term Dependencies. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 3297–3306. [Online] Appendix ● Appendix A Dataset Market1501 Criteria Model OSNet Improved OSNet DukeMTMCReID Rank-1 mAP Rank-1 mAP Latency 92.28 93.32 80.30 83.12 85.28 86.62 70.71 72.28 14.34 17.12 ● Appendix B HOTA Metric and Its Components for MCPT: 1. HOTA – Higher Order Tracking Accuracy: HOTA (Higher Order Tracking Accuracy) evaluates overall tracking performance by balancing detection accuracy (DetA) and association accuracy (AssA). Formula: 𝐻𝑂𝑇𝐴 = √𝐴𝑠𝑠𝐴 × 𝐷𝑒𝑡𝐴 2. DetA – Detection Accuracy: DetA measures the system’s ability to correctly detect objects, considering both False Positives and False Negatives. Formula: 𝐷𝑒𝑡𝐴 = 𝑇𝑃 𝑇𝑃 + 𝐹𝑁 + 𝐹𝑃 - TP: True Positives (correct detections) - FN: False Negatives (missed detections) - FP: False Positives (incorrect detections) 3. AssA – Association Accuracy: AssA evaluates the tracker’s ability to maintain consistent identities of objects over time. Formula: 𝐴𝑠𝑠𝐴 = 𝑇𝑃𝐴𝑠𝑠𝑜𝑐𝑖𝑎𝑡𝑖𝑜𝑛(𝑐) 𝑇𝑃𝐴𝑠𝑠𝑜𝑐𝑖𝑎𝑡𝑖𝑜𝑛(𝑐) + 𝐹𝑁𝐴𝑠𝑠𝑜𝑐𝑖𝑎𝑡𝑖𝑜𝑛(𝑐) + 𝐹𝑃𝐴𝑠𝑠𝑜𝑐𝑖𝑎𝑡𝑖𝑜𝑛(𝑐) where c is the set of all matched GT-predicted pairs ● Appendix C Figure 6. This figure presents a side-by-side comparison between the original OSNet and our Transformer-augmented OSNet at various network depths (conv3 to conv5). For each block, we visualize three views: the input image, the corresponding heatmap, and the activation map. In the left column, baseline OSNet shows reasonable spatial attention but tends to disperse activation across background or non-informative regions. In contrast, the right column (OSNet with Transformers) demonstrates a marked improvement in semantic focus, particularly after the Transformer blocks (trans3 and trans4), where attention becomes more concentrated around discriminative regions of the human body such as the torso and limbs. Notably: • After trans3, the activation map becomes more sharply localized than its conv3 counterpart, indicating improved feature selectivity. • After trans4, the model focuses even more robustly on identity-specific cues, filtering out background clutter. • The conv5 layer further refines global contextual understanding, but the impact of Transformerenhanced blocks earlier in the network clearly contributes to higher-quality embeddings. This visualization supports the efficacy of our MDHA and DGC integration in enabling deeper layers to capture richer local-global and topological relationships. ● Appendix D Figure 7. Dynamic Graph Convolution (DGC), inspired by EdgeConv, selects neighbors based on spatial proximity (left) and feature similarity (right). Local neighbors capture geometric structure, while semantic neighbors reflect contextual similarity via cosine distance. An MLP aggregates edge features to learn context-aware embeddings.
0
You can add this document to your study collection(s)
Sign in Available only to authorized usersYou can add this document to your saved list
Sign in Available only to authorized users(For complaints, use another form )