The Open Radio Access Network (O-RAN) is reshaping traditional cellular networks by introducing flexible, interoperable, innovative solutions through open interfaces and RAN Intelligent Controllers (RICs). The RICs - Near- and Non-real time (Near-RT and Non-RT) - leverage Artificial Intelligence/Machine Learning (AI/ML) models for making intelligent decisions, which are critical to enhancing network performance. The effectiveness of the AI/ML models heavily depends on the nature of the training data, with homogeneous and heterogeneous datasets requiring tailored approaches. Existing methods predominantly rely on single statistical measures to determine dataset characteristics, which may not fully capture the complexities of real-world data from the Radio Units (RUs). This paper proposes an Ensemble learning approach that combines multiple statistical distance measures to assess dataset nature robustly, enhancing the accuracy of AI/ML model training. The proposed approach integrates the strengths of various statistical tests such as - Levene's Test (LT), Bartlett's Test (BT), Mood's Median Test (MMT), and Cramér-von Mises Test (CMT). The proposed Ensemble approach achieves superior performance in distinguishing between homogeneous and heterogeneous datasets, leading to more reliable model training by leveraging the complementary properties of the above tests. The results demonstrate that the Ensemble approach outperforms individual measures and shows the most reliable solution for robust dataset identification. The proposed approach not only minimizes the risk of misclassification but also ensures that datasets are accurately identified and processed.
Assessing Datasets for O-RAN RICs: Hoarding vs Choosing
Marotta, Andrea;
2024-01-01
Abstract
The Open Radio Access Network (O-RAN) is reshaping traditional cellular networks by introducing flexible, interoperable, innovative solutions through open interfaces and RAN Intelligent Controllers (RICs). The RICs - Near- and Non-real time (Near-RT and Non-RT) - leverage Artificial Intelligence/Machine Learning (AI/ML) models for making intelligent decisions, which are critical to enhancing network performance. The effectiveness of the AI/ML models heavily depends on the nature of the training data, with homogeneous and heterogeneous datasets requiring tailored approaches. Existing methods predominantly rely on single statistical measures to determine dataset characteristics, which may not fully capture the complexities of real-world data from the Radio Units (RUs). This paper proposes an Ensemble learning approach that combines multiple statistical distance measures to assess dataset nature robustly, enhancing the accuracy of AI/ML model training. The proposed approach integrates the strengths of various statistical tests such as - Levene's Test (LT), Bartlett's Test (BT), Mood's Median Test (MMT), and Cramér-von Mises Test (CMT). The proposed Ensemble approach achieves superior performance in distinguishing between homogeneous and heterogeneous datasets, leading to more reliable model training by leveraging the complementary properties of the above tests. The results demonstrate that the Ensemble approach outperforms individual measures and shows the most reliable solution for robust dataset identification. The proposed approach not only minimizes the risk of misclassification but also ensures that datasets are accurately identified and processed.Pubblicazioni consigliate
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


