شماره مدرك
21201
شماره راهنما
18167
پديد آورنده
الكريش، عبدالله
عنوان
تحليل ترافيك شبكه براي شناسايي تهديدهاي سايبري بر پايه يادگيري عميق و انتخاب ويژگي
مقطع تحصيلي
كارشناسي ارشد
گرايش تحصيلي
هوش مصنوعي و رباتيك
محل تحصيل
اصفهان : دانشگاه صنعتي اصفهان
سال دفاع
1405
صفحه شمار
102 ص.
توصيفگر ها
شناسايي تهديدهاي سايبري , تحليل ترافيك شبكه , يادگيري عميق , كسب اطلاعات , انتخاب ويژگي بهينهشده
تاريخ ورود اطلاعات
1405/05/27
كتابنامه
كتابنامه
رشته تحصيلي
مهندسي كامپيوتر
دانشكده
مهندسي برق و كامپيوتر
تاريخ ويرايش اطلاعات
1405/05/27
كد ايرانداك
23243442
چكيده فارسي
گسترش سريع شبكههاي كامپيوتري و پيچيدگي روزافزون تهديدات سايبري (نظير بدافزارهاي چندريختي و تهديدات پيشرفتهي پايدار)، كارايي سامانههاي سنتي تشخيص نفوذِ مبتني بر امضا را بهشدت كاهش داده است. از سوي ديگر، روشهاي مبتني بر ناهنجاري نيز با توليد نرخ بالاي مثبت كاذب (10 تا 30 درصد)، منجر به چالشهايي چون خستگي از هشدار براي تحليلگران ميشوند. عليرغم انجام پژوهشهاي گسترده در زمينه معماريهاي تركيبي يادگيري عميق براي تشخيص نفوذ، خلأ پژوهشي قابلتوجهي در مقايسه نظاممند روشهاي كاهش بعد احساس ميشود؛ بهطوريكه تاكنون مطالعهاي به مقايسه دقيق انتخاب ويژگي مبتني بر بهره اطلاعاتي (IG) و تحليل مؤلفههاي اصلي (PCA) در شرايط كاملاً كنترلشده براي چارچوب CNN-LSTM نپرداخته است. در اين پاياننامه، يك چارچوب هوشمند تشخيص تهديدات سايبري مبتني بر معماري تركيبي CNN-LSTM پيشنهاد شده و بر روي مجموعهداده معيار CIC-IDS2017 (شامل بيش از 8/2ميليون ركورد جريان شبكه) مورد ارزيابي قرار گرفته است. اين چارچوب از يك روند پيشپردازش ششمرحلهاي و بدون نشت داده بهره ميبرد كه شامل تكنيكهايي نظير نرمالسازي و متوازنسازي كلاسها با استفاده از الگوريتم SMOTE (اعمالشده صرفاً روي دادههاي آموزش) است. به منظور ارزيابي، سه پيكربندي مختلف شامل: تمامي 78 ويژگي، 20 ويژگي برتر مبتني بر IG و 20 مؤلفه حاصل از PCA، روي معماري يكسان CNN-LSTM آموزش داده شده و با يكديگر مقايسه شدند. نتايج نشان داد كه مدل مبتني بر 20 ويژگي IG، بر روي مجموعه دادههاي آزمون مستقل، به دقت 64/98٪، معيار F1 برابر با 33/89٪، نرخ مثبت كاذب 08/4٪ و مقدار AUC-ROC معادل 9997/0دست يافته است. با اين حال، مقايسههاي صورت گرفته مشخص كرد كه هر سه پيكربندي عملكردي تقريباً مشابه دارند و فرضيه برتري محسوس IG بر PCA تأييد نشد؛ بلكه تفاوتهاي جزئيِ مشاهدهشده در معيار F1، عمدتاً ناشي از خطاي طبقهبندي در كلاسهاي نادر مانند Infiltration بوده است. افزون بر اين، ارزيابي مدل مبنا مبتني بر XGBoost نشان داد كه اين مدل با ثبت معيار F1 برابر با 36/98٪، عملكردي همسطح و حتي بهتر از مدل يادگيري عميق دارد. تحليل عملكرد مدل در كلاسهاي مختلف، يك سلسلهمراتب سهسطحي از دشواريِ تشخيص را نمايان ساخت؛ بهگونهاي كه شناسايي حملات حجمي تقريباً به حداكثر دقت ممكن رسيد، اما تشخيص حملات نادر همچنان چالشي اساسي در تحليل ترافيك شبكه به شمار ميرود. در نهايت، تحليل اثر حذف مؤلفهها ثابت كرد كه تكنيك SMOTE حياتيترين عامل در عملكرد مدل است؛ بهطوريكه حذف آن موجب افت 16 واحدي در معيار F1 ميشود، درحاليكه حذف لايههاي CNN و LSTM تأثير بهمراتب كمتري (بهترتيب 1/5و 0/3 واحد افت) در پي داشت.
چكيده انگليسي
The rapid expansion of computer networks and the increasing complexity of cyber threats, such as polymorphic malware and advanced persistent threats (APTs), have severely diminished the effectiveness of traditional signature-based intrusion detection systems (IDS). Conversely, anomaly-based methods often suffer from high false-positive rates (10–30%), leading to challenges like alert fatigue for security analysts. Despite extensive research on hybrid deep learning architectures for intrusion detection, there remains a notable research gap in the systematic comparison of dimensionality reduction techniques. Specifically, no study has yet provided a precise, controlled comparison between Information Gain (IG)-based feature selection and Principal Component Analysis (PCA) within a CNN-LSTM framework. In this thesis, an intelligent cyber-threat detection framework based on a hybrid CNN-LSTM architecture is proposed and evaluated using the benchmark CIC-IDS2017 dataset, which contains over 2.8 million network flow records. The framework employs a six-stage, data-leakage-free preprocessing pipeline, incorporating techniques such as normalization and class balancing via the SMOTE algorithm (applied strictly to the training data). For evaluation, three configurations—utilizing all 78 features, the top 20 features based on IG, and 20 components derived from PCA—were trained on the identical CNN-LSTM architecture and compared. The results demonstrated that the model based on the 20 IG features achieved an accuracy of 98.64%, an F1-score of 89.33%, a false-positive rate of 4.08%, and an AUC-ROC of 0.9997 on independent test sets. However, the comparative analysis revealed that all three configurations performed similarly, failing to confirm the hypothesis of a significant superiority of IG over PCA. Instead, minor observed differences in the F1-score were primarily attributed to classification errors in rare classes, such as ‘Infiltration’. Furthermore, benchmarking against an XGBoost model showed that the latter achieved competitive or even superior performance, with an F1-score of 98.36%. Analysis of class-specific performance revealed a three-tier hierarchy of detection difficulty: while high-volume attacks achieved near-maximal detection accuracy, identifying rare attacks remains a critical challenge in network traffic analysis. Finally, ablation studies demonstrated that the SMOTE technique is the most critical factor in model performance; removing it resulted in a 16-point drop in the F1-score, whereas the removal of the CNN and LSTM layers had significantly less impact (drops of 5.1 and 3.0 points, respectively).
استاد راهنما
علي فانيان , محمد داورپناه جزي
استاد داور
الهام محمودزاده , مينا اميري