• شماره مدرك
    21248
  • شماره راهنما
    18195
  • پديد آورنده

    الحمداوي، علي

  • عنوان

    به كارگيري الگوريتم‌هاي يادگيري ماشين در بهبود تشخيص و كاهش خطاهاي انساني در خدمات مراقبت‌هاي بهداشتي

  • مقطع تحصيلي
    كارشناسي ارشد
  • گرايش تحصيلي
    هوش مصنوعي و رباتيك
  • محل تحصيل
    اصفهان : دانشگاه صنعتي اصفهان
  • سال دفاع
    1405
  • صفحه شمار
    97ص.
  • توصيفگر ها

    يادگيري ماشين , بيماري‌هاي قلبي‌عروقي , ديابت ميليتوس , كاهش خطاي انساني , سيستم پشتيبان تصميم‌گيري باليني , جنگل تصادفي

  • تاريخ ورود اطلاعات
    1405/06/27
  • كتابنامه
    كتابنامه
  • رشته تحصيلي
    مهندسي كامپيوتر
  • دانشكده
    مهندسي برق و كامپيوتر
  • تاريخ ويرايش اطلاعات
    1405/06/28
  • كد ايرانداك
    23252865
  • چكيده فارسي
    بيماري‌هاي قلبي‌عروقي و ديابت از چالش‌هاي مهم سلامت عمومي به شمار مي‌آيند و تشخيص زودهنگام آن‌ها مي‌تواند به كاهش عوارض و هزينه‌هاي درمان كمك كند. پيچيدگي داده‌هاي باليني و احتمال بروز خطا در ارزيابي آن‌ها، ضرورت بررسي ابزارهاي پشتيبان تصميم‌گيري را افزايش داده است. پژوهش حاضر با هدف ارزيابي مقايسه‌اي الگوريتم‌هاي يادگيري ماشين در شناسايي بيماري قلبي و ديابت و بررسي ظرفيت آن‌ها براي پشتيباني از تشخيص انجام شد. در اين راستا، پنج الگوريتم جنگل تصادفي، شبكه عصبي مصنوعي، شبكه عصبي پيچشي يك‌بعدي (1D-CNN)، ماشين بردار پشتيبان با هسته تابع پايه شعاعي (SVM-RBF) و XGBoost در محيط پايتون پياده‌سازي و ارزيابي شدند. مجموعه‌داده بيماري قلبي UCI، پس از حذف شش ركورد ناقص، شامل 297 نمونه و 13 ويژگي باليني بود و نسخه متوازن مجموعه‌داده ديابت BRFSS 2015، تعداد 70٬692 نمونه و 21 ويژگي مرتبط با سلامت، سبك زندگي و مشخصات جمعيتي را در بر داشت. در هر دو مجموعه، مسئله به‌صورت طبقه‌بندي دودويي تعريف شد و داده‌ها با نسبت 80 درصد آموزش و 20 درصد آزمون، به‌صورت لايه‌اي و با بذر تصادفي ثابت 42 تقسيم شدند. استانداردسازي ويژگي‌ها با استفاده از پارامترهاي برآوردشده از داده‌هاي آموزش انجام گرفت و همان پارامترها بر داده‌هاي آزمون اعمال شدند تا از نشت اطلاعات جلوگيري شود. عملكرد مدل‌ها با معيارهاي دقت، صحت، بازخواني، امتياز F1 و سطح زير منحني ROC ارزيابي شد. نتايج نشان داد در مجموعه‌داده قلبي، مدل‌هاي SVM-RBF و XGBoost با دقت 85 درصد، رتبه نخست مشترك را كسب كردند و SVM-RBF با AUC برابر با 0/954، بالاترين توان تفكيك گزارش‌شده را داشت. دقت اين مدل نسبت به جنگل تصادفي و ANN، هركدام 2 درصد و نسبت به 1D-CNN حدود 4/08 درصد بهبود نسبي نشان داد، درحالي‌كه با XGBoost برابر بود. در مجموعه‌داده ديابت، بر اساس محاسبات ماتريس‌هاي درهم‌ريختگي، XGBoost با دقت 75/20 درصد و AUC برابر با 0/830، بالاترين دقت عددي را به دست آورد. بهبود نسبي دقت آن نسبت به 1D-CNN، ANN، SVM-RBF و جنگل تصادفي، به‌ترتيب حدود 0/01، 0/24، 0/29 و 0/35 درصد بود كه نشان‌دهنده نزديكي عملكرد مدل‌ها در اين مجموعه‌داده است. يافته‌ها از ظرفيت الگوريتم‌هاي بررسي‌شده براي توسعه ابزارهاي كمك‌تشخيصي حمايت مي‌كنند.
  • چكيده انگليسي
    Cardiovascular diseases an‎d diabetes are major public health challenges, an‎d their early detection can help reduce complications an‎d treatment costs. The complexity of clinical data an‎d the potential for errors in their interpretation highlight the need to investigate clinical decision-support tools. This study aimed to comparatively eva‎luate machine learning algorithms for identifying heart disease an‎d diabetes an‎d to examine their potential to support diagnostic decision-making. Five algorithms, Ran‎dom Forest (RF), Artificial Neural Network (ANN), One-Dimensional Convolutional Neural Network (1D-CNN), Support Vector Machine with a Radial Basis Function kernel (SVM-RBF), an‎d Extreme Gradient Boosting (XGBoost), were implemented an‎d eva‎luated in Python. After excluding six incomplete records, the UCI Heart Disease dataset comprised 297 samples with 13 clinical features. The balanced version of the BRFSS 2015 diabetes dataset contained 70,692 samples with 21 health-related, lifestyle, an‎d demographic features. Both tasks were formulated as binary classification problems. Each dataset was divided into training an‎d test sets using an 80:20 stratified split an‎d a fixed ran‎dom seed of 42. Stan‎dardization parameters were estimated exclusively from the training data an‎d subsequently applied to the test data to prevent data leakage. Model performance was eva‎luated using accuracy, precision, recall, F1-score, an‎d the area under the receiver operating characteristic curve (ROC-AUC). On the heart disease dataset, SVM-RBF an‎d XGBoost jointly achieved the highest accuracy of 85%, while SVM-RBF attained the highest reported AUC of 0.954. SVM-RBF demonstrated relative accuracy improvements of 2% over both RF an‎d ANN an‎d approximately 4.08% over 1D-CNN, while matching the accuracy of XGBoost. On the diabetes dataset, calculations based on the confusion matrices showed that XGBoost achieved the highest numerical accuracy, 75.20%, with an AUC of 0.830. Its relative accuracy improvements over 1D-CNN, ANN, SVM-RBF, an‎d RF were approximately 0.01%, 0.24%, 0.29%, an‎d 0.35%, respectively, indicating closely comparable performance across the models on this dataset. These findings support the potential of the eva‎luated algorithms for developing diagnostic support tools. However, the statistical significance of the observed performance differences an‎d any actual reduction in human diagnostic errors require further statistical an‎d clinical eva‎luation.
  • استاد راهنما
    محمدرضا احمدزاده , محمد داورپناه جزي
  • استاد داور
    الهام محمودزاده , مينا اميري