Klasifikasi Sepatu Sepak Bola Berdasarkan Peminatan, Brand, Dan Posisi Pemain Dengan Machine Learning
Keywords:
Klasifikasi, Sepatu Sepak Bola, Machine Learning, Random Forest, SMOTEAbstract
Pemilihan sepatu sepak bola yang tepat memengaruhi kenyamanan dan performa pemain, namun pemain amatir hingga semi-profesional sering mengalami kesulitan menentukan tipe sepatu yang sesuai karena kuatnya pengaruh merek dan kurangnya pertimbangan aspek teknis. Penelitian ini bertujuan mengklasifikasikan tipe sepatu sepak bola berdasarkan peminatan, brand, dan posisi pemain menggunakan algoritma Random Forest. Dataset awal berjumlah 3.971 data sepatu sepak bola. Setelah pembersihan data dan penanganan nilai kosong diperoleh 3.803 data, kemudian dibagi menjadi 3.042 data latih dan 761 data uji. Ketidakseimbangan kelas pada data latih ditangani menggunakan Synthetic Minority Over-sampling Technique (SMOTE) sehingga jumlah data latih menjadi 12.712. Hasil pengujian menunjukkan akurasi 99,61%, dengan 758 dari 761 data uji diklasifikasikan secara benar. Hasil feature importance menunjukkan BootsName sebagai fitur paling berpengaruh, diikuti Brand_BootsName dan BootsPosition. Model menunjukkan performa klasifikasi sangat baik. Model kemudian diimplementasikan dalam prototipe sistem berbasis web SMART BOOTS untuk menghasilkan klasifikasi dan rekomendasi sepatu berdasarkan karakteristik yang dimasukkan pengguna
References
N. Chmait and H. Westerbeek, “Artificial intelligence and machine learning in sport research: An introduction for non-data scientists,” Front. Sports Act. Living, vol. 3, Art. no. 682287, 2021, doi: 10.3389/fspor.2021.682287.
A. Felfernig, M. Wundara, T. N. T. Tran, V.-M. Le, S. Lubos, and S. Polat-Erdeniz, “Sports recommender systems: Overview and research directions,” J. Intell. Inf. Syst., vol. 62, pp. 1125–1164, 2024, doi: 10.1007/s10844-024-00857-w.
F. H. Yagin, U. C. H. Hasan, F. M. Clemente, O. Eken, G. Badicu, and M. Gulu, “Using machine learning to determine the positions of professional soccer players in terms of biomechanical variables,” Proc. Inst. Mech. Eng. P J. Sports Eng. Technol., vol. 239, no. 4, pp. 726–733, 2025, doi: 10.1177/17543371231199814.
I. Behravan and S. M. Razavi, “A novel machine learning method for estimating football players’ value in the transfer market,” Soft Comput., vol. 25, no. 3, pp. 2499–2511, 2021, doi: 10.1007/s00500-020-05319-3.
Y.-E. Yeh, “Prediction of optimized color design for sports shoes using an artificial neural network and genetic algorithm,” Appl. Sci., vol. 10, no. 5, Art. no. 1560, 2020, doi: 10.3390/app10051560.
J. W. Wannop, D. J. Stefanyshyn, R. B. Anderson, M. J. Coughlin, and R. Kent, “Development of a footwear sizing system in the National Football League,” Sports Health, vol. 11, no. 1, pp. 40–46, 2019, doi: 10.1177/1941738118789402.
L. Breiman, “Random forests,” Mach. Learn., vol. 45, pp. 5–32, 2001, doi: 10.1023/A:1010933404324.
A. Liaw and M. Wiener, “Classification and regression by randomForest,” R News, vol. 2, no. 3, pp. 18–22, 2002.
G. Biau and E. Scornet, “A random forest guided tour,” TEST, vol. 25, no. 2, pp. 197–227, 2016, doi: 10.1007/s11749-016-0481-7.
M. Fernández-Delgado, E. Cernadas, S. Barro, and D. Amorim, “Do we need hundreds of classifiers to solve real world classification problems?,” J. Mach. Learn. Res., vol. 15, pp. 3133–3181, 2014.
A.-L. Boulesteix, S. Janitza, J. Kruppa, and I. R. König, “Overview of random forest methodology and practical guidance with emphasis on computational biology and bioinformatics,” WIREs Data Min. Knowl. Discov., vol. 2, no. 6, pp. 493–507, 2012, doi: 10.1002/widm.1072.
P. Probst, M. N. Wright, and A.-L. Boulesteix, “Hyperparameters and tuning strategies for random forest,” WIREs Data Min. Knowl. Discov., vol. 9, no. 3, Art. no. e1301, 2019, doi: 10.1002/widm.1301.
N. V. Chawla, K. W. Bowyer, L. O. Hall, and W. P. Kegelmeyer, “SMOTE: Synthetic minority over-sampling technique,” J. Artif. Intell. Res., vol. 16, pp. 321–357, 2002, doi: 10.1613/jair.953.
J. D. Rodríguez, A. Pérez, and J. A. Lozano, “Sensitivity analysis of k-fold cross validation in prediction error estimation,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 32, no. 3, pp. 569–575, Mar. 2010, doi: 10.1109/TPAMI.2009.187.
A. Zikry and F. A. Siregar, “Performance analysis of logistic regression and SVM (support vector machine) algorithms on e-football mobile game review sentiment,” J. Sci. Technol. Innov., vol. 1, no. 3, pp. 212–224, 2026, doi: 10.65310/073zdt91.
R. Genuer, J.-M. Poggi, and C. Tuleau-Malot, “Variable selection using random forests,” Pattern Recognit. Lett., vol. 31, no. 14, pp. 2225–2236, 2010, doi: 10.1016/j.patrec.2010.03.014.
C. Strobl, A.-L. Boulesteix, T. Kneib, T. Augustin, and A. Zeileis, “Conditional variable importance for random forests,” BMC Bioinformatics, vol. 9, Art. no. 307, 2008, doi: 10.1186/1471-2105-9-307.
S. Janitza and R. Hornung, “On the overestimation of random forest’s out-of-bag error,” PLoS ONE, vol. 13, no. 8, Art. no. e0201904, 2018, doi: 10.1371/journal.pone.0201904.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Robby Kurniawan Syahfitra, Farid Akbar Siregar

This work is licensed under a Creative Commons Attribution 4.0 International License.













