Kavindu Oshadha.

Project 04 / 12 · NLP / Machine Learning

ResearchScope AI

A research-paper classification study and interactive application comparing six models across six academic categories.

PythonNLPTF-IDFscikit-learnLogistic Regression
ResearchScope AI research-paper classification interface
ResearchScope AI research-paper classification interface · Click to view full size

The project.

ResearchScope AI was developed by Group 20 — ROT NLP Solutions. The study compares six machine-learning and deep-learning models using a balanced dataset of 15,000 arXiv papers across six academic categories.

As Team Lead, my work covered preprocessing, TF-IDF, Logistic Regression and LSTM. The best-performing model, Logistic Regression, achieved 89.33% test accuracy in the reported evaluation.

The Streamlit application makes the research interactive: users can explore predictions, confidence scores, the top three categories, keyword insights and model comparisons. The work ran from June to August 2026.

What it does.

Six-model comparison

Logistic Regression, Linear SVM, XGBoost, LSTM, 1D CNN and DistilBERT are compared in the study.

Balanced research dataset

15,000 arXiv papers provide a balanced basis for evaluation across six categories.

Prediction insights

The application exposes confidence scores, top-three predictions and keyword insights.

Reported result

Logistic Regression achieved 89.33% test accuracy within the project’s evaluation setup.

How it comes together.

Prepare → Clean and preprocess the research-paper text for model input.

Represent & train → Build TF-IDF representations and train classical and neural approaches.

Evaluate & demonstrate → Compare the models and expose predictions through Streamlit.

Tools & technologies.

PythonNLPTF-IDFscikit-learnLogistic RegressionLinear SVMXGBoostLSTM1D CNNDistilBERTStreamlit

Explore the project.

Technology documentation

Official references for the tools used in this project.