Metadata-Version: 2.4
Name: utelearn
Version: 10.4.2026
Summary: A custom machine learning library built for educational purpose from HCMUTE - Vietnam.
Author-email: Dung Cai <dung.cai@hcmute.edu.vn>
Requires-Python: >=3.7
Description-Content-Type: text/markdown
Requires-Dist: numpy
Requires-Dist: matplotlib

# utelearn

`utelearn` is a custom machine learning library built from scratch for educational purposes at UTE.

## Installation

You can install the package directly from PyPI:

```bash
pip install utelearn
```

## Features

- **KMeans Clustering**: Custom implementation of the KMeans clustering algorithm using NumPy.
- Built-in support for data visualization with Matplotlib.

## Quick Start

Here are quick examples of how to use the machine learning libs from `utelearn`:

## K-Means 
### Kmeans Clustering Example

```python
import numpy as np
import matplotlib.pyplot as plt
from utelearn.kmeans import KMeans

# 1. Tạo dữ liệu mẫu ngẫu nhiên (2 cụm rõ rệt)
np.random.seed(42)
X = np.vstack([
    np.random.randn(50, 2) + np.array([2, 2]),
    np.random.randn(50, 2) + np.array([-2, -2])
])

# 2. Khởi tạo và huấn luyện mô hình K-Means với k=2
kmeans = KMeans(k=2, max_iters=100)
kmeans.fit(X)

# 3. In kết quả trung tâm cụm (centroids)
print("Tọa độ các tâm cụm (Centroids):")
print(kmeans.centroids)

# 4. Dự đoán nhãn cho các điểm dữ liệu (nếu lớp KMeans có hàm predict)
# Hoặc gán nhãn dựa vào khoảng cách gần nhất tới các centroids
labels = np.argmin([np.linalg.norm(X - c, axis=1) for c in kmeans.centroids], axis=0)

# 5. Vẽ biểu đồ trực quan
plt.figure(figsize=(8, 6))
plt.scatter(X[:, 0], X[:, 1], c=labels, cmap='viridis', marker='o', alpha=0.7, edgecolors='k', label='Dữ liệu mẫu')
plt.scatter(kmeans.centroids[:, 0], kmeans.centroids[:, 1], c='red', marker='X', s=250, edgecolors='k', label='Tâm cụm (Centroids)')

plt.title("Trực quan hóa thuật toán K-Means - Utelearn", fontsize=14)
plt.xlabel("Trục X", fontsize=12)
plt.ylabel("Trục Y", fontsize=12)
plt.legend()
plt.grid(True, linestyle='--', alpha=0.5)
plt.show()
```

## K-Nearest Neighbors (KNN)
### 1. KNN Regression & Classification Example

```python
import numpy as np
from utelearn import knn

# Generate synthetic training data
X_train = np.random.rand(100, 1)
y_train = 4 * X_train**2 + 3 * X_train + 2 + np.random.randn(100, 1) * 0.1

# Generate test points
X_test = np.linspace(0, 1, 10).reshape(-1, 1)

# Perform KNN regression with k=5
predictions = knn.knn_regression(X_train, y_train, X_test, k=5)
print("Regression Predictions:", predictions)

# Generate synthetic training and test features/labels
X_train = np.random.rand(100, 2)
y_train = np.random.randint(0, 3, size=100) # 3 classes
X_test = np.random.rand(10, 2)

# Perform KNN classification with k=5
predictions = knn.knn_classify(X_train, y_train, X_test, k=5, n_classes=3)
print("Classification Predictions:", predictions)

```

## K-Nearest Neighbors (KNN)
### 1. KNN Regression & Classification Example using Points Lib

```python
import numpy as np
import matplotlib.pyplot as plt
import matplotlib.colors as mcolors
from utelearn import knn
from utelearn import points as point

# 1. Configuration parameters
n_training_points = 100
n_test_points = 500
n_classes = 3
dimension = 2
k_neighbors = 5

# 2. Generate Training and Test Data using utelearn's points module (Spiral Distribution)
training_data = point.Spiral(n_training_points, n_classes, dimension)
test_data = point.Spiral(n_test_points, n_classes, dimension)

# 3. Perform KNN Classification using utelearn's knn module
print("Running KNN classification on test data using utelearn...")
predicted_labels = knn.knn_classify(
    X_train=training_data.P, 
    y_train=training_data.L, 
    X_test=test_data.P, 
    k=k_neighbors, 
    n_classes=n_classes
)

# 4. Plotting the Results
plt.figure(figsize=(15, 5))

# Subplot 1: Reference Training Data
plt.subplot(1, 3, 1)
plt.scatter(
    training_data.P[:, 0], 
    training_data.P[:, 1], 
    c=training_data.L, 
    s=20, 
    cmap=mcolors.ListedColormap(["red", "purple", "green"])
)
plt.title(f"Reference Data ({n_training_points * n_classes} points)")
plt.grid(True, alpha=0.3)

# Subplot 2: Unclassified Test Data
plt.subplot(1, 3, 2)
plt.scatter(
    test_data.P[:, 0], 
    test_data.P[:, 1], 
    color="blue", 
    s=15
)
plt.title("Unclassified Test Data")
plt.grid(True, alpha=0.3)

# Subplot 3: Classified Test Data via utelearn KNN
plt.subplot(1, 3, 3)
plt.scatter(
    test_data.P[:, 0], 
    test_data.P[:, 1], 
    c=predicted_labels, 
    s=15, 
    cmap=mcolors.ListedColormap(["red", "purple", "green"])
)
plt.title(f"Test Data After KNN (k={k_neighbors})")
plt.grid(True, alpha=0.3)

plt.tight_layout()
plt.show()
```

## Naives Bayes Classifier (using Gaussian probability density function)
### Classification Example using Points Lib

import numpy as np 
import matplotlib.pyplot as plt 
from utelearn import points as point  
from utelearn import naives_bayes as nb

# Configuration parameters
n_generated_points = 1000  
n_classes = 3 
dimension = 2  
N_total = n_generated_points * n_classes  

###############################################################
################### Generate Training Data ####################
###############################################################
Generated_Data = point.Circle(n_generated_points, n_classes, dimension)  
X = np.column_stack((Generated_Data.P[:, 0], Generated_Data.P[:, 1]))
y = Generated_Data.L 

# Train the Naive Bayes classifier using the custom library
means, vars, priors = nb.train_naive_bayes(X, y) 

############################################################## 
################### Generate Test Data ####################### 
############################################################## 
Generated_Data1 = point.Circle(n_generated_points, n_classes, dimension)  
X_test = np.column_stack((Generated_Data1.P[:, 0], Generated_Data1.P[:, 1]))
y_test = Generated_Data1.L 

# Predict on the test data 
y_pred = nb.predict_naive_bayes(X_test, means, vars, priors) 

# Print predictions and actual labels 
print("Predictions:", y_pred) 
print("Actual labels:", y_test) 

# Calculate and print accuracy 
accuracy = np.mean(y_pred == y_test) 
print(f'Accuracy: {accuracy * 100:.2f}%') 

########################################################################## 
####################  Visualization Section  #############################
##########################################################################
x_min, x_max = X[:, 0].min() - 1, X[:, 0].max() + 1 
y_min, y_max = X[:, 1].min() - 1, X[:, 1].max() + 1 
xx, yy = np.meshgrid(np.arange(x_min, x_max, 0.1), np.arange(y_min, y_max, 0.1)) 

# Predict labels for all points in the meshgrid 
grid_points = np.c_[xx.ravel(), yy.ravel()] 
Z = nb.predict_naive_bayes(grid_points, means, vars, priors) 
Z = Z.reshape(xx.shape) 

# Plotting side-by-side comparison
plt.figure(figsize=(12, 5))

# Subplot 1: Original Test Data
plt.subplot(1, 2, 1)
plt.scatter(X_test[:, 0], X_test[:, 1], c=y_test, cmap='coolwarm', edgecolors='k', marker='o', label='Test Data Points') 
plt.xlabel('Feature X_test[0]') 
plt.ylabel('Feature X_test[1]') 
plt.title(f'Test Data ({n_classes} Classes)') 

# Subplot 2: Decision Boundary
plt.subplot(1, 2, 2)
plt.contourf(xx, yy, Z, alpha=0.3, cmap='coolwarm') 
plt.scatter(X_test[:, 0], X_test[:, 1], c=y_pred, cmap='coolwarm', edgecolors='k', marker='o', label='Predictions') 
plt.xlabel('Feature X_test[0]') 
plt.ylabel('Feature X_test[1]') 
plt.title(f'Naive Bayes Decision Boundary ({n_classes} Classes)') 
plt.legend(loc='best') 

plt.tight_layout()
plt.show()

## Author

- **Dung Cai** (dung.cai@hcmute.edu.vn)
