AI-For-Beginners/3-NeuralNetworks/05-Frameworks/lab/LabFrameworks.ipynb

171 lines
4.5 KiB
Plaintext

{
"cells": [
{
"cell_type": "markdown",
"metadata": {},
"source": [
"# MNIST Digit Classification with our own Framework\n",
"\n",
"Lab Assignment from [AI for Beginners Curriculum](https://github.com/microsoft/ai-for-beginners).\n",
"\n",
"### Reading the Dataset\n",
"\n",
"This code download the dataset from the repository on the internet. You can also manually copy the dataset from `/data` directory of AI Curriculum repo."
]
},
{
"cell_type": "code",
"execution_count": 4,
"metadata": {
"tags": []
},
"outputs": [
{
"name": "stderr",
"output_type": "stream",
"text": [
" % Total % Received % Xferd Average Speed Time Time Time Current\n",
" Dload Upload Total Spent Left Speed\n",
"\n",
" 0 0 0 0 0 0 0 0 --:--:-- --:--:-- --:--:-- 0\n",
"100 9.9M 100 9.9M 0 0 9.9M 0 0:00:01 --:--:-- 0:00:01 15.8M\n"
]
}
],
"source": [
"!rm *.pkl\n",
"!wget https://raw.githubusercontent.com/microsoft/AI-For-Beginners/main/data/mnist.pkl.gz\n",
"!gzip -d mnist.pkl.gz"
]
},
{
"cell_type": "code",
"execution_count": 3,
"metadata": {},
"outputs": [],
"source": [
"import pickle\n",
"with open('mnist.pkl','rb') as f:\n",
" MNIST = pickle.load(f)"
]
},
{
"cell_type": "code",
"execution_count": 4,
"metadata": {},
"outputs": [],
"source": [
"labels = MNIST['Train']['Labels']\n",
"data = MNIST['Train']['Features']"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"Let's see what is the shape of data that we have:"
]
},
{
"cell_type": "code",
"execution_count": 5,
"metadata": {},
"outputs": [
{
"data": {
"text/plain": [
"(42000, 784)"
]
},
"execution_count": 5,
"metadata": {},
"output_type": "execute_result"
}
],
"source": [
"data.shape"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Splitting the Data\n",
"\n",
"We will use Scikit Learn to split the data between training and test dataset:"
]
},
{
"cell_type": "code",
"execution_count": 6,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Train samples: 33600, test samples: 8400\n"
]
}
],
"source": [
"from sklearn.model_selection import train_test_split\n",
"\n",
"features_train, features_test, labels_train, labels_test = train_test_split(data,labels,test_size=0.2)\n",
"\n",
"print(f\"Train samples: {len(features_train)}, test samples: {len(features_test)}\")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Instructions\n",
"\n",
"1. Take the framework code from the lesson and paste it into this notebook, or (even better) into a separate Python module\n",
"1. Define and train one-layered perceptron, observing training and validation accuracy during training\n",
"1. Try to understand if overfitting took place, and adjust layer parameters to improve accuracy\n",
"1. Repeat previous steps for 2- and 3-layered perceptrons. Try to experiment with different activation functions between layers.\n",
"1. Try to answer the following questions:\n",
" - Does the inter-layer activation function affect network performance?\n",
" - Do we need 2- or 3-layered network for this task?\n",
" - Did you experience any problems training the network? Especially as the number of layers increased.\n",
" - How do weights of the network behave during training? You may plot max abs value of weights vs. epoch to understand the relation."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": []
}
],
"metadata": {
"kernelspec": {
"display_name": "Python 3.7.4 64-bit (conda)",
"metadata": {
"interpreter": {
"hash": "86193a1ab0ba47eac1c69c1756090baa3b420b3eea7d4aafab8b85f8b312f0c5"
}
},
"name": "python3"
},
"language_info": {
"codemirror_mode": {
"name": "ipython",
"version": 3
},
"file_extension": ".py",
"mimetype": "text/x-python",
"name": "python",
"nbconvert_exporter": "python",
"pygments_lexer": "ipython3",
"version": "3.9.5"
},
"orig_nbformat": 2
},
"nbformat": 4,
"nbformat_minor": 2
}