Add image attributions
This commit is contained in:
parent
82dea7e1a4
commit
a96da961f9
|
|
@ -12,7 +12,7 @@ Originally, computers were invented by [Charles Babbage](https://en.wikipedia.or
|
|||
|
||||

|
||||
|
||||
> Photo by TODO
|
||||
> Photo by [Vickie Soshnikova](http://twitter.com/vickievalerie)
|
||||
|
||||
> ✅ Defining the age of a person from his or her photograph is a task that cannot be explicitly programmed, because we do not know how we come up with a number inside our head when we do it.
|
||||
|
||||
|
|
@ -30,7 +30,11 @@ The task of solving a specific human-like problem, such as determining a person'
|
|||
|
||||
One of the problems when dealing with the term **[Intelligence](https://en.wikipedia.org/wiki/Intelligence)** is that there is no clear definition of this term. One can argue that intelligence is connected to **abstract thinking**, or to **self-awareness**, but we cannot properly define it.
|
||||
|
||||
 | To see the ambiguity of a term *intelligence*, try answering a question: "Is a cat intelligent?". Different people tend to give different answers to this question, as there is no universally accepted test to prove the assertion true or not. And if you think there is - try running your cat through an IQ test...
|
||||

|
||||
|
||||
> [Photo](https://unsplash.com/photos/75715CVEJhI) by [Amber Kipp](https://unsplash.com/@sadmax) from Unsplash
|
||||
|
||||
To see the ambiguity of a term *intelligence*, try answering a question: "Is a cat intelligent?". Different people tend to give different answers to this question, as there is no universally accepted test to prove the assertion true or not. And if you think there is - try running your cat through an IQ test...
|
||||
|
||||
✅ Think for a minute about how you define intelligence. Is a crow who can solve a maze and get at some food intelligent? Is a child intelligent?
|
||||
|
||||
|
|
@ -86,7 +90,7 @@ Artificial Intelligence was started as a field in the middle of the twentieth ce
|
|||
|
||||
<img alt="Brief History of AI" src="images/history-of-ai.png" width="70%"/>
|
||||
|
||||
> Image by TODO
|
||||
> Image by [Dmitry Soshnikov](http://soshnikov.com)
|
||||
|
||||
As time passed, computing resources became cheaper, and more data has become available, so neural network approaches started demonstrating great performance in competing with human beings in many areas, such as computer vision or speech understanding. In the last decade, the term Artificial Intelligence has been mostly used as a synonym for Neural Networks, because most of the AI successes that we hear about are based on them.
|
||||
|
||||
|
|
@ -106,7 +110,7 @@ Similarly, we can see how the approach towards creating “talking programs” (
|
|||
|
||||
<img alt="the Turing test's evolution" src="images/turing-test-evol.png" width="70%"/>
|
||||
|
||||
> Image by TODO
|
||||
> Image by Dmitry Soshnikov, [photo](https://unsplash.com/photos/r8LmVbUKgns) by [Marina Abrosimova](https://unsplash.com/@abrosimova_marina_foto), Unsplash
|
||||
|
||||
## Recent AI Research
|
||||
|
||||
|
|
@ -114,7 +118,7 @@ The huge recent growth in neural network research started around 2010, when larg
|
|||
|
||||

|
||||
|
||||
> TODO image attribution
|
||||
> Image by [Dmitry Soshnikov](http://soshnikov.com)
|
||||
|
||||
In 2012, [Convolutional Neural Networks](../4-ComputerVision/07-ConvNets/README.md) were first used in image classification, which led to a significant drop in classification errors (from almost 30% to 16.4%). In 2015, ResNet architecture from Microsoft Research [achieved human-level accuracy](https://doi.org/10.1109/ICCV.2015.123).
|
||||
|
||||
|
|
|
|||
|
|
@ -2,6 +2,9 @@
|
|||
|
||||

|
||||
|
||||
> Sketchnote by [Tomomi Imura](https://twitter.com/girlie_mac)
|
||||
|
||||
|
||||
In the early days of AI, top-down approach to creating intelligent systems was popular. The idea was to extract the knowledge from people into some machine-readable form, and then use it to automatically solve problems. This approach was based on two big ideas:
|
||||
|
||||
* Knowledge Representation
|
||||
|
|
@ -31,6 +34,8 @@ Thus, the problem of **knowledge representation** is to find some effective way
|
|||
|
||||

|
||||
|
||||
> Image by [Dmitry Soshnikov](http://soshnikov.com)
|
||||
|
||||
We can classify different computer knowledge representation methods in the following categories:
|
||||
|
||||
* **Network representations** are based on the fact that we have a network of interrelated concepts inside our head. We can try to reproduce the same networks as a graph inside a computer - so-called **semantic network**.
|
||||
|
|
@ -82,6 +87,8 @@ As an example, let's consider the following expert system of determining an anim
|
|||
|
||||

|
||||
|
||||
> Image by [Dmitry Soshnikov](http://soshnikov.com)
|
||||
|
||||
This diagram is called **AND-OR tree**, and it is a graphical representation of a set of production rules. Drawing a tree is useful at the beginning of extracting knowledge from the expert, and to represent the knowledge inside the computer it is more convenient to use rules:
|
||||
```
|
||||
IF animal eats meat
|
||||
|
|
@ -148,6 +155,8 @@ In a more complex case, if we want to define a list of creators, we can use some
|
|||
|
||||
<img src="images/triplet-complex.png" width="40%"/>
|
||||
|
||||
> Diagrams above by [Dmitry Soshnikov](http://soshnikov.com)
|
||||
|
||||
Progress of building Semantic Web was somehow slowed down by the success of search engines and natural language processing techniques, which allow extracting structured data from text. However, in some areas there are still significant efforts to maintain ontologies and knowledgebases. A few projects worth noting:
|
||||
* [WikiData](https://wikidata.org/) is a collection of machine readable knowledgebases associated with Wikipedia. Most of the data is mined from Wikipedia *InfoBoxes*, pieces of structured content inside Wikipedia pages. You can [query](https://query.wikidata.org/) wikidata in SPARQL, a special query language for Semantic Web. Here is a sample query that displays most popular eye colors among humans:
|
||||
```sparql
|
||||
|
|
@ -167,7 +176,7 @@ GROUP BY ?eyeColorLabel
|
|||
|
||||
<img src="images/protege.png" width="70%"/>
|
||||
|
||||
*Web Protégé editor open with Romanov Family ontology*
|
||||
*Web Protégé editor open with Romanov Family ontology. Screenshot by Dmitry Soshnikov*
|
||||
|
||||
See [FamilyOntology.ipynb](FamilyOntology.ipynb) for an example of using Semantic Web techniques to reason about family relationships. We will take a family tree represented in common GEDCOM format, and an ontology of family relationships, and build a graph of all family relationships for given set of individuals.
|
||||
|
||||
|
|
|
|||
|
|
@ -5,6 +5,8 @@ One of the first attempts to implement something similar to a modern neural netw
|
|||
<img src='images/Rosenblatt-wikipedia.jpg' alt='Frank Rosenblatt'/> | <img src='images/Mark_I_perceptron_wikipedia.jpg' alt='The Mark 1 Perceptron' />
|
||||
-----|-----
|
||||
|
||||
> Images [from Wikipedia](https://en.wikipedia.org/wiki/Perceptron)
|
||||
|
||||
An input image was represented by 20x20 photocell array, so the neural network had 400 inputs and one binary output. Simple network contained one neuron, also called **threshold logic unit**. Neural network weights were potentiometers that required manual adjustment during the training phase.
|
||||
|
||||
> New York Times wrote about perceptron at that time:
|
||||
|
|
@ -64,3 +66,10 @@ def train(positive_examples, negative_examples, num_iterations = 100, eta = 1):
|
|||
## [Proceed to Notebook](Perceptron.ipynb)
|
||||
|
||||
To see how we can use perceptron to solve some toy as well as real-life problems, and to continue learning - go to [Perceptron](Perceptron.ipynb) notebook.
|
||||
|
||||
## [Lab](lab/README.md)
|
||||
|
||||
In this lesson, we have implemented perceptron for binary classification task, and we have used it to classify between two handwritten digits. In this lab, you are asked to solve the problem of digit classification entirely, i.e. determine which digit is most likely to correspond to a given image.
|
||||
|
||||
* [Instructions](lab/README.md)
|
||||
* [Notebook](lab/PerceptronMultiClass.ipynb)
|
||||
|
|
|
|||
|
|
@ -59,3 +59,10 @@ We will cover back prop in much more detail in our notebook example.
|
|||
## [Proceed to Notebook](OwnFramework.ipynb)
|
||||
|
||||
In the accompanying notebook, we will implement our own framework for building and training multi-layered perceptrons. You will be able to see in detail how modern neural networks operate. Proceed to [OwnFramework](OwnFramework.ipynb) notebook.
|
||||
|
||||
## [Lab](lab/README.md)
|
||||
|
||||
In this lesson, we have our own neural network library, and we have used for a simple two-dimensional classification task. In this lab, you are asked to use this framework to solve MNIST handwritten digit classification.
|
||||
|
||||
* [Instructions](lab/README.md)
|
||||
* [Notebook](lab/MyFW_MNIST.ipynb)
|
||||
|
|
@ -0,0 +1,170 @@
|
|||
{
|
||||
"cells": [
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"# MNIST Digit Classification with our own Framework\n",
|
||||
"\n",
|
||||
"Lab Assignment from [AI for Beginners Curriculum](https://github.com/microsoft/ai-for-beginners).\n",
|
||||
"\n",
|
||||
"### Reading the Dataset\n",
|
||||
"\n",
|
||||
"This code download the dataset from the repository on the internet. You can also manually copy the dataset from `/data` directory of AI Curriculum repo."
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 4,
|
||||
"metadata": {
|
||||
"tags": []
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stderr",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
" % Total % Received % Xferd Average Speed Time Time Time Current\n",
|
||||
" Dload Upload Total Spent Left Speed\n",
|
||||
"\n",
|
||||
" 0 0 0 0 0 0 0 0 --:--:-- --:--:-- --:--:-- 0\n",
|
||||
"100 9.9M 100 9.9M 0 0 9.9M 0 0:00:01 --:--:-- 0:00:01 15.8M\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"!rm *.pkl\n",
|
||||
"!wget https://raw.githubusercontent.com/microsoft/AI-For-Beginners/main/data/mnist.pkl.gz\n",
|
||||
"!gzip -d mnist.pkl.gz"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 3,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"import pickle\n",
|
||||
"with open('mnist.pkl','rb') as f:\n",
|
||||
" MNIST = pickle.load(f)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"labels = MNIST['Train']['Labels']\n",
|
||||
"data = MNIST['Train']['Features']"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Let's see what is the shape of data that we have:"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 5,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"(42000, 784)"
|
||||
]
|
||||
},
|
||||
"execution_count": 5,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"data.shape"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"### Splitting the Data\n",
|
||||
"\n",
|
||||
"We will use Scikit Learn to split the data between training and test dataset:"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 6,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Train samples: 33600, test samples: 8400\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"from sklearn.model_selection import train_test_split\n",
|
||||
"\n",
|
||||
"features_train, features_test, labels_train, labels_test = train_test_split(data,labels,test_size=0.2)\n",
|
||||
"\n",
|
||||
"print(f\"Train samples: {len(features_train)}, test samples: {len(features_test)}\")"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"### Instructions\n",
|
||||
"\n",
|
||||
"1. Take the framework code from the lesson and paste it into this notebook, or (even better) into a separate Python module\n",
|
||||
"1. Define and train one-layered perceptron, observing training and validation accuracy during training\n",
|
||||
"1. Try to understand if overfitting took place, and adjust layer parameters to improve accuracy\n",
|
||||
"1. Repeat previous steps for 2- and 3-layered perceptrons. Try to experiment with different activation functions between layers.\n",
|
||||
"1. Try to answer the following questions:\n",
|
||||
" - Does the inter-layer activation function affect network performance?\n",
|
||||
" - Do we need 2- or 3-layered network for this task?\n",
|
||||
" - Did you experience any problems training the network? Especially as the number of layers increased.\n",
|
||||
" - How do weights of the network behave during training? You may plot max abs value of weights vs. epoch to understand the relation."
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": []
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
"kernelspec": {
|
||||
"display_name": "Python 3.7.4 64-bit (conda)",
|
||||
"metadata": {
|
||||
"interpreter": {
|
||||
"hash": "86193a1ab0ba47eac1c69c1756090baa3b420b3eea7d4aafab8b85f8b312f0c5"
|
||||
}
|
||||
},
|
||||
"name": "python3"
|
||||
},
|
||||
"language_info": {
|
||||
"codemirror_mode": {
|
||||
"name": "ipython",
|
||||
"version": 3
|
||||
},
|
||||
"file_extension": ".py",
|
||||
"mimetype": "text/x-python",
|
||||
"name": "python",
|
||||
"nbconvert_exporter": "python",
|
||||
"pygments_lexer": "ipython3",
|
||||
"version": "3.9.5"
|
||||
},
|
||||
"orig_nbformat": 2
|
||||
},
|
||||
"nbformat": 4,
|
||||
"nbformat_minor": 2
|
||||
}
|
||||
|
|
@ -0,0 +1,21 @@
|
|||
# MNIST Classification with Our Own Framework
|
||||
|
||||
Lab Assignment from [AI for Beginners Curriculum](https://github.com/microsoft/ai-for-beginners).
|
||||
|
||||
## Task
|
||||
|
||||
Solve the MNIST handwritten digit classification problem using 1-, 2- and 3-layered perceptron. Use the neural network framework we have developed in the lesson.
|
||||
|
||||
|
||||
## Stating Notebook
|
||||
|
||||
Start the lab by opening [MyFW_MNIST.ipynb](MyFW_MNIST.ipynb)
|
||||
|
||||
## Questions
|
||||
|
||||
As a result of this lab, try to answer the following questions:
|
||||
|
||||
- Does the inter-layer activation function affect network performance?
|
||||
- Do we need 2- or 3-layered network for this task?
|
||||
- Did you experience any problems training the network? Especially as the number of layers increased.
|
||||
- How do weights of the network behave during training? You may plot max abs value of weights vs. epoch to understand the relation.
|
||||
|
|
@ -0,0 +1,170 @@
|
|||
{
|
||||
"cells": [
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"# MNIST Digit Classification with our own Framework\n",
|
||||
"\n",
|
||||
"Lab Assignment from [AI for Beginners Curriculum](https://github.com/microsoft/ai-for-beginners).\n",
|
||||
"\n",
|
||||
"### Reading the Dataset\n",
|
||||
"\n",
|
||||
"This code download the dataset from the repository on the internet. You can also manually copy the dataset from `/data` directory of AI Curriculum repo."
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 4,
|
||||
"metadata": {
|
||||
"tags": []
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stderr",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
" % Total % Received % Xferd Average Speed Time Time Time Current\n",
|
||||
" Dload Upload Total Spent Left Speed\n",
|
||||
"\n",
|
||||
" 0 0 0 0 0 0 0 0 --:--:-- --:--:-- --:--:-- 0\n",
|
||||
"100 9.9M 100 9.9M 0 0 9.9M 0 0:00:01 --:--:-- 0:00:01 15.8M\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"!rm *.pkl\n",
|
||||
"!wget https://raw.githubusercontent.com/microsoft/AI-For-Beginners/main/data/mnist.pkl.gz\n",
|
||||
"!gzip -d mnist.pkl.gz"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 3,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"import pickle\n",
|
||||
"with open('mnist.pkl','rb') as f:\n",
|
||||
" MNIST = pickle.load(f)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"labels = MNIST['Train']['Labels']\n",
|
||||
"data = MNIST['Train']['Features']"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Let's see what is the shape of data that we have:"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 5,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"(42000, 784)"
|
||||
]
|
||||
},
|
||||
"execution_count": 5,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"data.shape"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"### Splitting the Data\n",
|
||||
"\n",
|
||||
"We will use Scikit Learn to split the data between training and test dataset:"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 6,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Train samples: 33600, test samples: 8400\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"from sklearn.model_selection import train_test_split\n",
|
||||
"\n",
|
||||
"features_train, features_test, labels_train, labels_test = train_test_split(data,labels,test_size=0.2)\n",
|
||||
"\n",
|
||||
"print(f\"Train samples: {len(features_train)}, test samples: {len(features_test)}\")"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"### Instructions\n",
|
||||
"\n",
|
||||
"1. Take the framework code from the lesson and paste it into this notebook, or (even better) into a separate Python module\n",
|
||||
"1. Define and train one-layered perceptron, observing training and validation accuracy during training\n",
|
||||
"1. Try to understand if overfitting took place, and adjust layer parameters to improve accuracy\n",
|
||||
"1. Repeat previous steps for 2- and 3-layered perceptrons. Try to experiment with different activation functions between layers.\n",
|
||||
"1. Try to answer the following questions:\n",
|
||||
" - Does the inter-layer activation function affect network performance?\n",
|
||||
" - Do we need 2- or 3-layered network for this task?\n",
|
||||
" - Did you experience any problems training the network? Especially as the number of layers increased.\n",
|
||||
" - How do weights of the network behave during training? You may plot max abs value of weights vs. epoch to understand the relation."
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": []
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
"kernelspec": {
|
||||
"display_name": "Python 3.7.4 64-bit (conda)",
|
||||
"metadata": {
|
||||
"interpreter": {
|
||||
"hash": "86193a1ab0ba47eac1c69c1756090baa3b420b3eea7d4aafab8b85f8b312f0c5"
|
||||
}
|
||||
},
|
||||
"name": "python3"
|
||||
},
|
||||
"language_info": {
|
||||
"codemirror_mode": {
|
||||
"name": "ipython",
|
||||
"version": 3
|
||||
},
|
||||
"file_extension": ".py",
|
||||
"mimetype": "text/x-python",
|
||||
"name": "python",
|
||||
"nbconvert_exporter": "python",
|
||||
"pygments_lexer": "ipython3",
|
||||
"version": "3.9.5"
|
||||
},
|
||||
"orig_nbformat": 2
|
||||
},
|
||||
"nbformat": 4,
|
||||
"nbformat_minor": 2
|
||||
}
|
||||
|
|
@ -0,0 +1,26 @@
|
|||
# Classification with PyTorch/TensorFlow
|
||||
|
||||
Lab Assignment from [AI for Beginners Curriculum](https://github.com/microsoft/ai-for-beginners).
|
||||
|
||||
## Task
|
||||
|
||||
Solve two classification problems using single- and multi-layered fully-connected networks using PyTorch or TensorFlow:
|
||||
|
||||
1. Iris classification problem
|
||||
1. MNIST handwritten digit classification problem
|
||||
|
||||
Try different network architectures to achieve the best accuracy you can get.
|
||||
|
||||
|
||||
## Stating Notebook
|
||||
|
||||
Start the lab by opening [MyFW_MNIST.ipynb](MyFW_MNIST.ipynb)
|
||||
|
||||
## Questions
|
||||
|
||||
As a result of this lab, try to answer the following questions:
|
||||
|
||||
- Does the inter-layer activation function affect network performance?
|
||||
- Do we need 2- or 3-layered network for this task?
|
||||
- Did you experience any problems training the network? Especially as the number of layers increased.
|
||||
- How do weights of the network behave during training? You may plot max abs value of weights vs. epoch to understand the relation.
|
||||
|
|
@ -26,11 +26,11 @@ From biology we know that our brain consists of neural cells, each of them havin
|
|||
|
||||
 | 
|
||||
----|----
|
||||
Real Neuron | Artificial Neuron
|
||||
Real Neuron *([Image](https://en.wikipedia.org/wiki/Synapse#/media/File:SynapseSchematic_lines.svg) from Wikipedia)* | Artificial Neuron *(Image by Author)*
|
||||
|
||||
Thus, simplest mathematical model of a neuron contains several inputs X<sub>1</sub>, ..., X<sub>N</sub> and an output Y, and a series of weights W<sub>1</sub>, ..., W<sub>N</sub>. An output is calculated as
|
||||
|
||||
<img src="http://www.sciweavers.org/tex2img.php?eq=Y%20%3D%20f%5Cleft%28%5Csum_%7Bi%3D1%7D%5EN%20X_iW_i%5Cright%29&bc=White&fc=Black&im=jpg&fs=12&ff=arev&edit=0" align="center" border="0" alt="Y = f\left(\sum_{i=1}^N X_iW_i\right)" width="131" height="53" />
|
||||
<img src="images/netout.png" alt="Y = f\left(\sum_{i=1}^N X_iW_i\right)" width="131" height="53" align="center"/>
|
||||
|
||||
where f is some non-linear **activation function**.
|
||||
|
||||
|
|
|
|||
Binary file not shown.
|
After Width: | Height: | Size: 6.2 KiB |
|
|
@ -10,12 +10,16 @@ As you can see, VGG follows traditional pyramid architecture, which is a sequenc
|
|||
|
||||

|
||||
|
||||
> Image from [Researchgate](https://www.researchgate.net/figure/Vgg16-model-structure-To-get-the-VGG-NIN-model-we-replace-the-2-nd-4-th-6-th-7-th_fig2_335194493)
|
||||
|
||||
### ResNet
|
||||
|
||||
ResNet is a family of models proposed by Microsoft Research in 2015. The main idea of ResNet is to use **residual blocks**:
|
||||
|
||||
<img src="images/resnet-block.png" width="300"/>
|
||||
|
||||
> Image from [this paper](https://arxiv.org/pdf/1512.03385.pdf)
|
||||
|
||||
The reason for using identity pass-through is to have our layer predict **the difference** between the result of a previous layer and the output of the residual block - hence the name *residual*. Those blocks are much easier to train, and one can construct networks with several hundreds of those blocks (most common variants are ResNet-52, ResNet-101 and ResNet-152).
|
||||
|
||||
You can also think of this network as being able to adjust its complexity to the dataset. Initially, when you are starting to train the network, weights values are small, and most of the signal goes through passthrough identity layers. As training progresses and weights become larger, the significance of network parameters grow, and the networks adjusts to accommodate required expressive power to correctly classify training images.
|
||||
|
|
@ -26,6 +30,8 @@ Google Inception architecture takes this idea one step further, and builds each
|
|||
|
||||
<img src="images/inception.png" width="400"/>
|
||||
|
||||
> Image from [Researchgate](https://www.researchgate.net/figure/Inception-module-with-dimension-reductions-left-and-schema-for-Inception-ResNet-v1_fig2_355547454)
|
||||
|
||||
Here, we need to emphasize the role of 1x1 convolutions, because at first they do not make sense. Why would we need to run through the image with 1x1 filter? However, you need to remember that convolution filter also works with several depth channels (originally - RGB colors, in subsequent layers - channels for different filters), and 1x1 convolution is used to mix those input channels together using different trainable weights. It can be also viewed as downsampling (pooling) over channel dimension.
|
||||
|
||||
Here is [a good blog post](https://medium.com/analytics-vidhya/talented-mr-1x1-comprehensive-look-at-1x1-convolution-in-deep-learning-f6b355825578) on the subject, and [original paper](https://arxiv.org/pdf/1312.4400.pdf).
|
||||
|
|
|
|||
|
|
@ -13,6 +13,8 @@ For example, if we apply 3x3 vertical edge and horizontal edge filters to the MN
|
|||
|
||||
<img src="images/lmfilters.jpg" width="500" align="center"/>
|
||||
|
||||
> Image of Leung-Malik Filter Bank, from [here](https://www.robots.ox.ac.uk/~vgg/research/texclass/filters.html)
|
||||
|
||||
However, while we can design the filters to extract some patterns manually, we can also design the network in such a way that it will learn the patterns automatically. It is one of the main ideas behind the CNN.
|
||||
|
||||
## Main ideas behind CNN
|
||||
|
|
@ -24,6 +26,7 @@ The way CNNs work is based on the following important ideas:
|
|||
|
||||

|
||||
|
||||
> Image from [this paper](https://www.semanticscholar.org/paper/Computer-vision-based-pedestrian-trajectory-Hislop-Lynch/26e6f74853fc9bbb7487b06dc2cf095d36c9021d), based on [this research](https://dl.acm.org/doi/abs/10.1145/1553374.1553453)
|
||||
|
||||
## Continue in Notebook
|
||||
|
||||
|
|
@ -42,6 +45,8 @@ As an example, let's look at the architecture of VGG-16, a network that achieved
|
|||
|
||||

|
||||
|
||||
> Image from [Researchgate](https://www.researchgate.net/figure/Vgg16-model-structure-To-get-the-VGG-NIN-model-we-replace-the-2-nd-4-th-6-th-7-th_fig2_335194493)
|
||||
|
||||
[**Often Used CNN Architectures**](CNN_Architectures.md)
|
||||
|
||||
## CNNs for Other Tasks
|
||||
|
|
|
|||
|
|
@ -14,6 +14,10 @@ Both Keras and PyTorch contain functions to easily load pre-trained neural netwo
|
|||
* **ResNet** is a family of models proposed by Microsoft Research in 2015. They have more layers, and thus take more resources.
|
||||
* **MobileNet** is a family of models with reduced size, suitable for mobile devices. Use them if you are short in resources, and can sacrifice a little bit of accuracy.
|
||||
|
||||
Here are sample features extracted from a picture of a cat by VGG-16 network:
|
||||
|
||||

|
||||
|
||||
## Cats vs. Dogs Dataset
|
||||
|
||||
In this example, we will use a dataset of [Cats and Dogs](https://www.microsoft.com/en-us/download/details.aspx?id=54765&WT.mc_id=academic-33554-dmitryso), which is very close to a real-life image classification scenario.
|
||||
|
|
@ -24,3 +28,4 @@ Let's see transfer learning in action in corresponding notebooks:
|
|||
|
||||
* [Transfer Learning - PyTorch](TransferLearningPyTorch.ipynb)
|
||||
* [Transfer Learning - TensorFlow](TransferLearningTF.ipynb)
|
||||
|
||||
|
|
|
|||
Binary file not shown.
|
After Width: | Height: | Size: 9.2 KiB |
|
|
@ -8,7 +8,7 @@ Since we are training autoencoder to capture as much of the information from the
|
|||
|
||||

|
||||
|
||||
*Image from [Keras blog](https://blog.keras.io/building-autoencoders-in-keras.html)*
|
||||
> Image from [Keras blog](https://blog.keras.io/building-autoencoders-in-keras.html)
|
||||
|
||||
## Scenarios for using Autoencoders
|
||||
|
||||
|
|
@ -35,6 +35,8 @@ To summarize:
|
|||
|
||||
<img src="images/vae.png" width="50%">
|
||||
|
||||
> Image from [this blog post](https://ijdykeman.github.io/ml/2016/12/21/cvae.html) by Isaak Dykeman
|
||||
|
||||
Variational auto-encoders use complex loss function that consists of two parts:
|
||||
* **Reconstruction loss** is the loss function that shows how close reconstructed image is to the target (can be MSE). It is the same loss function as in normal autoencoders.
|
||||
* **KL loss**, which ensures that latent variable distributions stays close to normal distribution. It is based on the notion of [Kullback-Leibler divergence](https://www.countbayesie.com/blog/2017/5/9/kullback-leibler-divergence-explained) - a metric to estimate how similar two statistical distributions are.
|
||||
|
|
@ -43,10 +45,13 @@ One important advantage of VAEs is that they allow us to generate new images rel
|
|||
|
||||
<img src="images/vaemnist.png" width="50%"/>
|
||||
|
||||
> Image generated by author
|
||||
|
||||
Observe how images blend into each other, as we start getting latent vectors from the different portions of the latent parameter space. We can also visualize this space in 2D:
|
||||
|
||||
<img src="images/vaemnist-diag.png" width="50%"/>
|
||||
|
||||
> Image generated by author
|
||||
## Continue to Notebooks
|
||||
|
||||
* [Autoencoders in TensorFlow](AutoencodersTF.ipynb)
|
||||
|
|
@ -61,4 +66,5 @@ Observe how images blend into each other, as we start getting latent vectors fro
|
|||
|
||||
* [Building Autoencoders in Keras](https://blog.keras.io/building-autoencoders-in-keras.html)
|
||||
* [Blog post on NeuroHive](https://neurohive.io/ru/osnovy-data-science/variacionnyj-avtojenkoder-vae/)
|
||||
* [Variational Autoencoders Explained](https://kvfrans.com/variational-autoencoders-explained/)
|
||||
* [Variational Autoencoders Explained](https://kvfrans.com/variational-autoencoders-explained/)
|
||||
* [Conditional Variational Autoencoders](https://ijdykeman.github.io/ml/2016/12/21/cvae.html)
|
||||
|
|
@ -8,6 +8,8 @@ The main idea of GAN is to have two neural networks that will be trained against
|
|||
|
||||
<img src="images/gan_architecture.png" width="70%"/>
|
||||
|
||||
> Image by [Dmitry Soshnikov](http://soshnikov.com)
|
||||
|
||||
* **Generator** is a network that takes some random vector, and produces the image as a result
|
||||
* **Discriminator** is a network that takes an image, and it should tell whether it is a real image (from training dataset), or it was generated by a generator. It is essentially an image classifier.
|
||||
|
||||
|
|
@ -27,6 +29,8 @@ Generator is slightly more tricky. You can consider it to be a reversed discrimi
|
|||
|
||||
<img src="images/gan_arch_detail.png" width="70%"/>
|
||||
|
||||
> Image by [Dmitry Soshnikov](http://soshnikov.com)
|
||||
|
||||
### Training the GAN
|
||||
|
||||
GANs are called **adversarial** because there is a constant competition between generator and discriminator. During this cometition, both generator and discriminator improve, thus the network learns to produce better and better pictures.
|
||||
|
|
|
|||
|
|
@ -16,6 +16,8 @@ If we want to solve Natural Language Processing (NLP) tasks with neural networks
|
|||
|
||||
<img alt="Image showing diagram mapping a character to an ASCII and binary representation" src="images/ascii-character-map.png" width="50%"/>
|
||||
|
||||
> [Image source](https://www.seobility.net/en/wiki/ASCII)
|
||||
|
||||
We understand what each letter **represents**, and how all characters come together to form the words of a sentence. However, computers by themselves do not have such an understanding, and neural network has to learn the meaning during training.
|
||||
|
||||
Therefore, we can use different approaches when representing text:
|
||||
|
|
@ -36,6 +38,8 @@ When solving tasks like text classification, we need to be able to represent tex
|
|||
|
||||
<img src="images/bow.png" width="90%"/>
|
||||
|
||||
> Image by author
|
||||
|
||||
BOW essentially represents which words appear in text and in which quantities, which can indeed be a good indication of what the text is about. For example, news article on politics is likely to contains words such as *president* and *country*, while scientific publication would have something like *collider*, *discovered*, etc. Thus, word frequencies can in many cases be a good indicator of text content.
|
||||
|
||||
The problem with BOW is that certain common words, such as *and*, *is*, etc. appear in most of the texts, and they have highest frequencies, masking out the words that are really important. We may lower the importance of those words by taking into account the frequency at which words occur in the whole document collection. This is the main idea behind TF/IDF approach, which is covered in more detail in the notebooks below.
|
||||
|
|
|
|||
|
|
@ -10,6 +10,8 @@ By using embedding layer as a first layer in our classifier network, we can swit
|
|||
|
||||

|
||||
|
||||
> Image by author
|
||||
|
||||
## Continue in Notebooks
|
||||
|
||||
* [Embeddings with PyTorch](EmbeddingsPyTorch.ipynb)
|
||||
|
|
@ -28,6 +30,8 @@ CBoW is faster, while skip-gram is slower, but does a better job of representing
|
|||
|
||||

|
||||
|
||||
> Image from [this paper](https://arxiv.org/pdf/1301.3781.pdf)
|
||||
|
||||
Word2Vec pre-trained embeddings (as well as other similar models, such as GloVe) can also be used in place of embedding layer in neural networks. However, we need to deal with vocabularies, because the vocabulary used to pre-train Word2Vec/GloVe is likely to differ from the vocabulary in our text corpus. Have a look into Notebooks to see how this problem can be resolved.
|
||||
|
||||
## Contextual Embeddings
|
||||
|
|
@ -39,3 +43,7 @@ For example word 'play' in those two different sentences have quite different me
|
|||
- John wants to **play** with his friends.
|
||||
|
||||
The pretrained embeddings above represent both of these meanings of the word 'play' in the same embedding. To overcome this limitation, we need to build embeddings based on the **language model**, which is trained on a large corpus of text, and *knows* how words can be put together in different contexts. Discussing contextual embeddings is out of scope for this tutorial, but we will come back to them when talking about language models later in the course.
|
||||
|
||||
## References
|
||||
|
||||
* Paper on Word2Vec: [Efficient Estimation of Word Representations in Vector Space](https://arxiv.org/pdf/1301.3781.pdf)
|
||||
|
|
@ -11,6 +11,8 @@ In our previous examples, we have been using pre-trained semantic embeddings, bu
|
|||
|
||||

|
||||
|
||||
> Image from [this paper](https://arxiv.org/pdf/1301.3781.pdf)
|
||||
|
||||
The idea of CBoW is exactly predicting a missing word, however, to do this we take a small sliding window of text tokens (we can denote them from W<sub>-2</sub> to W<sub>2</sub>), and train a model to predict the central word W<sub>0</sub> from few surrounding words.
|
||||
|
||||
## More Info
|
||||
|
|
|
|||
|
|
@ -6,6 +6,8 @@ To capture the meaning of text sequence, we need to use another neural network a
|
|||
|
||||

|
||||
|
||||
> Image by author
|
||||
|
||||
Given the input sequence of tokens X<sub>0</sub>,...,X<sub>n</sub>, RNN creates a sequence of neural network blocks, and trains this sequence end-to-end using back propagation. Each network block takes a pair (X<sub>i</sub>,S<sub>i</sub>) as an input, and produces S<sub>i+1</sub> as a result. Final state S<sub>n</sub> or (output Y<sub>n</sub>) goes into a linear classifier to produce the result. All network blocks share the same weights, and are trained end-to-end using one backpropagation pass.
|
||||
|
||||
Because state vectors S<sub>0</sub>,...,S<sub>n</sub> are passed through the network, it is able to learn the sequential dependencies between words. For example, when the word *not* appears somewhere in the sequence, it can learn to negate certain elements within the state vector, resulting in negation.
|
||||
|
|
@ -20,6 +22,8 @@ Simple RNN cell has two weight matrices inside: one transforms input symbol (let
|
|||
|
||||
<img alt="RNN Cell Anatomy" src="images/rnn-anatomy.png" width="50%"/>
|
||||
|
||||
> Image by author
|
||||
|
||||
In many cases, input tokens are passed through the embedding layer before entering the RNN to lower the dimensionality. In this case, if the dimension of the input vectors is *emb_size*, and state vector is *hid_size* - the size of W is *emb_size*×*hid_size*, and the size of H is *hid_size*×*hid_size*.
|
||||
|
||||
## Long Short Term Memory (LSTM)
|
||||
|
|
@ -55,4 +59,3 @@ Recurrent network, one-directional or bidirectional, captures certain patterns w
|
|||
## RNNs for other tasks
|
||||
|
||||
In this unit, we have seen that RNNs can be used for sequence classification, but in fact, they can handle many more tasks, such as text generation, machine translation, and more. We will consider those tasks in the next unit.
|
||||
|
||||
|
|
|
|||
|
|
@ -7,7 +7,8 @@ In RNN architecture we discussed in the previous unit, each RNN unit produced ne
|
|||
This allows for different neural architectures that are shown in the picture below:
|
||||
|
||||

|
||||
*Image from blog post [Unreasonable Effectiveness of Recurrent Neural Networks](http://karpathy.github.io/2015/05/21/rnn-effectiveness/) by [Andrej Karpaty](http://karpathy.github.io/)*
|
||||
|
||||
> Image from blog post [Unreasonable Effectiveness of Recurrent Neural Networks](http://karpathy.github.io/2015/05/21/rnn-effectiveness/) by [Andrej Karpaty](http://karpathy.github.io/)
|
||||
|
||||
* **One-to-one** is a traditional neural network with one input and one output
|
||||
* **One-to-many** is a generative architecture that accepts one input value, and generates a sequence of output values. For example, if we want to train **image captioning** network that would produce a textual description of a picture, we can a picture as input, pass it through CNN to obtain hidden state, and then have recurrent chain generate caption word-by-word
|
||||
|
|
@ -22,8 +23,9 @@ The way we will train RNN to generate text is the following. On each step, we wi
|
|||
|
||||
When generating text (during inference), we start with some **prompt**, which is passed through RNN cells to generate intermediate state, and then from this state the generation starts. We generate one character at a time, and pass the state and the generated character to another RNN cell to generate the next one, until we generate enough characters.
|
||||
|
||||
<img src="images/rnn-generate-ing.png" width="60%"/>
|
||||
<img src="images/rnn-generate-inf.png" width="60%"/>
|
||||
|
||||
> Image by author
|
||||
## Continue to Notebooks
|
||||
|
||||
* [Generative Networks with PyTorch](GenerativePyTorch.ipynb)
|
||||
|
|
|
|||
|
|
@ -10,19 +10,20 @@ With RNNs, sequence-to-sequence is implemented by two recurrent networks, where
|
|||
**Attention Mechanisms** provide means of weighting the contextual impact of each input vector on each output prediction of the RNN. The way it is implemented is by creating shortcuts between intermediate states of the input RNN, and output RNN. In this manner, when generating output symbol y<sub>t</sub>, we will take into account all input hidden states h<sub>i</sub>, with different weight coefficients α<sub>t,i</sub>.
|
||||
|
||||

|
||||
*The encoder-decoder model with additive attention mechanism in [Bahdanau et al., 2015](https://arxiv.org/pdf/1409.0473.pdf), cited from [this blog post](https://lilianweng.github.io/lil-log/2018/06/24/attention-attention.html)*
|
||||
|
||||
> The encoder-decoder model with additive attention mechanism in [Bahdanau et al., 2015](https://arxiv.org/pdf/1409.0473.pdf), cited from [this blog post](https://lilianweng.github.io/lil-log/2018/06/24/attention-attention.html)
|
||||
|
||||
Attention matrix {α<sub>i,j</sub>} would represent the degree which certain input words play in generation of a given word in the output sequence. Below is the example of such a matrix:
|
||||
|
||||

|
||||
|
||||
*Figure taken from [Bahdanau et al., 2015](https://arxiv.org/pdf/1409.0473.pdf) (Fig.3)*
|
||||
> Figure from [Bahdanau et al., 2015](https://arxiv.org/pdf/1409.0473.pdf) (Fig.3)
|
||||
|
||||
Attention mechanisms are responsible for much of the current or near current state of the art in Natural language processing. Adding attention however greatly increases the number of model parameters which led to scaling issues with RNNs. A key constraint of scaling RNNs is that the recurrent nature of the models makes it challenging to batch and parallelize training. In an RNN each element of a sequence needs to be processed in sequential order which means it cannot be easily parallelized.
|
||||
|
||||

|
||||
|
||||
*Figure taken from [Google Blog](https://research.googleblog.com/2016/09/a-neural-network-for-machine.html)*
|
||||
> Figure from [Google Blog](https://research.googleblog.com/2016/09/a-neural-network-for-machine.html)
|
||||
|
||||
Adoption of attention mechanisms combined with this constraint led to the creation of the now State of the Art Transformer Models that we know and use today from BERT to Open-GPT3.
|
||||
|
||||
|
|
@ -44,6 +45,8 @@ We then mix the token position with token embedding vector. To transform positio
|
|||
|
||||
<img src="images/pos-embedding.png" width="50%"/>
|
||||
|
||||
> Image by author
|
||||
|
||||
The result we get with positional embedding embeds both original token and its position within sequence.
|
||||
|
||||
### Multi-Head Self-attention
|
||||
|
|
@ -51,7 +54,8 @@ The result we get with positional embedding embeds both original token and its p
|
|||
Next, we need to capture some patterns within our sequence. To do this, transformers use **self-attention** mechanism, which is essentially attention applied to the same sequence as input and output. Applying self-attention allows us to take into account **context** within the sentence, and see which words are inter-related. For example, it allows us to see which words are referred to by coreferences, such as *it*, and also take the context into account:
|
||||
|
||||

|
||||
*Image from the [Google Blog](https://research.googleblog.com/2017/08/transformer-novel-neural-network.html)*
|
||||
|
||||
> Image from the [Google Blog](https://research.googleblog.com/2017/08/transformer-novel-neural-network.html)
|
||||
|
||||
In transformers, we use **Multi-Head Attention**, in order to give network the power to capture several different types of dependencies, eg. long-term vs. short-term word relations, co-reference vs. something else, etc.
|
||||
|
||||
|
|
@ -75,6 +79,8 @@ Since each input position is mapped independently to each output position, trans
|
|||
|
||||

|
||||
|
||||
> Image [source](http://jalammar.github.io/illustrated-bert/)
|
||||
|
||||
There are many variations of Transformer architectures including BERT, DistilBERT. BigBird, OpenGPT3 and more that can be fine tuned. The [HuggingFace package](https://github.com/huggingface/) provides repository for training many of these architectures with both PyTorch and TensorFlow.
|
||||
|
||||
## Continue to Notebooks
|
||||
|
|
|
|||
|
|
@ -48,7 +48,7 @@ For a gentle introduction to *AI in the Cloud* topic you may consider taking the
|
|||
<tr><td>3</td><td>Perceptron</td>
|
||||
<td><a href="3-NeuralNetworks/03-Perceptron/README.md">Text</a>
|
||||
<td colspan="2"><a href="3-NeuralNetworks/03-Perceptron/Perceptron.ipynb">Notebook</a></td><td><a href="3-NeuralNetworks/03-Perceptron/lab/README.md">Lab</a></td></tr>
|
||||
<tr><td>4 </td><td>Multi-Layered Perceptron and Creating our own Framework</td><td><a href="3-NeuralNetworks/04-OwnFramework/README.md">Text</a></td><td colspan="2"><a href="3-NeuralNetworks/04-OwnFramework/OwnFramework.ipynb">Notebook</a><td></td></tr>
|
||||
<tr><td>4 </td><td>Multi-Layered Perceptron and Creating our own Framework</td><td><a href="3-NeuralNetworks/04-OwnFramework/README.md">Text</a></td><td colspan="2"><a href="3-NeuralNetworks/04-OwnFramework/OwnFramework.ipynb">Notebook</a><td><a href="3-NeuralNetworks/04-OwnFramework/lab/README.md">Lab</a></td></tr>
|
||||
<tr><td>5</td>
|
||||
<td>Intro to Frameworks (PyTorch/TensorFlow)<br/>Overfitting</td>
|
||||
<td><a href="3-NeuralNetworks/05-Frameworks/README.md">Text</a><br/><a href="3-NeuralNetworks/05-Frameworks/Overfitting.md">Text</a></td>
|
||||
|
|
|
|||
Loading…
Reference in New Issue