Merge pull request #514 from microsoft/update-translations
🌐 Update translations via Co-op Translator
This commit is contained in:
commit
338be1a743
|
|
@ -1,8 +1,8 @@
|
|||
<!--
|
||||
CO_OP_TRANSLATOR_METADATA:
|
||||
{
|
||||
"original_hash": "f3a6b0ddf7e6e3f33b2a543baf086dc9",
|
||||
"translation_date": "2025-08-24T20:43:36+00:00",
|
||||
"original_hash": "07191303b7ea2aff1d47e2b0fe4bb862",
|
||||
"translation_date": "2025-08-31T14:18:57+00:00",
|
||||
"source_file": "README.md",
|
||||
"language_code": "fr"
|
||||
}
|
||||
|
|
@ -23,86 +23,97 @@ CO_OP_TRANSLATOR_METADATA:
|
|||
|
||||
# Intelligence Artificielle pour Débutants - Un Programme
|
||||
|
||||
| ](./lessons/sketchnotes/ai-overview.png)|
|
||||
|:---:|
|
||||
| AI For Beginners - _Sketchnote par [@girlie_mac](https://twitter.com/girlie_mac)_ |
|
||||
||
|
||||
|:---:|
|
||||
| AI pour Débutants - _Sketchnote par [@girlie_mac](https://twitter.com/girlie_mac)_ |
|
||||
|
||||
Explorez le monde de l'**Intelligence Artificielle** (IA) avec notre programme de 12 semaines et 24 leçons ! Il inclut des leçons pratiques, des quiz et des laboratoires. Ce programme est adapté aux débutants et couvre des outils comme TensorFlow et PyTorch, ainsi que des questions d'éthique en IA.
|
||||
Explorez le monde de l'**Intelligence Artificielle** (IA) avec notre programme de 12 semaines et 24 leçons ! Il inclut des leçons pratiques, des quiz et des laboratoires. Ce programme est adapté aux débutants et couvre des outils comme TensorFlow et PyTorch, ainsi que des questions d'éthique en IA.
|
||||
|
||||
## Ce que vous apprendrez
|
||||
### 🌐 Support Multilingue
|
||||
|
||||
**[Carte mentale du cours](http://soshnikov.com/courses/ai-for-beginners/mindmap.html)**
|
||||
#### Supporté via GitHub Action (Automatisé & Toujours à Jour)
|
||||
|
||||
Dans ce programme, vous apprendrez :
|
||||
[Français](./README.md) | [Espagnol](../es/README.md) | [Allemand](../de/README.md) | [Russe](../ru/README.md) | [Arabe](../ar/README.md) | [Persan (Farsi)](../fa/README.md) | [Ourdou](../ur/README.md) | [Chinois (Simplifié)](../zh/README.md) | [Chinois (Traditionnel, Macao)](../mo/README.md) | [Chinois (Traditionnel, Hong Kong)](../hk/README.md) | [Chinois (Traditionnel, Taïwan)](../tw/README.md) | [Japonais](../ja/README.md) | [Coréen](../ko/README.md) | [Hindi](../hi/README.md) | [Bengali](../bn/README.md) | [Marathi](../mr/README.md) | [Népalais](../ne/README.md) | [Punjabi (Gurmukhi)](../pa/README.md) | [Portugais (Portugal)](../pt/README.md) | [Portugais (Brésil)](../br/README.md) | [Italien](../it/README.md) | [Polonais](../pl/README.md) | [Turc](../tr/README.md) | [Grec](../el/README.md) | [Thaï](../th/README.md) | [Suédois](../sv/README.md) | [Danois](../da/README.md) | [Norvégien](../no/README.md) | [Finnois](../fi/README.md) | [Néerlandais](../nl/README.md) | [Hébreu](../he/README.md) | [Vietnamien](../vi/README.md) | [Indonésien](../id/README.md) | [Malais](../ms/README.md) | [Tagalog (Filipino)](../tl/README.md) | [Swahili](../sw/README.md) | [Hongrois](../hu/README.md) | [Tchèque](../cs/README.md) | [Slovaque](../sk/README.md) | [Roumain](../ro/README.md) | [Bulgare](../bg/README.md) | [Serbe (Cyrillique)](../sr/README.md) | [Croate](../hr/README.md) | [Slovène](../sl/README.md) | [Ukrainien](../uk/README.md) | [Birman (Myanmar)](../my/README.md)
|
||||
|
||||
* Différentes approches de l'Intelligence Artificielle, y compris l'approche symbolique "classique" avec la **Représentation des Connaissances** et le raisonnement ([GOFAI](https://en.wikipedia.org/wiki/Symbolic_artificial_intelligence)).
|
||||
* Les **Réseaux Neuronaux** et le **Deep Learning**, qui sont au cœur de l'IA moderne. Nous illustrerons les concepts derrière ces sujets importants avec du code dans deux des frameworks les plus populaires - [TensorFlow](http://Tensorflow.org) et [PyTorch](http://pytorch.org).
|
||||
* Les **Architectures Neuronales** pour travailler avec les images et le texte. Nous couvrirons des modèles récents, mais il se peut que nous ne soyons pas totalement à jour avec les dernières avancées.
|
||||
* Des approches moins populaires de l'IA, comme les **Algorithmes Génétiques** et les **Systèmes Multi-Agents**.
|
||||
**Si vous souhaitez ajouter des langues supplémentaires, les langues supportées sont listées [ici](https://github.com/Azure/co-op-translator/blob/main/getting_started/supported-languages.md)**
|
||||
|
||||
Ce que nous ne couvrirons pas dans ce programme :
|
||||
## Rejoignez la Communauté
|
||||
[](https://discord.gg/kzRShWzttr)
|
||||
|
||||
> [Retrouvez toutes les ressources supplémentaires pour ce cours dans notre collection Microsoft Learn](https://learn.microsoft.com/en-us/collections/7w28iy2xrqzdj0?WT.mc_id=academic-77998-bethanycheum)
|
||||
## Ce que vous apprendrez
|
||||
|
||||
* Les cas d'utilisation de l'**IA en entreprise**. Pensez à suivre le parcours d'apprentissage [Introduction à l'IA pour les utilisateurs professionnels](https://docs.microsoft.com/learn/paths/introduction-ai-for-business-users/?WT.mc_id=academic-77998-bethanycheum) sur Microsoft Learn, ou [AI Business School](https://www.microsoft.com/ai/ai-business-school/?WT.mc_id=academic-77998-bethanycheum), développé en collaboration avec [INSEAD](https://www.insead.edu/).
|
||||
* Le **Machine Learning classique**, qui est bien décrit dans notre programme [Machine Learning pour Débutants](http://github.com/Microsoft/ML-for-Beginners).
|
||||
* Les applications pratiques d'IA construites avec **[Cognitive Services](https://azure.microsoft.com/services/cognitive-services/?WT.mc_id=academic-77998-bethanycheum)**. Pour cela, nous recommandons de commencer par les modules Microsoft Learn pour [vision](https://docs.microsoft.com/learn/paths/create-computer-vision-solutions-azure-cognitive-services/?WT.mc_id=academic-77998-bethanycheum), [traitement du langage naturel](https://docs.microsoft.com/learn/paths/explore-natural-language-processing/?WT.mc_id=academic-77998-bethanycheum), **[IA générative avec Azure OpenAI Service](https://learn.microsoft.com/en-us/training/paths/develop-ai-solutions-azure-openai/?WT.mc_id=academic-77998-bethanycheum)** et autres.
|
||||
* Les **Frameworks Cloud ML** spécifiques, comme [Azure Machine Learning](https://azure.microsoft.com/services/machine-learning/?WT.mc_id=academic-77998-bethanycheum), [Microsoft Fabric](https://learn.microsoft.com/en-us/training/paths/get-started-fabric/?WT.mc_id=academic-77998-bethanycheum), ou [Azure Databricks](https://docs.microsoft.com/learn/paths/data-engineer-azure-databricks?WT.mc_id=academic-77998-bethanycheum). Pensez à utiliser les parcours d'apprentissage [Créer et exploiter des solutions de machine learning avec Azure Machine Learning](https://docs.microsoft.com/learn/paths/build-ai-solutions-with-azure-ml-service/?WT.mc_id=academic-77998-bethanycheum) et [Créer et exploiter des solutions de machine learning avec Azure Databricks](https://docs.microsoft.com/learn/paths/build-operate-machine-learning-solutions-azure-databricks/?WT.mc_id=academic-77998-bethanycheum).
|
||||
* L'**IA conversationnelle** et les **Chat Bots**. Il existe un parcours d'apprentissage séparé [Créer des solutions d'IA conversationnelle](https://docs.microsoft.com/learn/paths/create-conversational-ai-solutions/?WT.mc_id=academic-77998-bethanycheum), et vous pouvez également consulter [cet article de blog](https://soshnikov.com/azure/hello-bot-conversational-ai-on-microsoft-platform/) pour plus de détails.
|
||||
* Les **Mathématiques avancées** derrière le deep learning. Pour cela, nous recommandons [Deep Learning](https://www.amazon.com/Deep-Learning-Adaptive-Computation-Machine/dp/0262035618) par Ian Goodfellow, Yoshua Bengio et Aaron Courville, également disponible en ligne à [https://www.deeplearningbook.org/](https://www.deeplearningbook.org/).
|
||||
**[Carte mentale du cours](http://soshnikov.com/courses/ai-for-beginners/mindmap.html)**
|
||||
|
||||
Pour une introduction douce aux sujets liés à l'_IA dans le Cloud_, vous pouvez envisager de suivre le parcours d'apprentissage [Commencer avec l'intelligence artificielle sur Azure](https://docs.microsoft.com/learn/paths/get-started-with-artificial-intelligence-on-azure/?WT.mc_id=academic-77998-bethanycheum).
|
||||
Dans ce programme, vous apprendrez :
|
||||
|
||||
# Contenu
|
||||
* Différentes approches de l'Intelligence Artificielle, y compris l'approche symbolique "classique" avec la **Représentation des Connaissances** et le raisonnement ([GOFAI](https://en.wikipedia.org/wiki/Symbolic_artificial_intelligence)).
|
||||
* Les **Réseaux Neuronaux** et le **Deep Learning**, qui sont au cœur de l'IA moderne. Nous illustrerons les concepts derrière ces sujets importants avec du code dans deux des frameworks les plus populaires - [TensorFlow](http://Tensorflow.org) et [PyTorch](http://pytorch.org).
|
||||
* Les **Architectures Neuronales** pour travailler avec les images et le texte. Nous couvrirons des modèles récents, mais il se peut que nous manquions un peu des derniers modèles à la pointe.
|
||||
* Des approches moins populaires de l'IA, comme les **Algorithmes Génétiques** et les **Systèmes Multi-Agents**.
|
||||
|
||||
| | Lien de la leçon | PyTorch/Keras/TensorFlow | Lab |
|
||||
| :-: | :------------------------------------------------------------------------------------------------------------------------------------------: | :---------------------------------------------------------------------------------------------: | ------------------------------------------------------------------------------ |
|
||||
| 0 | [Configuration du cours](./lessons/0-course-setup/setup.md) | [Configurer votre environnement de développement](./lessons/0-course-setup/how-to-run.md) | |
|
||||
| I | [**Introduction à l'IA**](./lessons/1-Intro/README.md) | | |
|
||||
| 01 | [Introduction et histoire de l'IA](./lessons/1-Intro/README.md) | - | - |
|
||||
| II | **IA Symbolique** |
|
||||
| 02 | [Représentation des connaissances et systèmes experts](./lessons/2-Symbolic/README.md) | [Systèmes experts](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/2-Symbolic/Animals.ipynb) / [Ontologie](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/2-Symbolic/FamilyOntology.ipynb) /[Graphes de concepts](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/2-Symbolic/MSConceptGraph.ipynb) | |
|
||||
| III | [**Introduction aux réseaux neuronaux**](./lessons/3-NeuralNetworks/README.md) |||
|
||||
| 03 | [Perceptron](./lessons/3-NeuralNetworks/03-Perceptron/README.md) | [Notebook](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/3-NeuralNetworks/03-Perceptron/Perceptron.ipynb) | [Lab](./lessons/3-NeuralNetworks/03-Perceptron/lab/README.md) |
|
||||
| 04 | [Perceptron multicouche et création de notre propre framework](./lessons/3-NeuralNetworks/04-OwnFramework/README.md) | [Notebook](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/3-NeuralNetworks/04-OwnFramework/OwnFramework.ipynb) | [Lab](./lessons/3-NeuralNetworks/04-OwnFramework/lab/README.md) |
|
||||
| 05 | [Introduction aux frameworks (PyTorch/TensorFlow) et surapprentissage](./lessons/3-NeuralNetworks/05-Frameworks/README.md) | [PyTorch](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/3-NeuralNetworks/05-Frameworks/IntroPyTorch.ipynb) / [Keras](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/3-NeuralNetworks/05-Frameworks/IntroKeras.ipynb) / [TensorFlow](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/3-NeuralNetworks/05-Frameworks/IntroKerasTF.ipynb) | [Lab](./lessons/3-NeuralNetworks/05-Frameworks/lab/README.md) |
|
||||
Ce que nous ne couvrirons pas dans ce programme :
|
||||
|
||||
> [Retrouvez toutes les ressources supplémentaires pour ce cours dans notre collection Microsoft Learn](https://learn.microsoft.com/en-us/collections/7w28iy2xrqzdj0?WT.mc_id=academic-77998-bethanycheum)
|
||||
|
||||
* Les cas d'utilisation de l'**IA en Entreprise**. Pensez à suivre le parcours d'apprentissage [Introduction à l'IA pour les utilisateurs professionnels](https://docs.microsoft.com/learn/paths/introduction-ai-for-business-users/?WT.mc_id=academic-77998-bethanycheum) sur Microsoft Learn, ou [AI Business School](https://www.microsoft.com/ai/ai-business-school/?WT.mc_id=academic-77998-bethanycheum), développé en coopération avec [INSEAD](https://www.insead.edu/).
|
||||
* Le **Machine Learning Classique**, qui est bien décrit dans notre programme [Machine Learning pour Débutants](http://github.com/Microsoft/ML-for-Beginners).
|
||||
* Les applications pratiques de l'IA construites avec **[Cognitive Services](https://azure.microsoft.com/services/cognitive-services/?WT.mc_id=academic-77998-bethanycheum)**. Pour cela, nous vous recommandons de commencer par les modules Microsoft Learn pour [vision](https://docs.microsoft.com/learn/paths/create-computer-vision-solutions-azure-cognitive-services/?WT.mc_id=academic-77998-bethanycheum), [traitement du langage naturel](https://docs.microsoft.com/learn/paths/explore-natural-language-processing/?WT.mc_id=academic-77998-bethanycheum), **[IA Générative avec Azure OpenAI Service](https://learn.microsoft.com/en-us/training/paths/develop-ai-solutions-azure-openai/?WT.mc_id=academic-77998-bethanycheum)** et autres.
|
||||
* Les **Frameworks Cloud ML** spécifiques, comme [Azure Machine Learning](https://azure.microsoft.com/services/machine-learning/?WT.mc_id=academic-77998-bethanycheum), [Microsoft Fabric](https://learn.microsoft.com/en-us/training/paths/get-started-fabric/?WT.mc_id=academic-77998-bethanycheum), ou [Azure Databricks](https://docs.microsoft.com/learn/paths/data-engineer-azure-databricks?WT.mc_id=academic-77998-bethanycheum). Pensez à utiliser les parcours d'apprentissage [Créer et exploiter des solutions de machine learning avec Azure Machine Learning](https://docs.microsoft.com/learn/paths/build-ai-solutions-with-azure-ml-service/?WT.mc_id=academic-77998-bethanycheum) et [Créer et exploiter des solutions de machine learning avec Azure Databricks](https://docs.microsoft.com/learn/paths/build-operate-machine-learning-solutions-azure-databricks/?WT.mc_id=academic-77998-bethanycheum).
|
||||
* L'**IA Conversationnelle** et les **Chat Bots**. Il existe un parcours d'apprentissage séparé [Créer des solutions d'IA conversationnelle](https://docs.microsoft.com/learn/paths/create-conversational-ai-solutions/?WT.mc_id=academic-77998-bethanycheum), et vous pouvez également consulter [cet article de blog](https://soshnikov.com/azure/hello-bot-conversational-ai-on-microsoft-platform/) pour plus de détails.
|
||||
* Les **Mathématiques Approfondies** derrière le deep learning. Pour cela, nous vous recommandons [Deep Learning](https://www.amazon.com/Deep-Learning-Adaptive-Computation-Machine/dp/0262035618) par Ian Goodfellow, Yoshua Bengio et Aaron Courville, également disponible en ligne à [https://www.deeplearningbook.org/](https://www.deeplearningbook.org/).
|
||||
|
||||
Pour une introduction douce aux sujets _IA dans le Cloud_, vous pouvez envisager de suivre le parcours d'apprentissage [Commencer avec l'intelligence artificielle sur Azure](https://docs.microsoft.com/learn/paths/get-started-with-artificial-intelligence-on-azure/?WT.mc_id=academic-77998-bethanycheum).
|
||||
|
||||
# Contenu
|
||||
|
||||
| | Lien de la Leçon | PyTorch/Keras/TensorFlow | Lab |
|
||||
| :-: | :------------------------------------------------------------------------------------------------------------------------------------------: | :---------------------------------------------------------------------------------------------: | ------------------------------------------------------------------------------ |
|
||||
| 0 | [Configuration du Cours](./lessons/0-course-setup/setup.md) | [Configurer votre environnement de développement](./lessons/0-course-setup/how-to-run.md) | |
|
||||
| I | [**Introduction à l'IA**](./lessons/1-Intro/README.md) | | |
|
||||
| 01 | [Introduction et Histoire de l'IA](./lessons/1-Intro/README.md) | - | - |
|
||||
| II | **IA Symbolique** |
|
||||
| 02 | [Représentation des Connaissances et Systèmes Experts](./lessons/2-Symbolic/README.md) | [Systèmes Experts](./lessons/2-Symbolic/Animals.ipynb) / [Ontologie](./lessons/2-Symbolic/FamilyOntology.ipynb) /[Graphique Conceptuel](./lessons/2-Symbolic/MSConceptGraph.ipynb) | |
|
||||
| III | [**Introduction aux Réseaux Neuronaux**](./lessons/3-NeuralNetworks/README.md) |||
|
||||
| 03 | [Perceptron](./lessons/3-NeuralNetworks/03-Perceptron/README.md) | [Notebook](./lessons/3-NeuralNetworks/03-Perceptron/Perceptron.ipynb) | [Lab](./lessons/3-NeuralNetworks/03-Perceptron/lab/README.md) |
|
||||
| 04 | [Perceptron Multicouche et Création de notre propre Framework](./lessons/3-NeuralNetworks/04-OwnFramework/README.md) | [Notebook](./lessons/3-NeuralNetworks/04-OwnFramework/OwnFramework.ipynb) | [Lab](./lessons/3-NeuralNetworks/04-OwnFramework/lab/README.md) |
|
||||
| 05 | [Introduction aux frameworks (PyTorch/TensorFlow) et surapprentissage](./lessons/3-NeuralNetworks/05-Frameworks/README.md) | [PyTorch](./lessons/3-NeuralNetworks/05-Frameworks/IntroPyTorch.ipynb) / [Keras](./lessons/3-NeuralNetworks/05-Frameworks/IntroKeras.ipynb) / [TensorFlow](./lessons/3-NeuralNetworks/05-Frameworks/IntroKerasTF.ipynb) | [Lab](./lessons/3-NeuralNetworks/05-Frameworks/lab/README.md) |
|
||||
| IV | [**Vision par ordinateur**](./lessons/4-ComputerVision/README.md) | [PyTorch](https://docs.microsoft.com/learn/modules/intro-computer-vision-pytorch/?WT.mc_id=academic-77998-cacaste) / [TensorFlow](https://docs.microsoft.com/learn/modules/intro-computer-vision-TensorFlow/?WT.mc_id=academic-77998-cacaste)| [Explorer la vision par ordinateur sur Microsoft Azure](https://learn.microsoft.com/en-us/collections/7w28iy2xrqzdj0?WT.mc_id=academic-77998-bethanycheum) |
|
||||
| 06 | [Introduction à la vision par ordinateur. OpenCV](./lessons/4-ComputerVision/06-IntroCV/README.md) | [Notebook](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/4-ComputerVision/06-IntroCV/OpenCV.ipynb) | [Lab](./lessons/4-ComputerVision/06-IntroCV/lab/README.md) |
|
||||
| 07 | [Réseaux neuronaux convolutionnels](./lessons/4-ComputerVision/07-ConvNets/README.md) & [Architectures CNN](./lessons/4-ComputerVision/07-ConvNets/CNN_Architectures.md) | [PyTorch](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/4-ComputerVision/07-ConvNets/ConvNetsPyTorch.ipynb) /[TensorFlow](https://microsoft.github.io/AI-For-Beginners/lessons/4-ComputerVision/07-ConvNets/ConvNetsTF.ipynb) | [Lab](./lessons/4-ComputerVision/07-ConvNets/lab/README.md) |
|
||||
| 08 | [Réseaux pré-entraînés et apprentissage par transfert](./lessons/4-ComputerVision/08-TransferLearning/README.md) et [Astuces d'entraînement](./lessons/4-ComputerVision/08-TransferLearning/TrainingTricks.md) | [PyTorch](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/4-ComputerVision/08-TransferLearning/TransferLearningPyTorch.ipynb) / [TensorFlow](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/3-NeuralNetworks/05-Frameworks/IntroKerasTF.ipynb) | [Lab](./lessons/4-ComputerVision/08-TransferLearning/lab/README.md) |
|
||||
| 09 | [Autoencodeurs et VAEs](./lessons/4-ComputerVision/09-Autoencoders/README.md) | [PyTorch](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/4-ComputerVision/09-Autoencoders/AutoEncodersPyTorch.ipynb) / [TensorFlow](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/4-ComputerVision/09-Autoencoders/AutoencodersTF.ipynb) | |
|
||||
| 10 | [Réseaux antagonistes génératifs et transfert de style artistique](./lessons/4-ComputerVision/10-GANs/README.md) | [PyTorch](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/4-ComputerVision/10-GANs/GANPyTorch.ipynb) / [TensorFlow](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/4-ComputerVision/10-GANs/GANTF.ipynb) | |
|
||||
| 11 | [Détection d'objets](./lessons/4-ComputerVision/11-ObjectDetection/README.md) | [TensorFlow](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/4-ComputerVision/11-ObjectDetection/ObjectDetection.ipynb) | [Lab](./lessons/4-ComputerVision/11-ObjectDetection/lab/README.md) |
|
||||
| 12 | [Segmentation sémantique. U-Net](./lessons/4-ComputerVision/12-Segmentation/README.md) | [PyTorch](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/4-ComputerVision/12-Segmentation/SemanticSegmentationPytorch.ipynb) / [TensorFlow](../../(https:/github.com/microsoft/AI-For-Beginners/blob/main/lessons/4-ComputerVision/12-Segmentation/SemanticSegmentationTF.ipynb)) | |
|
||||
| 06 | [Introduction à la vision par ordinateur. OpenCV](./lessons/4-ComputerVision/06-IntroCV/README.md) | [Notebook](./lessons/4-ComputerVision/06-IntroCV/OpenCV.ipynb) | [Lab](./lessons/4-ComputerVision/06-IntroCV/lab/README.md) |
|
||||
| 07 | [Réseaux neuronaux convolutionnels](./lessons/4-ComputerVision/07-ConvNets/README.md) & [Architectures CNN](./lessons/4-ComputerVision/07-ConvNets/CNN_Architectures.md) | [PyTorch](./lessons/4-ComputerVision/07-ConvNets/ConvNetsPyTorch.ipynb) /[TensorFlow](./lessons/4-ComputerVision/07-ConvNets/ConvNetsTF.ipynb) | [Lab](./lessons/4-ComputerVision/07-ConvNets/lab/README.md) |
|
||||
| 08 | [Réseaux pré-entraînés et apprentissage par transfert](./lessons/4-ComputerVision/08-TransferLearning/README.md) et [Astuces d'entraînement](./lessons/4-ComputerVision/08-TransferLearning/TrainingTricks.md) | [PyTorch](./lessons/4-ComputerVision/08-TransferLearning/TransferLearningPyTorch.ipynb) / [TensorFlow](./lessons/3-NeuralNetworks/05-Frameworks/IntroKerasTF.ipynb) | [Lab](./lessons/4-ComputerVision/08-TransferLearning/lab/README.md) |
|
||||
| 09 | [Autoencodeurs et VAEs](./lessons/4-ComputerVision/09-Autoencoders/README.md) | [PyTorch](./lessons/4-ComputerVision/09-Autoencoders/AutoEncodersPyTorch.ipynb) / [TensorFlow](./lessons/4-ComputerVision/09-Autoencoders/AutoencodersTF.ipynb) | |
|
||||
| 10 | [Réseaux antagonistes génératifs et transfert de style artistique](./lessons/4-ComputerVision/10-GANs/README.md) | [PyTorch](./lessons/4-ComputerVision/10-GANs/GANPyTorch.ipynb) / [TensorFlow](./lessons/4-ComputerVision/10-GANs/GANTF.ipynb) | |
|
||||
| 11 | [Détection d'objets](./lessons/4-ComputerVision/11-ObjectDetection/README.md) | [TensorFlow](./lessons/4-ComputerVision/11-ObjectDetection/ObjectDetection.ipynb) | [Lab](./lessons/4-ComputerVision/11-ObjectDetection/lab/README.md) |
|
||||
| 12 | [Segmentation sémantique. U-Net](./lessons/4-ComputerVision/12-Segmentation/README.md) | [PyTorch](./lessons/4-ComputerVision/12-Segmentation/SemanticSegmentationPytorch.ipynb) / [TensorFlow](./lessons/4-ComputerVision/12-Segmentation/SemanticSegmentationTF.ipynb) | |
|
||||
| V | [**Traitement du langage naturel**](./lessons/5-NLP/README.md) | [PyTorch](https://docs.microsoft.com/learn/modules/intro-natural-language-processing-pytorch/?WT.mc_id=academic-77998-cacaste) /[TensorFlow](https://docs.microsoft.com/learn/modules/intro-natural-language-processing-TensorFlow/?WT.mc_id=academic-77998-cacaste) | [Explorer le traitement du langage naturel sur Microsoft Azure](https://learn.microsoft.com/en-us/collections/7w28iy2xrqzdj0?WT.mc_id=academic-77998-bethanycheum)|
|
||||
| 13 | [Représentation des textes. Bow/TF-IDF](./lessons/5-NLP/13-TextRep/README.md) | [PyTorch](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/5-NLP/13-TextRep/TextRepresentationPyTorch.ipynb) / [TensorFlow](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/5-NLP/13-TextRep/TextRepresentationTF.ipynb) | |
|
||||
| 14 | [Word embeddings sémantiques. Word2Vec et GloVe](./lessons/5-NLP/14-Embeddings/README.md) | [PyTorch](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/5-NLP/14-Embeddings/EmbeddingsPyTorch.ipynb) / [TensorFlow](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/5-NLP/14-Embeddings/EmbeddingsTF.ipynb) | |
|
||||
| 15 | [Modélisation du langage. Entraîner vos propres embeddings](./lessons/5-NLP/15-LanguageModeling/README.md) | [PyTorch](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/5-NLP/15-LanguageModeling/CBoW-PyTorch.ipynb) / [TensorFlow](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/5-NLP/15-LanguageModeling/CBoW-TF.ipynb) | [Lab](./lessons/5-NLP/15-LanguageModeling/lab/README.md) |
|
||||
| 16 | [Réseaux neuronaux récurrents](./lessons/5-NLP/16-RNN/README.md) | [PyTorch](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/5-NLP/16-RNN/RNNPyTorch.ipynb) / [TensorFlow](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/5-NLP/16-RNN/RNNTF.ipynb) | |
|
||||
| 17 | [Réseaux récurrents génératifs](./lessons/5-NLP/17-GenerativeNetworks/README.md) | [PyTorch](https://microsoft.github.io/AI-For-Beginners/lessons/5-NLP/17-GenerativeNetworks/GenerativePyTorch.md) / [TensorFlow](https://microsoft.github.io/AI-For-Beginners/lessons/5-NLP/17-GenerativeNetworks/GenerativeTF.md) | [Lab](./lessons/5-NLP/17-GenerativeNetworks/lab/README.md) |
|
||||
| 18 | [Transformers. BERT.](./lessons/5-NLP/18-Transformers/READMEtransformers.md) | [PyTorch](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/5-NLP/18-Transformers/TransformersPyTorch.ipynb) /[TensorFlow](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/5-NLP/18-Transformers/TransformersTF.ipynb) | |
|
||||
| 19 | [Reconnaissance d'entités nommées](./lessons/5-NLP/19-NER/README.md) | [TensorFlow](https://microsoft.github.io/AI-For-Beginners/lessons/5-NLP/19-NER/NER-TF.ipynb) | [Lab](./lessons/5-NLP/19-NER/lab/README.md) |
|
||||
| 20 | [Grands modèles de langage, programmation par prompts et tâches few-shot](./lessons/5-NLP/20-LangModels/READMELargeLang.md) | [PyTorch](https://microsoft.github.io/AI-For-Beginners/lessons/5-NLP/20-LangModels/GPT-PyTorch.ipynb) | |
|
||||
| 13 | [Représentation de texte. Bow/TF-IDF](./lessons/5-NLP/13-TextRep/README.md) | [PyTorch](./lessons/5-NLP/13-TextRep/TextRepresentationPyTorch.ipynb) / [TensorFlow](./lessons/5-NLP/13-TextRep/TextRepresentationTF.ipynb) | |
|
||||
| 14 | [Embeddings sémantiques de mots. Word2Vec et GloVe](./lessons/5-NLP/14-Embeddings/README.md) | [PyTorch](./lessons/5-NLP/14-Embeddings/EmbeddingsPyTorch.ipynb) / [TensorFlow](./lessons/5-NLP/14-Embeddings/EmbeddingsTF.ipynb) | |
|
||||
| 15 | [Modélisation de langage. Entraînez vos propres embeddings](./lessons/5-NLP/15-LanguageModeling/README.md) | [PyTorch](./lessons/5-NLP/15-LanguageModeling/CBoW-PyTorch.ipynb) / [TensorFlow](./lessons/5-NLP/15-LanguageModeling/CBoW-TF.ipynb) | [Lab](./lessons/5-NLP/15-LanguageModeling/lab/README.md) |
|
||||
| 16 | [Réseaux neuronaux récurrents](./lessons/5-NLP/16-RNN/README.md) | [PyTorch](./lessons/5-NLP/16-RNN/RNNPyTorch.ipynb) / [TensorFlow](./lessons/5-NLP/16-RNN/RNNTF.ipynb) | |
|
||||
| 17 | [Réseaux récurrents génératifs](./lessons/5-NLP/17-GenerativeNetworks/README.md) | [PyTorch](./lessons/5-NLP/17-GenerativeNetworks/GenerativePyTorch.md) / [TensorFlow](./lessons/5-NLP/17-GenerativeNetworks/GenerativeTF.md) | [Lab](./lessons/5-NLP/17-GenerativeNetworks/lab/README.md) |
|
||||
| 18 | [Transformers. BERT.](./lessons/5-NLP/18-Transformers/READMEtransformers.md) | [PyTorch](./lessons/5-NLP/18-Transformers/TransformersPyTorch.ipynb) /[TensorFlow](./lessons/5-NLP/18-Transformers/TransformersTF.ipynb) | |
|
||||
| 19 | [Reconnaissance d'entités nommées](./lessons/5-NLP/19-NER/README.md) | [TensorFlow](./lessons/5-NLP/19-NER/NER-TF.ipynb) | [Lab](./lessons/5-NLP/19-NER/lab/README.md) |
|
||||
| 20 | [Grands modèles de langage, programmation par prompts et tâches en apprentissage par petits échantillons](./lessons/5-NLP/20-LangModels/READMELargeLang.md) | [PyTorch](./lessons/5-NLP/20-LangModels/GPT-PyTorch.ipynb) | |
|
||||
| VI | **Autres techniques d'IA** || |
|
||||
| 21 | [Algorithmes génétiques](./lessons/6-Other/21-GeneticAlgorithms/README.md) | [Notebook](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/6-Other/21-GeneticAlgorithms/Genetic.ipynb) | |
|
||||
| 22 | [Apprentissage par renforcement profond](./lessons/6-Other/22-DeepRL/README.md) | [PyTorch](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/6-Other/22-DeepRL/CartPole-RL-PyTorch.ipynb) /[TensorFlow](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/6-Other/22-DeepRL/CartPole-RL-TF.ipynb) | [Lab](./lessons/6-Other/22-DeepRL/lab/README.md) |
|
||||
| 21 | [Algorithmes génétiques](./lessons/6-Other/21-GeneticAlgorithms/README.md) | [Notebook](./lessons/6-Other/21-GeneticAlgorithms/Genetic.ipynb) | |
|
||||
| 22 | [Apprentissage par renforcement profond](./lessons/6-Other/22-DeepRL/README.md) | [PyTorch](./lessons/6-Other/22-DeepRL/CartPole-RL-PyTorch.ipynb) /[TensorFlow](./lessons/6-Other/22-DeepRL/CartPole-RL-TF.ipynb) | [Lab](./lessons/6-Other/22-DeepRL/lab/README.md) |
|
||||
| 23 | [Systèmes multi-agents](./lessons/6-Other/23-MultiagentSystems/README.md) | | |
|
||||
| VII | **Éthique de l'IA** | | |
|
||||
| 24 | [Éthique de l'IA et IA responsable](./lessons/7-Ethics/README.md) | [Microsoft Learn : Principes d'IA responsable](https://docs.microsoft.com/learn/paths/responsible-ai-business-principles/?WT.mc_id=academic-77998-cacaste) | |
|
||||
| 24 | [Éthique de l'IA et IA responsable](./lessons/7-Ethics/README.md) | [Microsoft Learn : Principes de l'IA responsable](https://docs.microsoft.com/learn/paths/responsible-ai-business-principles/?WT.mc_id=academic-77998-cacaste) | |
|
||||
| IX | **Extras** | | |
|
||||
| 25 | [Réseaux multi-modaux, CLIP et VQGAN](./lessons/X-Extras/X1-MultiModal/README.md) | [Notebook](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/X-Extras/X1-MultiModal/Clip.ipynb) | |
|
||||
| 25 | [Réseaux multi-modaux, CLIP et VQGAN](./lessons/X-Extras/X1-MultiModal/README.md) | [Notebook](./lessons/X-Extras/X1-MultiModal/Clip.ipynb) | |
|
||||
|
||||
## Chaque leçon contient
|
||||
|
||||
* Du matériel de pré-lecture
|
||||
* Des notebooks Jupyter exécutables, souvent spécifiques au framework (**PyTorch** ou **TensorFlow**). Ces notebooks contiennent également beaucoup de contenu théorique, donc pour comprendre le sujet, il est nécessaire de parcourir au moins une version du notebook (PyTorch ou TensorFlow).
|
||||
* **Labs** disponibles pour certains sujets, qui vous permettent d'appliquer les connaissances acquises à un problème spécifique.
|
||||
* Matériel de pré-lecture
|
||||
* Notebooks Jupyter exécutables, souvent spécifiques au framework (**PyTorch** ou **TensorFlow**). Le notebook exécutable contient également beaucoup de contenu théorique, donc pour comprendre le sujet, vous devez parcourir au moins une version du notebook (PyTorch ou TensorFlow).
|
||||
* **Labs** disponibles pour certains sujets, qui vous permettent d'appliquer le contenu appris à un problème spécifique.
|
||||
* Certaines sections contiennent des liens vers des modules [**MS Learn**](https://learn.microsoft.com/en-us/collections/7w28iy2xrqzdj0?WT.mc_id=academic-77998-bethanycheum) qui couvrent des sujets connexes.
|
||||
|
||||
## Pour commencer
|
||||
|
||||
- Nous avons créé une [leçon d'installation](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/0-course-setup/setup.md) pour vous aider à configurer votre environnement de développement. - Pour les éducateurs, nous avons également créé une [leçon de configuration des programmes](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/0-course-setup/for-teachers.md) !
|
||||
- Comment [exécuter le code dans VSCode ou un Codespace](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/0-course-setup/how-to-run.md)
|
||||
- Nous avons créé une [leçon de configuration](./lessons/0-course-setup/setup.md) pour vous aider à configurer votre environnement de développement. - Pour les éducateurs, nous avons également créé une [leçon de configuration des programmes](./lessons/0-course-setup/for-teachers.md) !
|
||||
- Comment [exécuter le code dans VSCode ou Codepace](./lessons/0-course-setup/how-to-run.md)
|
||||
|
||||
Suivez ces étapes :
|
||||
|
||||
|
|
@ -110,46 +121,48 @@ Forkez le dépôt : Cliquez sur le bouton "Fork" en haut à droite de cette page
|
|||
|
||||
Clonez le dépôt : `git clone https://github.com/microsoft/AI-For-Beginners.git`
|
||||
|
||||
N'oubliez pas d'ajouter une étoile (🌟) à ce dépôt pour le retrouver plus facilement plus tard.
|
||||
N'oubliez pas de mettre une étoile (🌟) à ce dépôt pour le retrouver plus facilement plus tard.
|
||||
|
||||
## Rencontrez d'autres apprenants
|
||||
|
||||
Rejoignez notre [serveur Discord officiel sur l'IA](https://aka.ms/genai-discord?WT.mc_id=academic-105485-bethanycheum) pour rencontrer et échanger avec d'autres apprenants suivant ce cours et obtenir du soutien.
|
||||
|
||||
Si vous avez des retours ou des questions sur les produits pendant votre apprentissage, visitez notre [forum des développeurs Azure AI Foundry](https://aka.ms/foundry/forum)
|
||||
Si vous avez des retours sur le produit ou des questions pendant la construction, visitez notre [forum des développeurs Azure AI Foundry](https://aka.ms/foundry/forum)
|
||||
|
||||
## Quiz
|
||||
> **Une note à propos des quiz** : Tous les quiz se trouvent dans le dossier Quiz-app dans etc\quiz-app. Ils sont liés depuis les leçons, et l'application de quiz peut être exécutée localement ou déployée sur Azure ; suivez les instructions dans le dossier `quiz-app`. Ils sont progressivement en cours de localisation.
|
||||
> **Une note à propos des quiz** : Tous les quiz se trouvent dans le dossier Quiz-app sous etc\quiz-app, ou [en ligne ici](https://ff-quizzes.netlify.app/). Ils sont liés depuis les leçons. L'application de quiz peut être exécutée localement ou déployée sur Azure ; suivez les instructions dans le dossier `quiz-app`. Leur localisation est en cours de réalisation progressivement.
|
||||
## Besoin d'aide
|
||||
|
||||
Vous avez des suggestions ou avez trouvé des erreurs de code ou d'orthographe ? Ouvrez une issue ou créez une pull request.
|
||||
Vous avez des suggestions ou avez trouvé des erreurs d'orthographe ou de code ? Ouvrez une issue ou créez une pull request.
|
||||
|
||||
## Remerciements spéciaux
|
||||
|
||||
* **✍️ Auteur principal :** [Dmitry Soshnikov](http://soshnikov.com), PhD
|
||||
* **🔥 Éditrice :** [Jen Looper](https://twitter.com/jenlooper), PhD
|
||||
* **🎨 Illustratrice de sketchnotes :** [Tomomi Imura](https://twitter.com/girlie_mac)
|
||||
* **✅ Créatrice de quiz :** [Lateefah Bello](https://github.com/CinnamonXI), [MLSA](https://studentambassadors.microsoft.com/)
|
||||
* **🎨 Illustratrice des sketchnotes :** [Tomomi Imura](https://twitter.com/girlie_mac)
|
||||
* **✅ Créatrice du quiz :** [Lateefah Bello](https://github.com/CinnamonXI), [MLSA](https://studentambassadors.microsoft.com/)
|
||||
* **🙏 Contributeurs principaux :** [Evgenii Pishchik](https://github.com/Pe4enIks)
|
||||
|
||||
## Autres programmes
|
||||
|
||||
Notre équipe produit d'autres programmes ! Découvrez :
|
||||
Notre équipe produit d'autres programmes ! Découvrez-les :
|
||||
|
||||
- [IA générative pour débutants](https://aka.ms/genai-beginners)
|
||||
- [IA générative pour débutants .NET](https://github.com/microsoft/Generative-AI-for-beginners-dotnet)
|
||||
- [IA générative avec JavaScript](https://github.com/microsoft/generative-ai-with-javascript)
|
||||
- [IA générative avec Java](https://github.com/microsoft/Generative-AI-for-beginners-java)
|
||||
- [IA pour débutants](https://aka.ms/ai-beginners)
|
||||
- [Science des données pour débutants](https://aka.ms/datascience-beginners)
|
||||
- [Apprentissage automatique pour débutants](https://aka.ms/ml-beginners)
|
||||
- [Cybersécurité pour débutants](https://github.com/microsoft/Security-101)
|
||||
- [Développement web pour débutants](https://aka.ms/webdev-beginners)
|
||||
- [IoT pour débutants](https://aka.ms/iot-beginners)
|
||||
- [Développement XR pour débutants](https://github.com/microsoft/xr-development-for-beginners)
|
||||
- [Maîtriser GitHub Copilot pour une utilisation agentique](https://github.com/microsoft/Mastering-GitHub-Copilot-for-Paired-Programming)
|
||||
- [Maîtriser GitHub Copilot pour les développeurs C#/.NET](https://github.com/microsoft/mastering-github-copilot-for-dotnet-csharp-developers)
|
||||
- [Choisissez votre propre aventure avec Copilot](https://github.com/microsoft/CopilotAdventures)
|
||||
- [Generative AI for Beginners](https://aka.ms/genai-beginners)
|
||||
- [Generative AI for Beginners .NET](https://github.com/microsoft/Generative-AI-for-beginners-dotnet)
|
||||
- [Generative AI with JavaScript](https://github.com/microsoft/generative-ai-with-javascript)
|
||||
- [Generative AI with Java](https://github.com/microsoft/Generative-AI-for-beginners-java)
|
||||
- [AI for Beginners](https://aka.ms/ai-beginners)
|
||||
- [Data Science for Beginners](https://aka.ms/datascience-beginners)
|
||||
- [ML for Beginners](https://aka.ms/ml-beginners)
|
||||
- [Cybersecurity for Beginners](https://github.com/microsoft/Security-101)
|
||||
- [Web Dev for Beginners](https://aka.ms/webdev-beginners)
|
||||
- [IoT for Beginners](https://aka.ms/iot-beginners)
|
||||
- [XR Development for Beginners](https://github.com/microsoft/xr-development-for-beginners)
|
||||
- [Mastering GitHub Copilot for Agentic use](https://github.com/microsoft/Mastering-GitHub-Copilot-for-Paired-Programming)
|
||||
- [Mastering GitHub Copilot for C#/.NET Developers](https://github.com/microsoft/mastering-github-copilot-for-dotnet-csharp-developers)
|
||||
- [Choose Your Own Copilot Adventure](https://github.com/microsoft/CopilotAdventures)
|
||||
|
||||
---
|
||||
|
||||
**Avertissement** :
|
||||
Ce document a été traduit à l'aide du service de traduction automatique [Co-op Translator](https://github.com/Azure/co-op-translator). Bien que nous nous efforcions d'assurer l'exactitude, veuillez noter que les traductions automatisées peuvent contenir des erreurs ou des inexactitudes. Le document original dans sa langue d'origine doit être considéré comme la source faisant autorité. Pour des informations critiques, il est recommandé de recourir à une traduction humaine professionnelle. Nous déclinons toute responsabilité en cas de malentendus ou d'interprétations erronées résultant de l'utilisation de cette traduction.
|
||||
|
|
@ -0,0 +1,478 @@
|
|||
{
|
||||
"cells": [
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {
|
||||
"collapsed": true
|
||||
},
|
||||
"source": [
|
||||
"# Mise en œuvre d'un système expert pour les animaux\n",
|
||||
"\n",
|
||||
"Un exemple tiré du [programme d'études AI for Beginners](http://github.com/microsoft/ai-for-beginners).\n",
|
||||
"\n",
|
||||
"Dans cet exemple, nous allons mettre en œuvre un système simple basé sur la connaissance pour identifier un animal en fonction de certaines caractéristiques physiques. Le système peut être représenté par l'arbre AND-OR suivant (il s'agit d'une partie de l'arbre complet, nous pouvons facilement ajouter d'autres règles) :\n",
|
||||
"\n",
|
||||
"\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Notre propre shell de systèmes experts avec inférence rétrograde\n",
|
||||
"\n",
|
||||
"Essayons de définir un langage simple pour la représentation des connaissances basé sur des règles de production. Nous utiliserons des classes Python comme mots-clés pour définir les règles. Il y aurait essentiellement 3 types de classes :\n",
|
||||
"* `Ask` représente une question qui doit être posée à l'utilisateur. Elle contient l'ensemble des réponses possibles.\n",
|
||||
"* `If` représente une règle, et c'est juste une simplification syntaxique pour stocker le contenu de la règle.\n",
|
||||
"* `AND`/`OR` sont des classes pour représenter les branches ET/OU de l'arbre. Elles se contentent de stocker la liste des arguments à l'intérieur. Pour simplifier le code, toutes les fonctionnalités sont définies dans la classe parente `Content`.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 1,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"class Ask():\n",
|
||||
" def __init__(self,choices=['y','n']):\n",
|
||||
" self.choices = choices\n",
|
||||
" def ask(self):\n",
|
||||
" if max([len(x) for x in self.choices])>1:\n",
|
||||
" for i,x in enumerate(self.choices):\n",
|
||||
" print(\"{0}. {1}\".format(i,x),flush=True)\n",
|
||||
" x = int(input())\n",
|
||||
" return self.choices[x]\n",
|
||||
" else:\n",
|
||||
" print(\"/\".join(self.choices),flush=True)\n",
|
||||
" return input()\n",
|
||||
"\n",
|
||||
"class Content():\n",
|
||||
" def __init__(self,x):\n",
|
||||
" self.x=x\n",
|
||||
" \n",
|
||||
"class If(Content):\n",
|
||||
" pass\n",
|
||||
"\n",
|
||||
"class AND(Content):\n",
|
||||
" pass\n",
|
||||
"\n",
|
||||
"class OR(Content):\n",
|
||||
" pass"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Dans notre système, la mémoire de travail contiendrait la liste des **faits** sous forme de **paires attribut-valeur**. La base de connaissances peut être définie comme un grand dictionnaire qui associe des actions (nouveaux faits devant être insérés dans la mémoire de travail) à des conditions, exprimées sous forme d'expressions ET-OU. De plus, certains faits peuvent être `Demandés`.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 2,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"rules = {\n",
|
||||
" 'default': Ask(['y','n']),\n",
|
||||
" 'color' : Ask(['red-brown','black and white','other']),\n",
|
||||
" 'pattern' : Ask(['dark stripes','dark spots']),\n",
|
||||
" 'mammal': If(OR(['hair','gives milk'])),\n",
|
||||
" 'carnivor': If(OR([AND(['sharp teeth','claws','forward-looking eyes']),'eats meat'])),\n",
|
||||
" 'ungulate': If(['mammal',OR(['has hooves','chews cud'])]),\n",
|
||||
" 'bird': If(OR(['feathers',AND(['flies','lies eggs'])])),\n",
|
||||
" 'animal:monkey' : If(['mammal','carnivor','color:red-brown','pattern:dark spots']),\n",
|
||||
" 'animal:tiger' : If(['mammal','carnivor','color:red-brown','pattern:dark stripes']),\n",
|
||||
" 'animal:giraffe' : If(['ungulate','long neck','long legs','pattern:dark spots']),\n",
|
||||
" 'animal:zebra' : If(['ungulate','pattern:dark stripes']),\n",
|
||||
" 'animal:ostrich' : If(['bird','long nech','color:black and white','cannot fly']),\n",
|
||||
" 'animal:pinguin' : If(['bird','swims','color:black and white','cannot fly']),\n",
|
||||
" 'animal:albatross' : If(['bird','flies well'])\n",
|
||||
"}"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Pour effectuer l'inférence à rebours, nous allons définir la classe `Knowledgebase`. Elle contiendra :\n",
|
||||
"* Une `mémoire` de travail - un dictionnaire qui associe des attributs à des valeurs\n",
|
||||
"* Les `règles` de la base de connaissances dans le format défini ci-dessus\n",
|
||||
"\n",
|
||||
"Les deux méthodes principales sont :\n",
|
||||
"* `get` pour obtenir la valeur d'un attribut, en effectuant une inférence si nécessaire. Par exemple, `get('color')` obtiendra la valeur d'un champ de couleur (il posera la question si nécessaire et stockera la valeur pour une utilisation ultérieure dans la mémoire de travail). Si nous demandons `get('color:blue')`, il demandera une couleur, puis retournera une valeur `y`/`n` en fonction de la couleur.\n",
|
||||
"* `eval` effectue l'inférence proprement dite, c'est-à-dire qu'elle parcourt l'arbre AND/OR, évalue les sous-objectifs, etc.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 3,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"class KnowledgeBase():\n",
|
||||
" def __init__(self,rules):\n",
|
||||
" self.rules = rules\n",
|
||||
" self.memory = {}\n",
|
||||
" \n",
|
||||
" def get(self,name):\n",
|
||||
" if ':' in name:\n",
|
||||
" k,v = name.split(':')\n",
|
||||
" vv = self.get(k)\n",
|
||||
" return 'y' if v==vv else 'n'\n",
|
||||
" if name in self.memory.keys():\n",
|
||||
" return self.memory[name]\n",
|
||||
" for fld in self.rules.keys():\n",
|
||||
" if fld==name or fld.startswith(name+\":\"):\n",
|
||||
" # print(\" + proving {}\".format(fld))\n",
|
||||
" value = 'y' if fld==name else fld.split(':')[1]\n",
|
||||
" res = self.eval(self.rules[fld],field=name)\n",
|
||||
" if res!='y' and res!='n' and value=='y':\n",
|
||||
" self.memory[name] = res\n",
|
||||
" return res\n",
|
||||
" if res=='y':\n",
|
||||
" self.memory[name] = value\n",
|
||||
" return value\n",
|
||||
" # field is not found, using default\n",
|
||||
" res = self.eval(self.rules['default'],field=name)\n",
|
||||
" self.memory[name]=res\n",
|
||||
" return res\n",
|
||||
" \n",
|
||||
" def eval(self,expr,field=None):\n",
|
||||
" # print(\" + eval {}\".format(expr))\n",
|
||||
" if isinstance(expr,Ask):\n",
|
||||
" print(field)\n",
|
||||
" return expr.ask()\n",
|
||||
" elif isinstance(expr,If):\n",
|
||||
" return self.eval(expr.x)\n",
|
||||
" elif isinstance(expr,AND) or isinstance(expr,list):\n",
|
||||
" expr = expr.x if isinstance(expr,AND) else expr\n",
|
||||
" for x in expr:\n",
|
||||
" if self.eval(x)=='n':\n",
|
||||
" return 'n'\n",
|
||||
" return 'y'\n",
|
||||
" elif isinstance(expr,OR):\n",
|
||||
" for x in expr.x:\n",
|
||||
" if self.eval(x)=='y':\n",
|
||||
" return 'y'\n",
|
||||
" return 'n'\n",
|
||||
" elif isinstance(expr,str):\n",
|
||||
" return self.get(expr)\n",
|
||||
" else:\n",
|
||||
" print(\"Unknown expr: {}\".format(expr))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Maintenant, définissons notre base de connaissances sur les animaux et effectuons la consultation. Notez que cet appel vous posera des questions. Vous pouvez répondre en tapant `y`/`n` pour les questions oui-non, ou en spécifiant un numéro (0..N) pour les questions avec des réponses à choix multiples plus longues.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 4,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"hair\n",
|
||||
"y/n\n",
|
||||
"sharp teeth\n",
|
||||
"y/n\n",
|
||||
"claws\n",
|
||||
"y/n\n",
|
||||
"forward-looking eyes\n",
|
||||
"y/n\n",
|
||||
"color\n",
|
||||
"0. red-brown\n",
|
||||
"1. black and white\n",
|
||||
"2. other\n",
|
||||
"has hooves\n",
|
||||
"y/n\n",
|
||||
"long neck\n",
|
||||
"y/n\n",
|
||||
"long legs\n",
|
||||
"y/n\n",
|
||||
"pattern\n",
|
||||
"0. dark stripes\n",
|
||||
"1. dark spots\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"'giraffe'"
|
||||
]
|
||||
},
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"kb = KnowledgeBase(rules)\n",
|
||||
"kb.get('animal')"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Utilisation de PyKnow pour l'inférence avant\n",
|
||||
"\n",
|
||||
"Dans l'exemple suivant, nous allons essayer de mettre en œuvre l'inférence avant en utilisant l'une des bibliothèques de représentation des connaissances, [PyKnow](https://github.com/buguroo/pyknow/). **PyKnow** est une bibliothèque permettant de créer des systèmes d'inférence avant en Python, conçue pour être similaire au système classique ancien [CLIPS](http://www.clipsrules.net/index.html).\n",
|
||||
"\n",
|
||||
"Nous aurions également pu implémenter nous-mêmes le chaînage avant sans trop de difficultés, mais les implémentations naïves ne sont généralement pas très efficaces. Pour un appariement des règles plus performant, un algorithme spécial appelé [Rete](https://en.wikipedia.org/wiki/Rete_algorithm) est utilisé.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 5,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Collecting git+https://github.com/buguroo/pyknow/\n",
|
||||
" Cloning https://github.com/buguroo/pyknow/ to /tmp/pip-req-build-3cqeulyl\n",
|
||||
" Running command git clone --filter=blob:none --quiet https://github.com/buguroo/pyknow/ /tmp/pip-req-build-3cqeulyl\n",
|
||||
" Resolved https://github.com/buguroo/pyknow/ to commit 48818336f2e9a126f1964f2d8dc22d37ff800fe8\n",
|
||||
" Preparing metadata (setup.py) ... \u001b[?25ldone\n",
|
||||
"\u001b[?25hCollecting frozendict==1.2\n",
|
||||
" Using cached frozendict-1.2.tar.gz (2.6 kB)\n",
|
||||
" Preparing metadata (setup.py) ... \u001b[?25ldone\n",
|
||||
"\u001b[?25hCollecting schema==0.6.7\n",
|
||||
" Using cached schema-0.6.7-py2.py3-none-any.whl (14 kB)\n",
|
||||
"Building wheels for collected packages: pyknow, frozendict\n",
|
||||
" Building wheel for pyknow (setup.py) ... \u001b[?25ldone\n",
|
||||
"\u001b[?25h Created wheel for pyknow: filename=pyknow-1.7.0-py3-none-any.whl size=34228 sha256=b7de5b09292c4007667c72f69b98d5a1b5f7324ff15f9dd8e077c3d5f7aade42\n",
|
||||
" Stored in directory: /tmp/pip-ephem-wheel-cache-k7jpave7/wheels/81/1a/d3/f6c15dbe1955598a37755215f2a10449e7418500d7bd4b9508\n",
|
||||
" Building wheel for frozendict (setup.py) ... \u001b[?25ldone\n",
|
||||
"\u001b[?25h Created wheel for frozendict: filename=frozendict-1.2-py3-none-any.whl size=3148 sha256=2863d55c240d2409cddf05ccfe600591f8478681549fc97555c47c90dc6bb160\n",
|
||||
" Stored in directory: /home/rg/.cache/pip/wheels/49/ac/f8/cb8120244e710bdb479c86198b03c7b08c3c2d3d2bf448fd6e\n",
|
||||
"Successfully built pyknow frozendict\n",
|
||||
"Installing collected packages: schema, frozendict, pyknow\n",
|
||||
"Successfully installed frozendict-1.2 pyknow-1.7.0 schema-0.6.7\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"import sys\n",
|
||||
"!{sys.executable} -m pip install git+https://github.com/buguroo/pyknow/"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 13,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"from pyknow import *\n",
|
||||
"#import pyknow"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Nous définirons notre système comme une classe qui hérite de `KnowledgeEngine`. Chaque règle est définie par une fonction distincte avec l'annotation `@Rule`, qui spécifie quand la règle doit s'exécuter. À l'intérieur de la règle, nous pouvons ajouter de nouveaux faits en utilisant la fonction `declare`, et l'ajout de ces faits entraînera l'appel de certaines autres règles par le moteur d'inférence avant.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 14,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"class Animals(KnowledgeEngine):\n",
|
||||
" @Rule(OR(\n",
|
||||
" AND(Fact('sharp teeth'),Fact('claws'),Fact('forward looking eyes')),\n",
|
||||
" Fact('eats meat')))\n",
|
||||
" def cornivor(self):\n",
|
||||
" self.declare(Fact('carnivor'))\n",
|
||||
" \n",
|
||||
" @Rule(OR(Fact('hair'),Fact('gives milk')))\n",
|
||||
" def mammal(self):\n",
|
||||
" self.declare(Fact('mammal'))\n",
|
||||
"\n",
|
||||
" @Rule(Fact('mammal'),\n",
|
||||
" OR(Fact('has hooves'),Fact('chews cud')))\n",
|
||||
" def hooves(self):\n",
|
||||
" self.declare('ungulate')\n",
|
||||
" \n",
|
||||
" @Rule(OR(Fact('feathers'),AND(Fact('flies'),Fact('lays eggs'))))\n",
|
||||
" def bird(self):\n",
|
||||
" self.declare('bird')\n",
|
||||
" \n",
|
||||
" @Rule(Fact('mammal'),Fact('carnivor'),\n",
|
||||
" Fact(color='red-brown'),\n",
|
||||
" Fact(pattern='dark spots'))\n",
|
||||
" def monkey(self):\n",
|
||||
" self.declare(Fact(animal='monkey'))\n",
|
||||
"\n",
|
||||
" @Rule(Fact('mammal'),Fact('carnivor'),\n",
|
||||
" Fact(color='red-brown'),\n",
|
||||
" Fact(pattern='dark stripes'))\n",
|
||||
" def tiger(self):\n",
|
||||
" self.declare(Fact(animal='tiger'))\n",
|
||||
"\n",
|
||||
" @Rule(Fact('ungulate'),\n",
|
||||
" Fact('long neck'),\n",
|
||||
" Fact('long legs'),\n",
|
||||
" Fact(pattern='dark spots'))\n",
|
||||
" def giraffe(self):\n",
|
||||
" self.declare(Fact(animal='giraffe'))\n",
|
||||
"\n",
|
||||
" @Rule(Fact('ungulate'),\n",
|
||||
" Fact(pattern='dark stripes'))\n",
|
||||
" def zebra(self):\n",
|
||||
" self.declare(Fact(animal='zebra'))\n",
|
||||
"\n",
|
||||
" @Rule(Fact('bird'),\n",
|
||||
" Fact('long neck'),\n",
|
||||
" Fact('cannot fly'),\n",
|
||||
" Fact(color='black and white'))\n",
|
||||
" def straus(self):\n",
|
||||
" self.declare(Fact(animal='ostrich'))\n",
|
||||
"\n",
|
||||
" @Rule(Fact('bird'),\n",
|
||||
" Fact('swims'),\n",
|
||||
" Fact('cannot fly'),\n",
|
||||
" Fact(color='black and white'))\n",
|
||||
" def pinguin(self):\n",
|
||||
" self.declare(Fact(animal='pinguin'))\n",
|
||||
"\n",
|
||||
" @Rule(Fact('bird'),\n",
|
||||
" Fact('flies well'))\n",
|
||||
" def albatros(self):\n",
|
||||
" self.declare(Fact(animal='albatross'))\n",
|
||||
" \n",
|
||||
" @Rule(Fact(animal=MATCH.a))\n",
|
||||
" def print_result(self,a):\n",
|
||||
" print('Animal is {}'.format(a))\n",
|
||||
" \n",
|
||||
" def factz(self,l):\n",
|
||||
" for x in l:\n",
|
||||
" self.declare(x)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Une fois que nous avons défini une base de connaissances, nous remplissons notre mémoire de travail avec quelques faits initiaux, puis nous appelons la méthode `run()` pour effectuer l'inférence. Vous pouvez voir qu'en conséquence, de nouveaux faits déduits sont ajoutés à la mémoire de travail, y compris le fait final concernant l'animal (si nous avons correctement configuré tous les faits initiaux).\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 15,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Animal is tiger\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"FactList([(0, InitialFact()),\n",
|
||||
" (1, Fact(color='red-brown')),\n",
|
||||
" (2, Fact(pattern='dark stripes')),\n",
|
||||
" (3, Fact('sharp teeth')),\n",
|
||||
" (4, Fact('claws')),\n",
|
||||
" (5, Fact('forward looking eyes')),\n",
|
||||
" (6, Fact('gives milk')),\n",
|
||||
" (7, Fact('mammal')),\n",
|
||||
" (8, Fact('carnivor')),\n",
|
||||
" (9, Fact(animal='tiger'))])"
|
||||
]
|
||||
},
|
||||
"execution_count": 15,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"ex1 = Animals()\n",
|
||||
"ex1.reset()\n",
|
||||
"ex1.factz([\n",
|
||||
" Fact(color='red-brown'),\n",
|
||||
" Fact(pattern='dark stripes'),\n",
|
||||
" Fact('sharp teeth'),\n",
|
||||
" Fact('claws'),\n",
|
||||
" Fact('forward looking eyes'),\n",
|
||||
" Fact('gives milk')])\n",
|
||||
"ex1.run()\n",
|
||||
"ex1.facts"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": []
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"\n---\n\n**Avertissement** : \nCe document a été traduit à l'aide du service de traduction automatique [Co-op Translator](https://github.com/Azure/co-op-translator). Bien que nous nous efforcions d'assurer l'exactitude, veuillez noter que les traductions automatisées peuvent contenir des erreurs ou des inexactitudes. Le document original dans sa langue d'origine doit être considéré comme la source faisant autorité. Pour des informations critiques, il est recommandé de recourir à une traduction professionnelle réalisée par un humain. Nous déclinons toute responsabilité en cas de malentendus ou d'interprétations erronées résultant de l'utilisation de cette traduction.\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
"kernelspec": {
|
||||
"display_name": "Python 3.7.4 64-bit (conda)",
|
||||
"metadata": {
|
||||
"interpreter": {
|
||||
"hash": "86193a1ab0ba47eac1c69c1756090baa3b420b3eea7d4aafab8b85f8b312f0c5"
|
||||
}
|
||||
},
|
||||
"name": "python3"
|
||||
},
|
||||
"language_info": {
|
||||
"codemirror_mode": {
|
||||
"name": "ipython",
|
||||
"version": 3
|
||||
},
|
||||
"file_extension": ".py",
|
||||
"mimetype": "text/x-python",
|
||||
"name": "python",
|
||||
"nbconvert_exporter": "python",
|
||||
"pygments_lexer": "ipython3",
|
||||
"version": "3.11.2"
|
||||
},
|
||||
"coopTranslator": {
|
||||
"original_hash": "ab2bd97b0453415b89a469284609a8ce",
|
||||
"translation_date": "2025-08-31T14:54:39+00:00",
|
||||
"source_file": "lessons/2-Symbolic/Animals.ipynb",
|
||||
"language_code": "fr"
|
||||
}
|
||||
},
|
||||
"nbformat": 4,
|
||||
"nbformat_minor": 2
|
||||
}
|
||||
|
|
@ -0,0 +1,595 @@
|
|||
{
|
||||
"cells": [
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {
|
||||
"collapsed": true
|
||||
},
|
||||
"source": [
|
||||
"# Ontologie des Relations Familiales\n",
|
||||
"\n",
|
||||
"Cet exemple fait partie du [programme AI for Beginners](http://github.com/microsoft/ai-for-beginners), et il s'inspire de [cet article de blog](https://habr.com/post/270857/).\n",
|
||||
"\n",
|
||||
"J'ai toujours trouvé difficile de me souvenir des différentes relations entre les membres d'une famille. Dans cet exemple, nous allons utiliser une ontologie qui définit les relations familiales, ainsi qu'un arbre généalogique réel, et montrer comment nous pouvons ensuite effectuer une inférence automatique pour trouver tous les membres de la famille.\n",
|
||||
"\n",
|
||||
"### Obtenir l'Arbre Généalogique\n",
|
||||
"\n",
|
||||
"À titre d'exemple, nous allons utiliser l'arbre généalogique de la [famille des Tsars Romanov](https://en.wikipedia.org/wiki/House_of_Romanov). Le format le plus courant pour décrire les relations familiales est le [GEDCOM](https://en.wikipedia.org/wiki/GEDCOM). Nous allons utiliser l'arbre généalogique de la famille Romanov au format GEDCOM :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 1,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"0 HEAD\n",
|
||||
"1 CHAR UTF8\n",
|
||||
"1 GEDC\n",
|
||||
"2 VERS 5.5\n",
|
||||
"0 @0@ INDI\n",
|
||||
"1 NAME Mihail Fedorovich /Romanov/\n",
|
||||
"1 SEX M\n",
|
||||
"1 BIRT\n",
|
||||
"2 DATE 1613\n",
|
||||
"1 DEAT \n",
|
||||
"2 DATE 1645\n",
|
||||
"1 FAMS @41@\n",
|
||||
"0 @1@ INDI\n",
|
||||
"1 NAME Evdokija Lukjanovna /Streshneva/\n",
|
||||
"1 SEX F\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"!head -15 data/tsars.ged"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Pour utiliser un fichier GEDCOM, nous pouvons utiliser la bibliothèque `python-gedcom` :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 2,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Collecting python-gedcom\n",
|
||||
" Downloading python_gedcom-1.0.0-py2.py3-none-any.whl (35 kB)\n",
|
||||
"Installing collected packages: python-gedcom\n",
|
||||
"Successfully installed python-gedcom-1.0.0\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"import sys\n",
|
||||
"!{sys.executable} -m pip install python-gedcom"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Cette bibliothèque élimine certains des problèmes techniques liés à l'analyse de fichiers, mais elle nous donne toujours un accès assez bas niveau à tous les individus et familles dans l'arbre. Voici comment nous pouvons analyser le fichier et afficher la liste de tous les individus :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 3,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"from gedcom.parser import Parser\n",
|
||||
"from gedcom.element.individual import IndividualElement\n",
|
||||
"from gedcom.element.family import FamilyElement\n",
|
||||
"g = Parser()\n",
|
||||
"g.parse_file('data/tsars.ged')"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 4,
|
||||
"metadata": {
|
||||
"scrolled": true,
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"[('@0@', ('Mihail Fedorovich', 'Romanov')),\n",
|
||||
" ('@1@', ('Evdokija Lukjanovna', 'Streshneva')),\n",
|
||||
" ('@2@', ('Aleksej Mihajlovich', 'Romanov')),\n",
|
||||
" ('@3@', ('Marija Ilinichna', 'Miloslavskaja')),\n",
|
||||
" ('@4@', ('Natalja Kirillovna', 'Naryshkina')),\n",
|
||||
" ('@5@', ('Marfa Matveevna', 'Apraksina')),\n",
|
||||
" ('@6@', ('Fedor Alekseevich', 'Romanov')),\n",
|
||||
" ('@7@', ('Sofja Aleksevna', 'Romanova')),\n",
|
||||
" ('@8@', ('Ivan V Alekseevich', 'Romanov')),\n",
|
||||
" ('@9@', ('Praskovja Fedorovna', 'Saltykova')),\n",
|
||||
" ('@10@', ('Ekaterina Ivanovna', 'Romanova')),\n",
|
||||
" ('@11@', ('Anna Ivanovna', 'Romanova')),\n",
|
||||
" ('@12@', ('Fridrih Vilgelm', 'Kurlandskij')),\n",
|
||||
" ('@13@', ('Karl Leopold', 'Meklenburg-Shverinskij')),\n",
|
||||
" ('@14@', ('Anna Leopoldovna', 'Meklenburg-Shverinskaja')),\n",
|
||||
" ('@15@', ('Anton Ulrih', 'Braunshvejg-Volfenbjuttelskij')),\n",
|
||||
" ('@16@', ('Ivan VI Antonovich', 'Braunshvejg-Volfenbjuttelskij')),\n",
|
||||
" ('@17@', ('Petr I Alekseevich', 'Romanov')),\n",
|
||||
" ('@18@', ('Evdokija Fedorovna', 'Lopuhina')),\n",
|
||||
" ('@19@', ('Ekaterina I Alekseevna', 'Mihajlova')),\n",
|
||||
" ('@20@', ('Aleksej Petrovich', 'Romanov')),\n",
|
||||
" ('@21@', ('Sharlotta Kristina', 'Braunshvejg-Volfenbjuttelskaja')),\n",
|
||||
" ('@22@', ('Petr II Alekseevich', 'Romanov')),\n",
|
||||
" ('@23@', ('Anna Petrovna', 'Romanova')),\n",
|
||||
" ('@24@', ('Elizaveta Petrovna', 'Romanova')),\n",
|
||||
" ('@25@', ('Karl Fridrih', 'Golshtejn-Gottorpskij')),\n",
|
||||
" ('@26@', ('Petr III Fedorovich', 'Romanov')),\n",
|
||||
" ('@27@', ('Ekaterina II', 'Alekseevna')),\n",
|
||||
" ('@28@', ('Pavel I Petrovich', 'Romanov')),\n",
|
||||
" ('@29@', ('Natalja Alekseevna', 'Gessen-Darmshtadskaja')),\n",
|
||||
" ('@30@', ('Marija Fedorovna', 'Vjurtembergskaja')),\n",
|
||||
" ('@31@', ('Aleksandr I Pavlovich', 'Romanov')),\n",
|
||||
" ('@32@', ('Elizaveta Alekseevna', 'Baden-Durlahskaja')),\n",
|
||||
" ('@33@', ('Nikolaj I Pavlovich', 'Romanov')),\n",
|
||||
" ('@34@', ('Aleksandra Fedorovna', 'Prusskaja')),\n",
|
||||
" ('@35@', ('Aleksandr II Nikolaevich', 'Romanov')),\n",
|
||||
" ('@36@', ('Marija Aleksandrovna', 'Gessenskaja')),\n",
|
||||
" ('@37@', ('Aleksandr III Aleksandrovich', 'Romanov')),\n",
|
||||
" ('@38@', ('Marija Fedorovna', 'Datskaja')),\n",
|
||||
" ('@39@', ('Nikolaj II Aleksandrovich', 'Romanov')),\n",
|
||||
" ('@40@', ('Aleksandra Fedorovna', 'Gessenskaja'))]"
|
||||
]
|
||||
},
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"d = g.get_element_dictionary()\n",
|
||||
"[ (k,v.get_name()) for k,v in d.items() if isinstance(v,IndividualElement)]"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Voici comment nous pouvons obtenir des informations sur les familles. Notez que cela nous donne une liste d'**identifiants**, et nous devons les convertir en noms si nous voulons plus de clarté :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 5,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"[('@41@', ['@0@', '@1@', '@2@']),\n",
|
||||
" ('@42@', ['@2@', '@3@', '@6@', '@7@', '@8@']),\n",
|
||||
" ('@43@', ['@8@', '@9@', '@10@', '@11@']),\n",
|
||||
" ('@44@', ['@13@', '@10@', '@14@']),\n",
|
||||
" ('@45@', ['@15@', '@14@', '@16@']),\n",
|
||||
" ('@46@', ['@2@', '@4@', '@17@']),\n",
|
||||
" ('@47@', ['@17@', '@18@', '@20@']),\n",
|
||||
" ('@48@', ['@20@', '@21@', '@22@']),\n",
|
||||
" ('@49@', ['@17@', '@19@', '@23@', '@24@']),\n",
|
||||
" ('@50@', ['@25@', '@23@', '@26@']),\n",
|
||||
" ('@51@', ['@26@', '@27@', '@28@']),\n",
|
||||
" ('@52@', ['@28@', '@30@', '@31@', '@33@']),\n",
|
||||
" ('@53@', ['@33@', '@34@', '@35@']),\n",
|
||||
" ('@54@', ['@35@', '@36@', '@37@']),\n",
|
||||
" ('@55@', ['@37@', '@38@', '@39@'])]"
|
||||
]
|
||||
},
|
||||
"execution_count": 5,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"d = g.get_element_dictionary()\n",
|
||||
"[ (k,[x.get_value() for x in v.get_child_elements()]) for k,v in d.items() if isinstance(v,FamilyElement)]"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"### Obtenir l'ontologie familiale\n",
|
||||
"\n",
|
||||
"Ensuite, examinons [l'ontologie familiale](https://raw.githubusercontent.com/blokhin/genealogical-trees/master/data/header.ttl) définie comme un ensemble de triplets du Web sémantique. Cette ontologie définit des relations telles que `isUncleOf`, `isCousinOf`, et bien d'autres. Toutes ces relations sont définies en termes de prédicats de base `isMotherOf`, `isFatherOf`, `isBrotherOf` et `isSisterOf`. Nous utiliserons un raisonnement automatique pour déduire toutes les autres relations à partir de l'ontologie.\n",
|
||||
"\n",
|
||||
"Voici un exemple de définition de la propriété `isAuntOf`, qui est définie comme une composition de `isSisterOf` et `isParentOf` (*Une tante est la sœur d'un parent*).\n",
|
||||
"\n",
|
||||
"```\n",
|
||||
"fhkb:isAuntOf a owl:ObjectProperty ;\n",
|
||||
" rdfs:domain fhkb:Woman ;\n",
|
||||
" rdfs:range fhkb:Person ;\n",
|
||||
" owl:propertyChainAxiom ( fhkb:isSisterOf fhkb:isParentOf ) .\n",
|
||||
"```\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 6,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"@prefix fhkb: <http://www.example.com/genealogy.owl#> .\n",
|
||||
"@prefix owl: <http://www.w3.org/2002/07/owl#> .\n",
|
||||
"@prefix rdf: <http://www.w3.org/1999/02/22-rdf-syntax-ns#> .\n",
|
||||
"@prefix rdfs: <http://www.w3.org/2000/01/rdf-schema#> .\n",
|
||||
"@prefix xml: <http://www.w3.org/XML/1998/namespace> .\n",
|
||||
"@prefix xsd: <http://www.w3.org/2001/XMLSchema#> .\n",
|
||||
"\n",
|
||||
"<http://www.example.com/genealogy.owl#> a owl:Ontology .\n",
|
||||
"\n",
|
||||
"fhkb:DomainEntity a owl:Class .\n",
|
||||
"\n",
|
||||
"fhkb:Man a owl:Class ;\n",
|
||||
" owl:equivalentClass [ a owl:Class ;\n",
|
||||
" owl:intersectionOf ( fhkb:Person [ a owl:Restriction ;\n",
|
||||
" owl:onProperty fhkb:hasSex ;\n",
|
||||
" owl:someValuesFrom fhkb:Male ] ) ] .\n",
|
||||
"\n",
|
||||
"fhkb:Woman a owl:Class ;\n",
|
||||
" owl:equivalentClass [ a owl:Class ;\n",
|
||||
" owl:intersectionOf ( fhkb:Person [ a owl:Restriction ;\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"!head -20 data/onto.ttl"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"### Construire une ontologie pour l'inférence\n",
|
||||
"\n",
|
||||
"Pour simplifier, nous allons créer un fichier d'ontologie unique qui inclura les règles originales de l'ontologie familiale, ainsi que les faits concernant les individus issus de notre fichier GEDCOM. Nous parcourrons le fichier GEDCOM pour extraire des informations sur les familles et les individus, et les convertir en triplets.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 7,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"!cp data/onto.ttl .\n",
|
||||
"\n",
|
||||
"gedcom_dict = g.get_element_dictionary()\n",
|
||||
"individuals, marriages = {}, {}\n",
|
||||
"\n",
|
||||
"def term2id(el):\n",
|
||||
" return \"i\" + el.get_pointer().replace('@', '').lower()\n",
|
||||
"\n",
|
||||
"out = open(\"onto.ttl\",\"a\")\n",
|
||||
"\n",
|
||||
"for k, v in gedcom_dict.items():\n",
|
||||
" if isinstance(v,IndividualElement):\n",
|
||||
" children, siblings = set(), set()\n",
|
||||
" idx = term2id(v)\n",
|
||||
"\n",
|
||||
" title = v.get_name()[0] + \" \" + v.get_name()[1]\n",
|
||||
" title = title.replace('\"', '').replace('[', '').replace(']', '').replace('(', '').replace(')', '').strip()\n",
|
||||
"\n",
|
||||
" own_families = g.get_families(v, 'FAMS')\n",
|
||||
" for fam in own_families:\n",
|
||||
" children |= set(term2id(i) for i in g.get_family_members(fam, \"CHIL\"))\n",
|
||||
"\n",
|
||||
" parent_families = g.get_families(v, 'FAMC')\n",
|
||||
" if len(parent_families):\n",
|
||||
" for member in g.get_family_members(parent_families[0], \"CHIL\"): # NB adoptive families i.e len(parent_families)>1 are not considered (TODO?)\n",
|
||||
" if member.get_pointer() == v.get_pointer():\n",
|
||||
" continue\n",
|
||||
" siblings.add(term2id(member))\n",
|
||||
"\n",
|
||||
" if idx in individuals:\n",
|
||||
" children |= individuals[idx].get('children', set())\n",
|
||||
" siblings |= individuals[idx].get('siblings', set())\n",
|
||||
" individuals[idx] = {'sex': v.get_gender().lower(), 'children': children, 'siblings': siblings, 'title': title}\n",
|
||||
"\n",
|
||||
" elif isinstance(v,FamilyElement):\n",
|
||||
" wife, husb, children = None, None, set()\n",
|
||||
" children = set(term2id(i) for i in g.get_family_members(v, \"CHIL\"))\n",
|
||||
"\n",
|
||||
" try:\n",
|
||||
" wife = g.get_family_members(v, \"WIFE\")[0]\n",
|
||||
" wife = term2id(wife)\n",
|
||||
" if wife in individuals: individuals[wife]['children'] |= children\n",
|
||||
" else: individuals[wife] = {'children': children}\n",
|
||||
" except IndexError: pass\n",
|
||||
" try:\n",
|
||||
" husb = g.get_family_members(v, \"HUSB\")[0]\n",
|
||||
" husb = term2id(husb)\n",
|
||||
" if husb in individuals: individuals[husb]['children'] |= children\n",
|
||||
" else: individuals[husb] = {'children': children}\n",
|
||||
" except IndexError: pass\n",
|
||||
"\n",
|
||||
" if wife and husb: marriages[wife + husb] = (term2id(v), wife, husb)\n",
|
||||
"\n",
|
||||
"for idx, val in individuals.items():\n",
|
||||
" added_terms = ''\n",
|
||||
" if val['sex'] == 'f':\n",
|
||||
" parent_predicate, sibl_predicate = \"isMotherOf\", \"isSisterOf\"\n",
|
||||
" else:\n",
|
||||
" parent_predicate, sibl_predicate = \"isFatherOf\", \"isBrotherOf\"\n",
|
||||
" if len(val['children']):\n",
|
||||
" added_terms += \" ;\\n fhkb:\" + parent_predicate + \" \" + \", \".join([\"fhkb:\" + i for i in val['children']])\n",
|
||||
" if len(val['siblings']):\n",
|
||||
" added_terms += \" ;\\n fhkb:\" + sibl_predicate + \" \" + \", \".join([\"fhkb:\" + i for i in val['siblings']])\n",
|
||||
" out.write(\"fhkb:%s a owl:NamedIndividual, owl:Thing%s ;\\n rdfs:label \\\"%s\\\" .\\n\" % (idx, added_terms, val['title']))\n",
|
||||
"\n",
|
||||
"for k, v in marriages.items():\n",
|
||||
" out.write(\"fhkb:%s a owl:NamedIndividual, owl:Thing ;\\n fhkb:hasFemalePartner fhkb:%s ;\\n fhkb:hasMalePartner fhkb:%s .\\n\" % v)\n",
|
||||
"\n",
|
||||
"out.write(\"[] a owl:AllDifferent ;\\n owl:distinctMembers (\")\n",
|
||||
"for idx in individuals.keys():\n",
|
||||
" out.write(\" fhkb:\" + idx)\n",
|
||||
"for k, v in marriages.items():\n",
|
||||
" out.write(\" fhkb:\" + v[0])\n",
|
||||
"out.write(\" ) .\")\n",
|
||||
"out.close()"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 8,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
" fhkb:hasFemalePartner fhkb:i34 ;\n",
|
||||
" fhkb:hasMalePartner fhkb:i33 .\n",
|
||||
"fhkb:i54 a owl:NamedIndividual, owl:Thing ;\n",
|
||||
" fhkb:hasFemalePartner fhkb:i36 ;\n",
|
||||
" fhkb:hasMalePartner fhkb:i35 .\n",
|
||||
"fhkb:i55 a owl:NamedIndividual, owl:Thing ;\n",
|
||||
" fhkb:hasFemalePartner fhkb:i38 ;\n",
|
||||
" fhkb:hasMalePartner fhkb:i37 .\n",
|
||||
"[] a owl:AllDifferent ;\n",
|
||||
" owl:distinctMembers ( fhkb:i0 fhkb:i1 fhkb:i2 fhkb:i3 fhkb:i4 fhkb:i5 fhkb:i6 fhkb:i7 fhkb:i8 fhkb:i9 fhkb:i10 fhkb:i11 fhkb:i12 fhkb:i13 fhkb:i14 fhkb:i15 fhkb:i16 fhkb:i17 fhkb:i18 fhkb:i19 fhkb:i20 fhkb:i21 fhkb:i22 fhkb:i23 fhkb:i24 fhkb:i25 fhkb:i26 fhkb:i27 fhkb:i28 fhkb:i29 fhkb:i30 fhkb:i31 fhkb:i32 fhkb:i33 fhkb:i34 fhkb:i35 fhkb:i36 fhkb:i37 fhkb:i38 fhkb:i39 fhkb:i40 fhkb:i41 fhkb:i42 fhkb:i43 fhkb:i44 fhkb:i45 fhkb:i46 fhkb:i47 fhkb:i48 fhkb:i49 fhkb:i50 fhkb:i51 fhkb:i52 fhkb:i53 fhkb:i54 fhkb:i55 ) ."
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"!tail onto.ttl"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"### Faire des inférences\n",
|
||||
"\n",
|
||||
"Nous souhaitons maintenant utiliser cette ontologie pour effectuer des inférences et des requêtes. Nous utiliserons [RDFLib](https://github.com/RDFLib), une bibliothèque permettant de lire des graphes RDF dans différents formats, de les interroger, etc.\n",
|
||||
"\n",
|
||||
"Pour les inférences logiques, nous utiliserons la bibliothèque [OWL-RL](https://github.com/RDFLib/OWL-RL), qui nous permet de construire la **Fermeture** du graphe RDF, c'est-à-dire d'ajouter tous les concepts et relations possibles qui peuvent être déduits.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 10,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Requirement already satisfied: rdflib in /home/rg/anaconda3/envs/ai4beg/lib/python3.11/site-packages (6.3.2)\n",
|
||||
"Requirement already satisfied: isodate<0.7.0,>=0.6.0 in /home/rg/anaconda3/envs/ai4beg/lib/python3.11/site-packages (from rdflib) (0.6.1)\n",
|
||||
"Requirement already satisfied: pyparsing<4,>=2.1.0 in /home/rg/anaconda3/envs/ai4beg/lib/python3.11/site-packages (from rdflib) (3.0.9)\n",
|
||||
"Requirement already satisfied: six in /home/rg/anaconda3/envs/ai4beg/lib/python3.11/site-packages (from isodate<0.7.0,>=0.6.0->rdflib) (1.16.0)\n",
|
||||
"Collecting git+https://github.com/RDFLib/OWL-RL.git\n",
|
||||
" Cloning https://github.com/RDFLib/OWL-RL.git to /tmp/pip-req-build-lbfzwi3m\n",
|
||||
" Running command git clone --filter=blob:none --quiet https://github.com/RDFLib/OWL-RL.git /tmp/pip-req-build-lbfzwi3m\n",
|
||||
" Resolved https://github.com/RDFLib/OWL-RL.git to commit a77e1791b88b54aace609bc6000aac14c7add4ff\n",
|
||||
" Preparing metadata (setup.py) ... \u001b[?25ldone\n",
|
||||
"\u001b[?25hRequirement already satisfied: rdflib>=6.0.2 in /home/rg/anaconda3/envs/ai4beg/lib/python3.11/site-packages (from owlrl==6.0.2) (6.3.2)\n",
|
||||
"Requirement already satisfied: isodate<0.7.0,>=0.6.0 in /home/rg/anaconda3/envs/ai4beg/lib/python3.11/site-packages (from rdflib>=6.0.2->owlrl==6.0.2) (0.6.1)\n",
|
||||
"Requirement already satisfied: pyparsing<4,>=2.1.0 in /home/rg/anaconda3/envs/ai4beg/lib/python3.11/site-packages (from rdflib>=6.0.2->owlrl==6.0.2) (3.0.9)\n",
|
||||
"Requirement already satisfied: six in /home/rg/anaconda3/envs/ai4beg/lib/python3.11/site-packages (from isodate<0.7.0,>=0.6.0->rdflib>=6.0.2->owlrl==6.0.2) (1.16.0)\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"!{sys.executable} -m pip install rdflib\n",
|
||||
"!{sys.executable} -m pip install git+https://github.com/RDFLib/OWL-RL.git"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Ouvrons le fichier d'ontologie et voyons combien de triplets il contient :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 11,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Triplets found:669\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"import rdflib\n",
|
||||
"from owlrl import DeductiveClosure, OWLRL_Extension\n",
|
||||
"\n",
|
||||
"g = rdflib.Graph()\n",
|
||||
"g.parse(\"onto.ttl\", format=\"turtle\")\n",
|
||||
"\n",
|
||||
"print(\"Triplets found:%d\" % len(g))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Maintenant, construisons la fermeture et voyons comment le nombre de triplets augmente :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 12,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Triplets after inference:4246\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"DeductiveClosure(OWLRL_Extension).expand(g)\n",
|
||||
"print(\"Triplets after inference:%d\" % len(g))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"### Interroger les relations familiales\n",
|
||||
"\n",
|
||||
"Nous pouvons maintenant interroger le graphe pour voir les différentes relations entre les personnes. Nous pouvons utiliser le langage **SPARQL** avec la méthode `query`. Dans notre cas, voyons tous les **oncles** dans notre arbre généalogique :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 13,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Fedor Alekseevich Romanov is uncle of Ekaterina Ivanovna Romanova\n",
|
||||
"Aleksandr I Pavlovich Romanov is uncle of Aleksandr II Nikolaevich Romanov\n",
|
||||
"Fedor Alekseevich Romanov is uncle of Anna Ivanovna Romanova\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"qres = g.query(\n",
|
||||
" \"\"\"SELECT DISTINCT ?aname ?bname\n",
|
||||
" WHERE {\n",
|
||||
" ?a fhkb:isUncleOf ?b .\n",
|
||||
" ?a rdfs:label ?aname .\n",
|
||||
" ?b rdfs:label ?bname .\n",
|
||||
" }\"\"\")\n",
|
||||
"\n",
|
||||
"for row in qres:\n",
|
||||
" print(\"%s is uncle of %s\" % row)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"N'hésitez pas à expérimenter avec d'autres relations familiales. Par exemple, vous pouvez examiner la relation `isAncestorOf`, qui définit de manière récursive tous les ancêtres d'une personne donnée.\n",
|
||||
"\n",
|
||||
"Enfin, passons au nettoyage !\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 14,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"!rm onto.ttl"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": []
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"\n---\n\n**Avertissement** : \nCe document a été traduit à l'aide du service de traduction automatique [Co-op Translator](https://github.com/Azure/co-op-translator). Bien que nous nous efforcions d'assurer l'exactitude, veuillez noter que les traductions automatisées peuvent contenir des erreurs ou des inexactitudes. Le document original dans sa langue d'origine doit être considéré comme la source faisant autorité. Pour des informations critiques, il est recommandé de recourir à une traduction professionnelle réalisée par un humain. Nous déclinons toute responsabilité en cas de malentendus ou d'interprétations erronées résultant de l'utilisation de cette traduction.\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
"interpreter": {
|
||||
"hash": "86193a1ab0ba47eac1c69c1756090baa3b420b3eea7d4aafab8b85f8b312f0c5"
|
||||
},
|
||||
"kernelspec": {
|
||||
"display_name": "Python 3.6",
|
||||
"language": "python",
|
||||
"name": "python3"
|
||||
},
|
||||
"language_info": {
|
||||
"codemirror_mode": {
|
||||
"name": "ipython",
|
||||
"version": 3
|
||||
},
|
||||
"file_extension": ".py",
|
||||
"mimetype": "text/x-python",
|
||||
"name": "python",
|
||||
"nbconvert_exporter": "python",
|
||||
"pygments_lexer": "ipython3",
|
||||
"version": "3.11.2"
|
||||
},
|
||||
"coopTranslator": {
|
||||
"original_hash": "6537d5597320e27b6052b4377b8ff8bb",
|
||||
"translation_date": "2025-08-31T14:53:22+00:00",
|
||||
"source_file": "lessons/2-Symbolic/FamilyOntology.ipynb",
|
||||
"language_code": "fr"
|
||||
}
|
||||
},
|
||||
"nbformat": 4,
|
||||
"nbformat_minor": 2
|
||||
}
|
||||
|
|
@ -0,0 +1,548 @@
|
|||
{
|
||||
"cells": [
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {
|
||||
"collapsed": true
|
||||
},
|
||||
"source": [
|
||||
"## Microsoft Concept Graph\n",
|
||||
"\n",
|
||||
"[Microsoft Concept Graph](https://concept.research.microsoft.com/) est une vaste taxonomie de termes extraits d'internet, avec des relations de type `is-a` entre les concepts.\n",
|
||||
"\n",
|
||||
"Le Context Graph est disponible sous deux formes :\n",
|
||||
" * Un fichier texte volumineux à télécharger\n",
|
||||
" * Une API REST\n",
|
||||
"\n",
|
||||
"Statistiques :\n",
|
||||
" * 5 401 933 concepts uniques,\n",
|
||||
" * 12 551 613 instances uniques,\n",
|
||||
" * 87 603 947 relations de type `is-a`\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Utilisation du service web\n",
|
||||
"\n",
|
||||
"Le service web propose différents appels pour estimer la probabilité qu'un concept appartienne à différents groupes. Plus d'informations sont disponibles [ici](https://concept.research.microsoft.com/Home/Api). \n",
|
||||
"Voici l'URL d'exemple pour effectuer un appel : `https://concept.research.microsoft.com/api/Concept/ScoreByProb?instance=microsoft&topK=10`\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 7,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"{'company': 0.6105356614382954,\n",
|
||||
" 'vendor': 0.08858636677518003,\n",
|
||||
" 'client': 0.048239124001183784,\n",
|
||||
" 'firm': 0.045476965571668145,\n",
|
||||
" 'large company': 0.043109401203511886,\n",
|
||||
" 'organization': 0.043010752688172046,\n",
|
||||
" 'corporation': 0.035908059583703265,\n",
|
||||
" 'brand': 0.03383644076156654,\n",
|
||||
" 'software company': 0.027522935779816515,\n",
|
||||
" 'technology company': 0.023774292196902438}"
|
||||
]
|
||||
},
|
||||
"execution_count": 7,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"import urllib\n",
|
||||
"import json\n",
|
||||
"import ssl\n",
|
||||
"\n",
|
||||
"def http(x):\n",
|
||||
" ssl._create_default_https_context = ssl._create_unverified_context\n",
|
||||
" response = urllib.request.urlopen(x)\n",
|
||||
" data = response.read()\n",
|
||||
" return data.decode('utf-8')\n",
|
||||
"\n",
|
||||
"def query(x):\n",
|
||||
" return json.loads(http(\"https://concept.research.microsoft.com/api/Concept/ScoreByProb?instance={}&topK=10\".format(urllib.parse.quote(x))))\n",
|
||||
"\n",
|
||||
"query('microsoft')"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Essayons de catégoriser les titres des actualités en utilisant des concepts parentaux. Pour obtenir les titres des actualités, nous utiliserons le service [NewsApi.org](http://newsapi.org). Vous devez obtenir votre propre clé API pour utiliser le service - rendez-vous sur le site web et inscrivez-vous au plan développeur gratuit.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 20,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"newsapi_key = '<your API key here>'\n",
|
||||
"def get_news(country='us'):\n",
|
||||
" res = json.loads(http(\"https://newsapi.org/v2/top-headlines?country={0}&apiKey={1}\".format(country,newsapi_key)))\n",
|
||||
" return res['articles']\n",
|
||||
"\n",
|
||||
"all_titles = [x['title'] for x in get_news('us')+get_news('gb')]"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 21,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"['Covid-19 Live Updates: Vaccines and Boosters News - The New York Times',\n",
|
||||
" 'Ukrainians Flee Mariupol as Russian Forces Push to Take Port City - The Wall Street Journal',\n",
|
||||
" 'Bond Yields Jump, Stock Futures Rise After Powell Says Fed Is Ready to Be More Aggressive - The Wall Street Journal',\n",
|
||||
" 'Putin critic Alexei Navalny found guilty by Russian court - New York Post ',\n",
|
||||
" \"Supreme Court nominee Ketanji Brown Jackson will face questions at confirmation hearing's second day - CNN\",\n",
|
||||
" '2 teachers killed at Swedish high school, student arrested - ABC News',\n",
|
||||
" 'Clues to Covid-19’s Next Moves Come From Sewers - The Wall Street Journal',\n",
|
||||
" 'Republicans to roll dice by grilling Jackson over child-pornography sentencing decisions | TheHill - The Hill',\n",
|
||||
" '‘Clear sign’ Putin considering using chemical weapons in Ukraine, claims President Biden - The Independent',\n",
|
||||
" 'NASA confirms there are 5,000 planets outside our solar system - Daily Mail',\n",
|
||||
" \"US stocks whipsawed overnight after Fed Chair Powell's remarks - Fox Business\",\n",
|
||||
" \"'We've learned absolutely nothing': Tests could again be in short supply if Covid surges - POLITICO\",\n",
|
||||
" \"Duchess of Cambridge swaps khaki jungle gear for Vampire's Wife dress on Belize trip - Daily Mail\",\n",
|
||||
" 'China searches for victims, flight recorders after first plane crash in 12 years - Reuters',\n",
|
||||
" 'Second superyacht linked to Russian oligarch Abramovich docks in Turkey - Reuters',\n",
|
||||
" 'Live updates: Russia stops talks with Japan over sanctions - The Associated Press - en Español',\n",
|
||||
" 'Powers Remain and Threats Lurk as Women’s Sweet 16 Is Set - The New York Times',\n",
|
||||
" 'Webb Space Telescope Begins Multi-Instrument Alignment - SciTechDaily',\n",
|
||||
" \"UConn vs UCF - NCAA women's tournament second-round highlights - March Madness\",\n",
|
||||
" 'Bucking Republican Trend, Indiana Governor Vetoes Transgender Sports Bill - The New York Times',\n",
|
||||
" \"Maggie Fox dead: Coronation Street and Shameless actress dies after 'sudden accident' - Mirror Online - The Mirror\",\n",
|
||||
" 'China plane crash – live: Search for survivors continues as witness describes moment flight fell from sky - The Independent',\n",
|
||||
" 'Daniel Morgan murder: damning report condemns Met police - The Guardian',\n",
|
||||
" 'What to expect from Rishi Sunak’s Spring Statement - BBC.com',\n",
|
||||
" 'UK and Republic of Ireland in line to host Euro 2028 after no one else bids - The Guardian',\n",
|
||||
" \"Friends beg Vladimir Putin's 'lover' to persuade him to end Ukraine invasion - The Mirror\",\n",
|
||||
" 'Brass Eye’s outtakes show the brutal TV comedy was the tip of an iceberg - The Guardian',\n",
|
||||
" \"Vladimir Putin threatens civilians to break Mariupol's spirit - The Times\",\n",
|
||||
" 'Shell U-turn on Cambo oilfield would threaten green targets, say campaigners - The Guardian',\n",
|
||||
" 'St Helens dog attack: Girl aged 17 months killed at home - BBC',\n",
|
||||
" \"PlayStation to buy 'Assassin's Creed' veteran Jade Raymond's Haven Studios - NME\",\n",
|
||||
" '‘Clear sign’ Putin considering using chemical weapons in Ukraine, claims President Biden - The Independent',\n",
|
||||
" 'NASA confirms there are 5,000 planets outside our solar system - Daily Mail',\n",
|
||||
" 'Nintendo Switch finally has folders • Eurogamer.net - Eurogamer.net',\n",
|
||||
" 'FA to “find a solution” as Liverpool fan group blasts “shambolic” Wembley travel - This Is Anfield',\n",
|
||||
" 'Manchester United transfer news LIVE Erik ten Hag latest and Man Utd manager updates - Manchester Evening News',\n",
|
||||
" 'Inflation raises cost of UK government borrowing in February; crude oil up again – business live - The Guardian',\n",
|
||||
" 'Alexei Navalny: Kremlin critic found guilty of large-scale fraud and contempt of court by Russian court - Sky News',\n",
|
||||
" \"UK prepares to nationalize Russia natural gas giant Gazprom's retail unit - Business Insider\",\n",
|
||||
" 'Zaghari-Ratcliffe: Hunt calls for inquiry into delay over Iran debt payment - The Guardian']"
|
||||
]
|
||||
},
|
||||
"execution_count": 21,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"all_titles"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Tout d'abord, nous voulons pouvoir extraire des noms des titres d'actualités. Nous utiliserons la bibliothèque `TextBlob` pour cela, ce qui simplifie beaucoup de tâches typiques de NLP comme celle-ci.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 15,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Requirement already satisfied: textblob in c:\\winapp\\miniconda3\\lib\\site-packages (0.17.1)\n",
|
||||
"Requirement already satisfied: nltk>=3.1 in c:\\winapp\\miniconda3\\lib\\site-packages (from textblob) (3.5)\n",
|
||||
"Requirement already satisfied: joblib in c:\\winapp\\miniconda3\\lib\\site-packages (from nltk>=3.1->textblob) (1.0.1)\n",
|
||||
"Requirement already satisfied: regex in c:\\winapp\\miniconda3\\lib\\site-packages (from nltk>=3.1->textblob) (2021.11.10)\n",
|
||||
"Requirement already satisfied: tqdm in c:\\winapp\\miniconda3\\lib\\site-packages (from nltk>=3.1->textblob) (4.61.2)\n",
|
||||
"Requirement already satisfied: click in c:\\winapp\\miniconda3\\lib\\site-packages (from nltk>=3.1->textblob) (8.0.3)\n",
|
||||
"Requirement already satisfied: colorama in c:\\winapp\\miniconda3\\lib\\site-packages (from click->nltk>=3.1->textblob) (0.4.4)\n",
|
||||
"Finished.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"name": "stderr",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"[nltk_data] Downloading package brown to\n",
|
||||
"[nltk_data] C:\\Users\\dmitryso\\AppData\\Roaming\\nltk_data...\n",
|
||||
"[nltk_data] Package brown is already up-to-date!\n",
|
||||
"[nltk_data] Downloading package punkt to\n",
|
||||
"[nltk_data] C:\\Users\\dmitryso\\AppData\\Roaming\\nltk_data...\n",
|
||||
"[nltk_data] Package punkt is already up-to-date!\n",
|
||||
"[nltk_data] Downloading package wordnet to\n",
|
||||
"[nltk_data] C:\\Users\\dmitryso\\AppData\\Roaming\\nltk_data...\n",
|
||||
"[nltk_data] Package wordnet is already up-to-date!\n",
|
||||
"[nltk_data] Downloading package averaged_perceptron_tagger to\n",
|
||||
"[nltk_data] C:\\Users\\dmitryso\\AppData\\Roaming\\nltk_data...\n",
|
||||
"[nltk_data] Package averaged_perceptron_tagger is already up-to-\n",
|
||||
"[nltk_data] date!\n",
|
||||
"[nltk_data] Downloading package conll2000 to\n",
|
||||
"[nltk_data] C:\\Users\\dmitryso\\AppData\\Roaming\\nltk_data...\n",
|
||||
"[nltk_data] Package conll2000 is already up-to-date!\n",
|
||||
"[nltk_data] Downloading package movie_reviews to\n",
|
||||
"[nltk_data] C:\\Users\\dmitryso\\AppData\\Roaming\\nltk_data...\n",
|
||||
"[nltk_data] Package movie_reviews is already up-to-date!\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"import sys\n",
|
||||
"!{sys.executable} -m pip install textblob\n",
|
||||
"!{sys.executable} -m textblob.download_corpora\n",
|
||||
"from textblob import TextBlob"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 22,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"{'covid-19 live updates': 1,\n",
|
||||
" 'vaccines': 1,\n",
|
||||
" 'boosters': 1,\n",
|
||||
" 'york': 4,\n",
|
||||
" 'ukrainians flee mariupol': 1,\n",
|
||||
" 'forces push': 1,\n",
|
||||
" 'port city': 1,\n",
|
||||
" 'wall street journal': 3,\n",
|
||||
" 'bond yields': 1,\n",
|
||||
" 'futures rise': 1,\n",
|
||||
" 'powell says fed': 1,\n",
|
||||
" 'ready': 1,\n",
|
||||
" 'be': 1,\n",
|
||||
" 'aggressive': 1,\n",
|
||||
" 'putin': 3,\n",
|
||||
" 'alexei navalny': 2,\n",
|
||||
" 'russian': 2,\n",
|
||||
" 'supreme court nominee': 1,\n",
|
||||
" 'ketanji brown jackson': 1,\n",
|
||||
" \"confirmation hearing 's\": 1,\n",
|
||||
" 'cnn': 1,\n",
|
||||
" 'swedish': 1,\n",
|
||||
" 'high school': 1,\n",
|
||||
" 'abc': 1,\n",
|
||||
" 'clues': 1,\n",
|
||||
" 'covid-19': 1,\n",
|
||||
" '’ s': 2,\n",
|
||||
" 'moves': 1,\n",
|
||||
" 'sewers': 1,\n",
|
||||
" 'roll dice': 1,\n",
|
||||
" 'jackson': 1,\n",
|
||||
" 'decisions |': 1,\n",
|
||||
" 'thehill': 1,\n",
|
||||
" 'clear': 2,\n",
|
||||
" 'chemical weapons': 2,\n",
|
||||
" 'ukraine': 3,\n",
|
||||
" 'claims president': 2,\n",
|
||||
" 'biden': 2,\n",
|
||||
" 'nasa': 2,\n",
|
||||
" 'solar system': 2,\n",
|
||||
" 'daily mail': 3,\n",
|
||||
" 'us stocks': 1,\n",
|
||||
" 'fed chair powell': 1,\n",
|
||||
" \"'s remarks\": 1,\n",
|
||||
" 'fox': 1,\n",
|
||||
" \"'we 've\": 1,\n",
|
||||
" 'tests': 1,\n",
|
||||
" 'covid': 1,\n",
|
||||
" 'politico': 1,\n",
|
||||
" 'duchess': 1,\n",
|
||||
" 'cambridge': 1,\n",
|
||||
" 'swaps khaki jungle gear': 1,\n",
|
||||
" 'vampire': 1,\n",
|
||||
" 'wife': 1,\n",
|
||||
" 'belize': 1,\n",
|
||||
" 'china': 2,\n",
|
||||
" 'flight recorders': 1,\n",
|
||||
" 'plane crash': 1,\n",
|
||||
" 'reuters': 2,\n",
|
||||
" 'russian oligarch': 1,\n",
|
||||
" 'abramovich': 1,\n",
|
||||
" 'live': 1,\n",
|
||||
" 'russia': 2,\n",
|
||||
" 'stops talks': 1,\n",
|
||||
" 'japan': 1,\n",
|
||||
" 'español': 1,\n",
|
||||
" 'powers remain': 1,\n",
|
||||
" 'threats lurk': 1,\n",
|
||||
" 'set': 1,\n",
|
||||
" 'webb': 1,\n",
|
||||
" 'telescope begins multi-instrument alignment': 1,\n",
|
||||
" 'scitechdaily': 1,\n",
|
||||
" 'uconn': 1,\n",
|
||||
" 'ucf': 1,\n",
|
||||
" 'ncaa': 1,\n",
|
||||
" \"women 's tournament second-round highlights\": 1,\n",
|
||||
" 'march madness': 1,\n",
|
||||
" 'bucking republican trend': 1,\n",
|
||||
" 'indiana': 1,\n",
|
||||
" 'vetoes transgender': 1,\n",
|
||||
" 'bill': 1,\n",
|
||||
" 'maggie fox': 1,\n",
|
||||
" 'coronation': 1,\n",
|
||||
" 'shameless': 1,\n",
|
||||
" \"'sudden accident\": 1,\n",
|
||||
" 'mirror online': 1,\n",
|
||||
" 'mirror': 2,\n",
|
||||
" 'plane crash –': 1,\n",
|
||||
" 'search': 1,\n",
|
||||
" 'moment flight': 1,\n",
|
||||
" 'daniel morgan': 1,\n",
|
||||
" 'report condemns': 1,\n",
|
||||
" 'met': 1,\n",
|
||||
" 'guardian': 6,\n",
|
||||
" 'rishi sunak': 1,\n",
|
||||
" '’ s spring': 1,\n",
|
||||
" 'statement': 1,\n",
|
||||
" 'bbc.com': 1,\n",
|
||||
" 'uk': 3,\n",
|
||||
" 'ireland': 1,\n",
|
||||
" 'euro': 1,\n",
|
||||
" 'vladimir putin': 2,\n",
|
||||
" \"'s 'lover\": 1,\n",
|
||||
" 'brass eye': 1,\n",
|
||||
" '’ s outtakes': 1,\n",
|
||||
" 'brutal tv comedy': 1,\n",
|
||||
" 'threatens civilians': 1,\n",
|
||||
" 'mariupol': 1,\n",
|
||||
" \"'s spirit\": 1,\n",
|
||||
" 'shell u-turn': 1,\n",
|
||||
" 'cambo': 1,\n",
|
||||
" 'green targets': 1,\n",
|
||||
" 'st helens': 1,\n",
|
||||
" 'dog attack': 1,\n",
|
||||
" 'girl': 1,\n",
|
||||
" 'bbc': 1,\n",
|
||||
" 'playstation': 1,\n",
|
||||
" \"'assassin 's\": 1,\n",
|
||||
" 'creed': 1,\n",
|
||||
" 'jade raymond': 1,\n",
|
||||
" 'haven studios': 1,\n",
|
||||
" 'nme': 1,\n",
|
||||
" 'nintendo switch': 1,\n",
|
||||
" 'folders •': 1,\n",
|
||||
" 'eurogamer.net': 2,\n",
|
||||
" 'fa': 1,\n",
|
||||
" 'solution ”': 1,\n",
|
||||
" 'liverpool': 1,\n",
|
||||
" 'fan group blasts “ shambolic ”': 1,\n",
|
||||
" 'wembley': 1,\n",
|
||||
" 'anfield': 1,\n",
|
||||
" 'manchester': 1,\n",
|
||||
" 'live erik': 1,\n",
|
||||
" 'hag': 1,\n",
|
||||
" 'utd': 1,\n",
|
||||
" 'manager updates': 1,\n",
|
||||
" 'manchester evening': 1,\n",
|
||||
" 'inflation': 1,\n",
|
||||
" 'government borrowing': 1,\n",
|
||||
" 'february': 1,\n",
|
||||
" 'crude oil': 1,\n",
|
||||
" '– business': 1,\n",
|
||||
" 'kremlin': 1,\n",
|
||||
" 'large-scale fraud': 1,\n",
|
||||
" 'sky': 1,\n",
|
||||
" 'natural gas': 1,\n",
|
||||
" 'gazprom': 1,\n",
|
||||
" 'retail unit': 1,\n",
|
||||
" 'insider': 1,\n",
|
||||
" 'zaghari-ratcliffe': 1,\n",
|
||||
" 'hunt': 1,\n",
|
||||
" 'iran': 1,\n",
|
||||
" 'debt payment': 1}"
|
||||
]
|
||||
},
|
||||
"execution_count": 22,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"w = {}\n",
|
||||
"for x in all_titles:\n",
|
||||
" for n in TextBlob(x).noun_phrases:\n",
|
||||
" if n in w:\n",
|
||||
" w[n].append(x)\n",
|
||||
" else:\n",
|
||||
" w[n]=[x]\n",
|
||||
"{ x:len(w[x]) for x in w.keys()}"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Nous pouvons voir que les noms ne nous donnent pas de grands groupes thématiques. Remplaçons les noms par des termes plus généraux obtenus à partir du graphe de concepts. Cela prendra du temps, car nous effectuons un appel REST pour chaque syntagme nominal.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 23,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"w = {}\n",
|
||||
"for x in all_titles:\n",
|
||||
" for noun in TextBlob(x).noun_phrases:\n",
|
||||
" terms = query(noun.replace(' ','%20'))\n",
|
||||
" for term in [u for u in terms.keys() if terms[u]>0.1]:\n",
|
||||
" if term in w:\n",
|
||||
" w[term].append(x)\n",
|
||||
" else:\n",
|
||||
" w[term]=[x]"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 24,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"{'city': 9,\n",
|
||||
" 'brand': 4,\n",
|
||||
" 'place': 9,\n",
|
||||
" 'town': 4,\n",
|
||||
" 'factor': 4,\n",
|
||||
" 'film': 4,\n",
|
||||
" 'nation': 11,\n",
|
||||
" 'state': 5,\n",
|
||||
" 'person': 4,\n",
|
||||
" 'organization': 5,\n",
|
||||
" 'publication': 10,\n",
|
||||
" 'market': 5,\n",
|
||||
" 'economy': 4,\n",
|
||||
" 'company': 6,\n",
|
||||
" 'newspaper': 6,\n",
|
||||
" 'relationship': 6}"
|
||||
]
|
||||
},
|
||||
"execution_count": 24,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"{ x:len(w[x]) for x in w.keys() if len(w[x])>3}"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 27,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"\n",
|
||||
"ECONOMY:\n",
|
||||
"China searches for victims, flight recorders after first plane crash in 12 years - Reuters\n",
|
||||
"Live updates: Russia stops talks with Japan over sanctions - The Associated Press - en Español\n",
|
||||
"China plane crash – live: Search for survivors continues as witness describes moment flight fell from sky - The Independent\n",
|
||||
"UK prepares to nationalize Russia natural gas giant Gazprom's retail unit - Business Insider\n",
|
||||
"\n",
|
||||
"NATION:\n",
|
||||
"‘Clear sign’ Putin considering using chemical weapons in Ukraine, claims President Biden - The Independent\n",
|
||||
"Duchess of Cambridge swaps khaki jungle gear for Vampire's Wife dress on Belize trip - Daily Mail\n",
|
||||
"China searches for victims, flight recorders after first plane crash in 12 years - Reuters\n",
|
||||
"Live updates: Russia stops talks with Japan over sanctions - The Associated Press - en Español\n",
|
||||
"Live updates: Russia stops talks with Japan over sanctions - The Associated Press - en Español\n",
|
||||
"China plane crash – live: Search for survivors continues as witness describes moment flight fell from sky - The Independent\n",
|
||||
"UK and Republic of Ireland in line to host Euro 2028 after no one else bids - The Guardian\n",
|
||||
"Friends beg Vladimir Putin's 'lover' to persuade him to end Ukraine invasion - The Mirror\n",
|
||||
"‘Clear sign’ Putin considering using chemical weapons in Ukraine, claims President Biden - The Independent\n",
|
||||
"UK prepares to nationalize Russia natural gas giant Gazprom's retail unit - Business Insider\n",
|
||||
"Zaghari-Ratcliffe: Hunt calls for inquiry into delay over Iran debt payment - The Guardian\n",
|
||||
"\n",
|
||||
"PERSON:\n",
|
||||
"‘Clear sign’ Putin considering using chemical weapons in Ukraine, claims President Biden - The Independent\n",
|
||||
"Duchess of Cambridge swaps khaki jungle gear for Vampire's Wife dress on Belize trip - Daily Mail\n",
|
||||
"Second superyacht linked to Russian oligarch Abramovich docks in Turkey - Reuters\n",
|
||||
"‘Clear sign’ Putin considering using chemical weapons in Ukraine, claims President Biden - The Independent\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"print('\\nECONOMY:\\n'+'\\n'.join(w['economy']))\n",
|
||||
"print('\\nNATION:\\n'+'\\n'.join(w['nation']))\n",
|
||||
"print('\\nPERSON:\\n'+'\\n'.join(w['person']))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"\n---\n\n**Avertissement** : \nCe document a été traduit à l'aide du service de traduction automatique [Co-op Translator](https://github.com/Azure/co-op-translator). Bien que nous nous efforcions d'assurer l'exactitude, veuillez noter que les traductions automatisées peuvent contenir des erreurs ou des inexactitudes. Le document original dans sa langue d'origine doit être considéré comme la source faisant autorité. Pour des informations critiques, il est recommandé de recourir à une traduction professionnelle réalisée par un humain. Nous déclinons toute responsabilité en cas de malentendus ou d'interprétations erronées résultant de l'utilisation de cette traduction.\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
"kernelspec": {
|
||||
"display_name": "Python 3.7.4 64-bit (conda)",
|
||||
"metadata": {
|
||||
"interpreter": {
|
||||
"hash": "86193a1ab0ba47eac1c69c1756090baa3b420b3eea7d4aafab8b85f8b312f0c5"
|
||||
}
|
||||
},
|
||||
"name": "python3"
|
||||
},
|
||||
"language_info": {
|
||||
"codemirror_mode": {
|
||||
"name": "ipython",
|
||||
"version": 3
|
||||
},
|
||||
"file_extension": ".py",
|
||||
"mimetype": "text/x-python",
|
||||
"name": "python",
|
||||
"nbconvert_exporter": "python",
|
||||
"pygments_lexer": "ipython3",
|
||||
"version": "3.9.5"
|
||||
},
|
||||
"coopTranslator": {
|
||||
"original_hash": "4087f998407d06ceb2947016ba4605d0",
|
||||
"translation_date": "2025-08-31T14:54:02+00:00",
|
||||
"source_file": "lessons/2-Symbolic/MSConceptGraph.ipynb",
|
||||
"language_code": "fr"
|
||||
}
|
||||
},
|
||||
"nbformat": 4,
|
||||
"nbformat_minor": 2
|
||||
}
|
||||
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
|
|
@ -1,15 +1,15 @@
|
|||
<!--
|
||||
CO_OP_TRANSLATOR_METADATA:
|
||||
{
|
||||
"original_hash": "7336583e4630220c835335da640016db",
|
||||
"translation_date": "2025-08-24T20:57:00+00:00",
|
||||
"original_hash": "ba5d1eb353d20d3e7181066b3c424b99",
|
||||
"translation_date": "2025-08-31T14:20:13+00:00",
|
||||
"source_file": "lessons/3-NeuralNetworks/03-Perceptron/lab/README.md",
|
||||
"language_code": "fr"
|
||||
}
|
||||
-->
|
||||
# Classification multi-classes avec Perceptron
|
||||
|
||||
Travail pratique issu du [Curriculum AI pour Débutants](https://github.com/microsoft/ai-for-beginners).
|
||||
Travail pratique tiré du [Curriculum AI pour les débutants](https://github.com/microsoft/ai-for-beginners).
|
||||
|
||||
## Tâche
|
||||
|
||||
|
|
@ -21,11 +21,13 @@ En utilisant le code que nous avons développé dans cette leçon pour la classi
|
|||
1. Entraînez 10 perceptrons différents pour la classification binaire (un pour chaque chiffre).
|
||||
1. Définissez une fonction qui classera un chiffre donné en entrée.
|
||||
|
||||
> **Conseil** : Si nous combinons les poids des 10 perceptrons dans une seule matrice, nous devrions pouvoir appliquer les 10 perceptrons aux chiffres en entrée par une seule multiplication matricielle. Le chiffre le plus probable peut ensuite être déterminé simplement en appliquant l'opération `argmax` sur le résultat.
|
||||
> **Conseil** : Si nous combinons les poids des 10 perceptrons dans une seule matrice, nous devrions être capables d'appliquer les 10 perceptrons aux chiffres en entrée par une seule multiplication matricielle. Le chiffre le plus probable peut ensuite être trouvé simplement en appliquant l'opération `argmax` sur la sortie.
|
||||
|
||||
## Notebook de départ
|
||||
|
||||
Commencez le travail pratique en ouvrant [PerceptronMultiClass.ipynb](../../../../../../lessons/3-NeuralNetworks/03-Perceptron/lab/PerceptronMultiClass.ipynb)
|
||||
Commencez le travail pratique en ouvrant [PerceptronMultiClass.ipynb](PerceptronMultiClass.ipynb)
|
||||
|
||||
---
|
||||
|
||||
**Avertissement** :
|
||||
Ce document a été traduit à l'aide du service de traduction automatique [Co-op Translator](https://github.com/Azure/co-op-translator). Bien que nous nous efforcions d'assurer l'exactitude, veuillez noter que les traductions automatisées peuvent contenir des erreurs ou des inexactitudes. Le document original dans sa langue d'origine doit être considéré comme la source faisant autorité. Pour des informations critiques, il est recommandé de recourir à une traduction humaine professionnelle. Nous déclinons toute responsabilité en cas de malentendus ou d'interprétations erronées résultant de l'utilisation de cette traduction.
|
||||
File diff suppressed because one or more lines are too long
|
|
@ -0,0 +1,183 @@
|
|||
{
|
||||
"cells": [
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"# Classification des chiffres MNIST avec notre propre framework\n",
|
||||
"\n",
|
||||
"Travail pratique issu du [programme AI for Beginners](https://github.com/microsoft/ai-for-beginners).\n",
|
||||
"\n",
|
||||
"### Lecture du jeu de données\n",
|
||||
"\n",
|
||||
"Ce code télécharge le jeu de données depuis le dépôt sur Internet. Vous pouvez également copier manuellement le jeu de données depuis le répertoire `/data` du dépôt AI Curriculum.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 4,
|
||||
"metadata": {
|
||||
"tags": []
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stderr",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
" % Total % Received % Xferd Average Speed Time Time Time Current\n",
|
||||
" Dload Upload Total Spent Left Speed\n",
|
||||
"\n",
|
||||
" 0 0 0 0 0 0 0 0 --:--:-- --:--:-- --:--:-- 0\n",
|
||||
"100 9.9M 100 9.9M 0 0 9.9M 0 0:00:01 --:--:-- 0:00:01 15.8M\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"!rm *.pkl\n",
|
||||
"!wget https://raw.githubusercontent.com/microsoft/AI-For-Beginners/main/data/mnist.pkl.gz\n",
|
||||
"!gzip -d mnist.pkl.gz"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 3,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"import pickle\n",
|
||||
"with open('mnist.pkl','rb') as f:\n",
|
||||
" MNIST = pickle.load(f)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"labels = MNIST['Train']['Labels']\n",
|
||||
"data = MNIST['Train']['Features']"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Voyons quelle est la forme des données que nous avons :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 5,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"(42000, 784)"
|
||||
]
|
||||
},
|
||||
"execution_count": 5,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"data.shape"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"### Séparation des données\n",
|
||||
"\n",
|
||||
"Nous utiliserons Scikit Learn pour diviser les données entre le jeu d'entraînement et le jeu de test :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 6,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Train samples: 33600, test samples: 8400\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"from sklearn.model_selection import train_test_split\n",
|
||||
"\n",
|
||||
"features_train, features_test, labels_train, labels_test = train_test_split(data,labels,test_size=0.2)\n",
|
||||
"\n",
|
||||
"print(f\"Train samples: {len(features_train)}, test samples: {len(features_test)}\")"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"### Instructions\n",
|
||||
"\n",
|
||||
"1. Prenez le code du framework de la leçon et collez-le dans ce notebook, ou (encore mieux) dans un module Python séparé.\n",
|
||||
"1. Définissez et entraînez un perceptron à une seule couche, en observant la précision de l'entraînement et de la validation pendant l'entraînement.\n",
|
||||
"1. Essayez de comprendre si un surapprentissage a eu lieu, et ajustez les paramètres de la couche pour améliorer la précision.\n",
|
||||
"1. Répétez les étapes précédentes pour des perceptrons à 2 et 3 couches. Essayez d'expérimenter avec différentes fonctions d'activation entre les couches.\n",
|
||||
"1. Essayez de répondre aux questions suivantes :\n",
|
||||
" - La fonction d'activation entre les couches affecte-t-elle les performances du réseau ?\n",
|
||||
" - Avons-nous besoin d'un réseau à 2 ou 3 couches pour cette tâche ?\n",
|
||||
" - Avez-vous rencontré des problèmes lors de l'entraînement du réseau ? En particulier lorsque le nombre de couches augmentait.\n",
|
||||
" - Comment se comportent les poids du réseau pendant l'entraînement ? Vous pouvez tracer la valeur absolue maximale des poids en fonction des époques pour comprendre la relation.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": []
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"\n---\n\n**Avertissement** : \nCe document a été traduit à l'aide du service de traduction automatique [Co-op Translator](https://github.com/Azure/co-op-translator). Bien que nous nous efforcions d'assurer l'exactitude, veuillez noter que les traductions automatisées peuvent contenir des erreurs ou des inexactitudes. Le document original dans sa langue d'origine doit être considéré comme la source faisant autorité. Pour des informations critiques, il est recommandé de recourir à une traduction professionnelle effectuée par un humain. Nous déclinons toute responsabilité en cas de malentendus ou d'interprétations erronées résultant de l'utilisation de cette traduction.\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
"kernelspec": {
|
||||
"display_name": "Python 3.7.4 64-bit (conda)",
|
||||
"metadata": {
|
||||
"interpreter": {
|
||||
"hash": "86193a1ab0ba47eac1c69c1756090baa3b420b3eea7d4aafab8b85f8b312f0c5"
|
||||
}
|
||||
},
|
||||
"name": "python3"
|
||||
},
|
||||
"language_info": {
|
||||
"codemirror_mode": {
|
||||
"name": "ipython",
|
||||
"version": 3
|
||||
},
|
||||
"file_extension": ".py",
|
||||
"mimetype": "text/x-python",
|
||||
"name": "python",
|
||||
"nbconvert_exporter": "python",
|
||||
"pygments_lexer": "ipython3",
|
||||
"version": "3.9.5"
|
||||
},
|
||||
"orig_nbformat": 2,
|
||||
"coopTranslator": {
|
||||
"original_hash": "6fa055f484eb5d6bdf41166a356d3abf",
|
||||
"translation_date": "2025-08-31T14:58:32+00:00",
|
||||
"source_file": "lessons/3-NeuralNetworks/04-OwnFramework/lab/MyFW_MNIST.ipynb",
|
||||
"language_code": "fr"
|
||||
}
|
||||
},
|
||||
"nbformat": 4,
|
||||
"nbformat_minor": 2
|
||||
}
|
||||
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
|
|
@ -0,0 +1,102 @@
|
|||
{
|
||||
"cells": [
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"**Votre objectif** sera d'utiliser le flux optique pour déterminer quelles parties de la vidéo contiennent des mouvements vers le haut, le bas, la gauche ou la droite.\n",
|
||||
"\n",
|
||||
"Commencez par obtenir les images de la vidéo comme expliqué dans le cours :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"# Code here"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Maintenant, calculez les cadres de flux optique dense comme décrit dans le cours, et convertissez le flux optique dense en coordonnées polaires :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"# Code here"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Construire un histogramme des directions pour chaque image du flux optique. Un histogramme montre combien de vecteurs se trouvent dans une certaine plage, et il doit distinguer les différentes directions de mouvement dans l'image.\n",
|
||||
"\n",
|
||||
"> Vous pouvez également vouloir annuler tous les vecteurs dont la magnitude est inférieure à un certain seuil. Cela permettra d'éliminer les petits mouvements parasites dans la vidéo, comme ceux des yeux et de la tête.\n",
|
||||
"\n",
|
||||
"Tracer les histogrammes pour certaines des images.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"# Code here"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"En regardant les histogrammes, il devrait être assez simple de déterminer la direction du mouvement. Vous devez sélectionner les barres qui correspondent aux directions haut/bas/gauche/droite, et qui sont au-dessus d'un certain seuil.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"# Code here"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Félicitations ! Si vous avez suivi toutes les étapes ci-dessus, vous avez terminé le laboratoire !\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"\n---\n\n**Avertissement** : \nCe document a été traduit à l'aide du service de traduction automatique [Co-op Translator](https://github.com/Azure/co-op-translator). Bien que nous nous efforcions d'assurer l'exactitude, veuillez noter que les traductions automatisées peuvent contenir des erreurs ou des inexactitudes. Le document original dans sa langue d'origine doit être considéré comme la source faisant autorité. Pour des informations critiques, il est recommandé de recourir à une traduction professionnelle effectuée par un humain. Nous déclinons toute responsabilité en cas de malentendus ou d'interprétations erronées résultant de l'utilisation de cette traduction.\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
"language_info": {
|
||||
"name": "python"
|
||||
},
|
||||
"orig_nbformat": 4,
|
||||
"coopTranslator": {
|
||||
"original_hash": "153d9e417e079bf62f8f693002d0deaf",
|
||||
"translation_date": "2025-08-31T14:42:38+00:00",
|
||||
"source_file": "lessons/4-ComputerVision/06-IntroCV/lab/MovementDetection.ipynb",
|
||||
"language_code": "fr"
|
||||
}
|
||||
},
|
||||
"nbformat": 4,
|
||||
"nbformat_minor": 2
|
||||
}
|
||||
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
|
|
@ -0,0 +1,577 @@
|
|||
{
|
||||
"cells": [
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"# Tâche de classification de texte\n",
|
||||
"\n",
|
||||
"Comme nous l'avons mentionné, nous allons nous concentrer sur une tâche simple de classification de texte basée sur le dataset **AG_NEWS**, qui consiste à classer les titres d'actualités dans l'une des 4 catégories : Monde, Sports, Économie et Sci/Tech.\n",
|
||||
"\n",
|
||||
"## Le Dataset\n",
|
||||
"\n",
|
||||
"Ce dataset est intégré dans le module [`torchtext`](https://github.com/pytorch/text), ce qui nous permet d'y accéder facilement.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 1,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"import torch\n",
|
||||
"import torchtext\n",
|
||||
"import os\n",
|
||||
"import collections\n",
|
||||
"os.makedirs('./data',exist_ok=True)\n",
|
||||
"train_dataset, test_dataset = torchtext.datasets.AG_NEWS(root='./data')\n",
|
||||
"classes = ['World', 'Sports', 'Business', 'Sci/Tech']"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Ici, `train_dataset` et `test_dataset` contiennent des collections qui renvoient respectivement des paires d'étiquette (numéro de classe) et de texte, par exemple :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 2,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"(3,\n",
|
||||
" \"Wall St. Bears Claw Back Into the Black (Reuters) Reuters - Short-sellers, Wall Street's dwindling\\\\band of ultra-cynics, are seeing green again.\")"
|
||||
]
|
||||
},
|
||||
"execution_count": 2,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"list(train_dataset)[0]"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Alors, imprimons les 10 premiers nouveaux titres de notre ensemble de données :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 5,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"**Sci/Tech** -> Wall St. Bears Claw Back Into the Black (Reuters) Reuters - Short-sellers, Wall Street's dwindling\\band of ultra-cynics, are seeing green again.\n",
|
||||
"**Sci/Tech** -> Carlyle Looks Toward Commercial Aerospace (Reuters) Reuters - Private investment firm Carlyle Group,\\which has a reputation for making well-timed and occasionally\\controversial plays in the defense industry, has quietly placed\\its bets on another part of the market.\n",
|
||||
"**Sci/Tech** -> Oil and Economy Cloud Stocks' Outlook (Reuters) Reuters - Soaring crude prices plus worries\\about the economy and the outlook for earnings are expected to\\hang over the stock market next week during the depth of the\\summer doldrums.\n",
|
||||
"**Sci/Tech** -> Iraq Halts Oil Exports from Main Southern Pipeline (Reuters) Reuters - Authorities have halted oil export\\flows from the main pipeline in southern Iraq after\\intelligence showed a rebel militia could strike\\infrastructure, an oil official said on Saturday.\n",
|
||||
"**Sci/Tech** -> Oil prices soar to all-time record, posing new menace to US economy (AFP) AFP - Tearaway world oil prices, toppling records and straining wallets, present a new economic menace barely three months before the US presidential elections.\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"for i,x in zip(range(5),train_dataset):\n",
|
||||
" print(f\"**{classes[x[0]]}** -> {x[1]}\")\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Parce que les ensembles de données sont des itérateurs, si nous voulons utiliser les données plusieurs fois, nous devons les convertir en liste :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 3,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"train_dataset, test_dataset = torchtext.datasets.AG_NEWS(root='./data')\n",
|
||||
"train_dataset = list(train_dataset)\n",
|
||||
"test_dataset = list(test_dataset)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Tokenisation\n",
|
||||
"\n",
|
||||
"Nous devons maintenant convertir le texte en **nombres** pouvant être représentés sous forme de tenseurs. Si nous souhaitons une représentation au niveau des mots, nous devons effectuer deux étapes :\n",
|
||||
"* utiliser un **tokeniseur** pour diviser le texte en **tokens**\n",
|
||||
"* construire un **vocabulaire** à partir de ces tokens.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"['he', 'said', 'hello']"
|
||||
]
|
||||
},
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"tokenizer = torchtext.data.utils.get_tokenizer('basic_english')\n",
|
||||
"tokenizer('He said: hello')"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 5,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"counter = collections.Counter()\n",
|
||||
"for (label, line) in train_dataset:\n",
|
||||
" counter.update(tokenizer(line))\n",
|
||||
"vocab = torchtext.vocab.vocab(counter, min_freq=1)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"En utilisant le vocabulaire, nous pouvons facilement encoder notre chaîne tokenisée en un ensemble de nombres :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 19,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Vocab size if 95810\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"[599, 3279, 97, 1220, 329, 225, 7368]"
|
||||
]
|
||||
},
|
||||
"execution_count": 19,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"vocab_size = len(vocab)\n",
|
||||
"print(f\"Vocab size if {vocab_size}\")\n",
|
||||
"\n",
|
||||
"stoi = vocab.get_stoi() # dict to convert tokens to indices\n",
|
||||
"\n",
|
||||
"def encode(x):\n",
|
||||
" return [stoi[s] for s in tokenizer(x)]\n",
|
||||
"\n",
|
||||
"encode('I love to play with my words')"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Représentation textuelle par sac de mots\n",
|
||||
"\n",
|
||||
"Parce que les mots véhiculent du sens, il est parfois possible de comprendre le sens d'un texte simplement en regardant les mots individuels, indépendamment de leur ordre dans la phrase. Par exemple, pour classifier des articles de presse, des mots comme *météo*, *neige* sont susceptibles d'indiquer une *prévision météorologique*, tandis que des mots comme *actions*, *dollar* pourraient correspondre à des *nouvelles financières*.\n",
|
||||
"\n",
|
||||
"La représentation vectorielle **Sac de mots** (BoW) est la méthode traditionnelle la plus couramment utilisée. Chaque mot est associé à un indice de vecteur, et l'élément du vecteur contient le nombre d'occurrences d'un mot dans un document donné.\n",
|
||||
"\n",
|
||||
" \n",
|
||||
"\n",
|
||||
"> **Note** : Vous pouvez également considérer BoW comme la somme de tous les vecteurs encodés en one-hot pour les mots individuels du texte.\n",
|
||||
"\n",
|
||||
"Voici un exemple de génération d'une représentation par sac de mots en utilisant la bibliothèque Python Scikit Learn :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 7,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"array([[1, 1, 0, 2, 0, 0, 0, 0, 0]], dtype=int64)"
|
||||
]
|
||||
},
|
||||
"execution_count": 7,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"from sklearn.feature_extraction.text import CountVectorizer\n",
|
||||
"vectorizer = CountVectorizer()\n",
|
||||
"corpus = [\n",
|
||||
" 'I like hot dogs.',\n",
|
||||
" 'The dog ran fast.',\n",
|
||||
" 'Its hot outside.',\n",
|
||||
" ]\n",
|
||||
"vectorizer.fit_transform(corpus)\n",
|
||||
"vectorizer.transform(['My dog likes hot dogs on a hot day.']).toarray()"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Pour calculer le vecteur sac-de-mots à partir de la représentation vectorielle de notre ensemble de données AG_NEWS, nous pouvons utiliser la fonction suivante :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 20,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"tensor([2., 1., 2., ..., 0., 0., 0.])\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"vocab_size = len(vocab)\n",
|
||||
"\n",
|
||||
"def to_bow(text,bow_vocab_size=vocab_size):\n",
|
||||
" res = torch.zeros(bow_vocab_size,dtype=torch.float32)\n",
|
||||
" for i in encode(text):\n",
|
||||
" if i<bow_vocab_size:\n",
|
||||
" res[i] += 1\n",
|
||||
" return res\n",
|
||||
"\n",
|
||||
"print(to_bow(train_dataset[0][1]))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"> **Remarque :** Ici, nous utilisons la variable globale `vocab_size` pour spécifier la taille par défaut du vocabulaire. Étant donné que la taille du vocabulaire est souvent assez grande, nous pouvons limiter la taille du vocabulaire aux mots les plus fréquents. Essayez de réduire la valeur de `vocab_size` et d'exécuter le code ci-dessous, et observez comment cela affecte la précision. Vous devriez vous attendre à une légère baisse de précision, mais pas dramatique, en échange d'une meilleure performance.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Entraîner un classificateur BoW\n",
|
||||
"\n",
|
||||
"Maintenant que nous avons appris à construire une représentation Bag-of-Words pour notre texte, entraînons un classificateur par-dessus. Tout d'abord, nous devons convertir notre ensemble de données pour l'entraînement de manière à ce que toutes les représentations vectorielles positionnelles soient transformées en représentation Bag-of-Words. Cela peut être réalisé en passant la fonction `bowify` comme paramètre `collate_fn` au `DataLoader` standard de torch :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 21,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"from torch.utils.data import DataLoader\n",
|
||||
"import numpy as np \n",
|
||||
"\n",
|
||||
"# this collate function gets list of batch_size tuples, and needs to \n",
|
||||
"# return a pair of label-feature tensors for the whole minibatch\n",
|
||||
"def bowify(b):\n",
|
||||
" return (\n",
|
||||
" torch.LongTensor([t[0]-1 for t in b]),\n",
|
||||
" torch.stack([to_bow(t[1]) for t in b])\n",
|
||||
" )\n",
|
||||
"\n",
|
||||
"train_loader = DataLoader(train_dataset, batch_size=16, collate_fn=bowify, shuffle=True)\n",
|
||||
"test_loader = DataLoader(test_dataset, batch_size=16, collate_fn=bowify, shuffle=True)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Définissons maintenant un réseau de neurones classificateur simple qui contient une couche linéaire. La taille du vecteur d'entrée est égale à `vocab_size`, et la taille de sortie correspond au nombre de classes (4). Étant donné que nous résolvons une tâche de classification, la fonction d'activation finale est `LogSoftmax()`.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 22,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"net = torch.nn.Sequential(torch.nn.Linear(vocab_size,4),torch.nn.LogSoftmax(dim=1))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Maintenant, nous allons définir une boucle d'entraînement standard avec PyTorch. Étant donné que notre ensemble de données est assez volumineux, pour notre objectif pédagogique, nous n'entraînerons que pendant une seule époque, et parfois même moins d'une époque (la spécification du paramètre `epoch_size` nous permet de limiter l'entraînement). Nous rapporterons également l'exactitude accumulée de l'entraînement pendant la formation ; la fréquence de rapport est spécifiée à l'aide du paramètre `report_freq`.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 24,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"def train_epoch(net,dataloader,lr=0.01,optimizer=None,loss_fn = torch.nn.NLLLoss(),epoch_size=None, report_freq=200):\n",
|
||||
" optimizer = optimizer or torch.optim.Adam(net.parameters(),lr=lr)\n",
|
||||
" net.train()\n",
|
||||
" total_loss,acc,count,i = 0,0,0,0\n",
|
||||
" for labels,features in dataloader:\n",
|
||||
" optimizer.zero_grad()\n",
|
||||
" out = net(features)\n",
|
||||
" loss = loss_fn(out,labels) #cross_entropy(out,labels)\n",
|
||||
" loss.backward()\n",
|
||||
" optimizer.step()\n",
|
||||
" total_loss+=loss\n",
|
||||
" _,predicted = torch.max(out,1)\n",
|
||||
" acc+=(predicted==labels).sum()\n",
|
||||
" count+=len(labels)\n",
|
||||
" i+=1\n",
|
||||
" if i%report_freq==0:\n",
|
||||
" print(f\"{count}: acc={acc.item()/count}\")\n",
|
||||
" if epoch_size and count>epoch_size:\n",
|
||||
" break\n",
|
||||
" return total_loss.item()/count, acc.item()/count"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 25,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"3200: acc=0.8028125\n",
|
||||
"6400: acc=0.8371875\n",
|
||||
"9600: acc=0.8534375\n",
|
||||
"12800: acc=0.85765625\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"(0.026090790722161722, 0.8620069296375267)"
|
||||
]
|
||||
},
|
||||
"execution_count": 25,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"train_epoch(net,train_loader,epoch_size=15000)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## BiGrams, TriGrams et N-Grams\n",
|
||||
"\n",
|
||||
"Une limitation de l'approche par sac de mots est que certains mots font partie d'expressions composées de plusieurs mots. Par exemple, le mot 'hot dog' a une signification complètement différente des mots 'hot' et 'dog' dans d'autres contextes. Si nous représentons toujours les mots 'hot' et 'dog' par les mêmes vecteurs, cela peut perturber notre modèle.\n",
|
||||
"\n",
|
||||
"Pour résoudre ce problème, les **représentations N-gram** sont souvent utilisées dans les méthodes de classification de documents, où la fréquence de chaque mot, bi-mot ou tri-mot constitue une caractéristique utile pour entraîner des classificateurs. Dans une représentation bigramme, par exemple, nous ajouterons toutes les paires de mots au vocabulaire, en plus des mots originaux.\n",
|
||||
"\n",
|
||||
"Voici un exemple de génération d'une représentation par sac de mots bigramme en utilisant Scikit Learn :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 26,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Vocabulary:\n",
|
||||
" {'i': 7, 'like': 11, 'hot': 4, 'dogs': 2, 'i like': 8, 'like hot': 12, 'hot dogs': 5, 'the': 16, 'dog': 0, 'ran': 14, 'fast': 3, 'the dog': 17, 'dog ran': 1, 'ran fast': 15, 'its': 9, 'outside': 13, 'its hot': 10, 'hot outside': 6}\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"array([[1, 0, 1, 0, 2, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]],\n",
|
||||
" dtype=int64)"
|
||||
]
|
||||
},
|
||||
"execution_count": 26,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"bigram_vectorizer = CountVectorizer(ngram_range=(1, 2), token_pattern=r'\\b\\w+\\b', min_df=1)\n",
|
||||
"corpus = [\n",
|
||||
" 'I like hot dogs.',\n",
|
||||
" 'The dog ran fast.',\n",
|
||||
" 'Its hot outside.',\n",
|
||||
" ]\n",
|
||||
"bigram_vectorizer.fit_transform(corpus)\n",
|
||||
"print(\"Vocabulary:\\n\",bigram_vectorizer.vocabulary_)\n",
|
||||
"bigram_vectorizer.transform(['My dog likes hot dogs on a hot day.']).toarray()\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Le principal inconvénient de l'approche N-gram est que la taille du vocabulaire commence à croître de manière extrêmement rapide. En pratique, il est nécessaire de combiner la représentation N-gram avec certaines techniques de réduction de dimensionnalité, comme les *embeddings*, que nous aborderons dans la prochaine unité.\n",
|
||||
"\n",
|
||||
"Pour utiliser la représentation N-gram dans notre jeu de données **AG News**, nous devons construire un vocabulaire spécifique aux n-grams :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 27,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Bigram vocabulary length = 1308842\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"counter = collections.Counter()\n",
|
||||
"for (label, line) in train_dataset:\n",
|
||||
" l = tokenizer(line)\n",
|
||||
" counter.update(torchtext.data.utils.ngrams_iterator(l,ngrams=2))\n",
|
||||
" \n",
|
||||
"bi_vocab = torchtext.vocab.vocab(counter, min_freq=1)\n",
|
||||
"\n",
|
||||
"print(\"Bigram vocabulary length = \",len(bi_vocab))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Nous pourrions alors utiliser le même code que ci-dessus pour entraîner le classificateur, cependant, cela serait très inefficace en termes de mémoire. Dans la prochaine unité, nous entraînerons un classificateur bigramme en utilisant des embeddings.\n",
|
||||
"\n",
|
||||
"> **Note:** Vous pouvez conserver uniquement les ngrams qui apparaissent dans le texte plus souvent qu'un nombre spécifié de fois. Cela garantira que les bigrammes peu fréquents seront omis et réduira considérablement la dimensionnalité. Pour ce faire, définissez le paramètre `min_freq` à une valeur plus élevée et observez le changement de longueur du vocabulaire.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Fréquence Terme-Fréquence Inverse de Document TF-IDF\n",
|
||||
"\n",
|
||||
"Dans la représentation BoW, les occurrences des mots sont pondérées de manière égale, quel que soit le mot lui-même. Cependant, il est évident que les mots fréquents, tels que *un*, *dans*, etc., sont beaucoup moins importants pour la classification que les termes spécialisés. En réalité, dans la plupart des tâches de NLP, certains mots sont plus pertinents que d'autres.\n",
|
||||
"\n",
|
||||
"**TF-IDF** signifie **fréquence terme–fréquence inverse de document**. C'est une variation du sac de mots, où au lieu d'une valeur binaire 0/1 indiquant la présence d'un mot dans un document, une valeur en virgule flottante est utilisée, qui est liée à la fréquence d'apparition du mot dans le corpus.\n",
|
||||
"\n",
|
||||
"Plus formellement, le poids $w_{ij}$ d'un mot $i$ dans le document $j$ est défini comme suit :\n",
|
||||
"$$\n",
|
||||
"w_{ij} = tf_{ij}\\times\\log({N\\over df_i})\n",
|
||||
"$$\n",
|
||||
"où\n",
|
||||
"* $tf_{ij}$ est le nombre d'occurrences de $i$ dans $j$, c'est-à-dire la valeur BoW que nous avons vue précédemment\n",
|
||||
"* $N$ est le nombre de documents dans la collection\n",
|
||||
"* $df_i$ est le nombre de documents contenant le mot $i$ dans l'ensemble de la collection\n",
|
||||
"\n",
|
||||
"La valeur TF-IDF $w_{ij}$ augmente proportionnellement au nombre de fois qu'un mot apparaît dans un document et est ajustée par le nombre de documents du corpus contenant ce mot, ce qui permet de corriger le fait que certains mots apparaissent plus fréquemment que d'autres. Par exemple, si le mot apparaît dans *chaque* document de la collection, $df_i=N$, et $w_{ij}=0$, ces termes seraient alors complètement ignorés.\n",
|
||||
"\n",
|
||||
"Vous pouvez facilement créer une vectorisation TF-IDF de texte en utilisant Scikit Learn :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 28,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"array([[0.43381609, 0. , 0.43381609, 0. , 0.65985664,\n",
|
||||
" 0.43381609, 0. , 0. , 0. , 0. ,\n",
|
||||
" 0. , 0. , 0. , 0. , 0. ,\n",
|
||||
" 0. ]])"
|
||||
]
|
||||
},
|
||||
"execution_count": 28,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"from sklearn.feature_extraction.text import TfidfVectorizer\n",
|
||||
"vectorizer = TfidfVectorizer(ngram_range=(1,2))\n",
|
||||
"vectorizer.fit_transform(corpus)\n",
|
||||
"vectorizer.transform(['My dog likes hot dogs on a hot day.']).toarray()"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Conclusion \n",
|
||||
"\n",
|
||||
"Cependant, bien que les représentations TF-IDF attribuent un poids de fréquence à différents mots, elles ne parviennent pas à représenter le sens ou l'ordre. Comme l'a dit le célèbre linguiste J. R. Firth en 1935 : « Le sens complet d'un mot est toujours contextuel, et aucune étude du sens en dehors du contexte ne peut être prise au sérieux. ». Nous apprendrons plus tard dans le cours comment capturer les informations contextuelles à partir du texte en utilisant la modélisation du langage.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"\n---\n\n**Avertissement** : \nCe document a été traduit à l'aide du service de traduction automatique [Co-op Translator](https://github.com/Azure/co-op-translator). Bien que nous nous efforcions d'assurer l'exactitude, veuillez noter que les traductions automatisées peuvent contenir des erreurs ou des inexactitudes. Le document original dans sa langue d'origine doit être considéré comme la source faisant autorité. Pour des informations critiques, il est recommandé de recourir à une traduction professionnelle réalisée par un humain. Nous déclinons toute responsabilité en cas de malentendus ou d'interprétations erronées résultant de l'utilisation de cette traduction.\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
"interpreter": {
|
||||
"hash": "16af2a8bbb083ea23e5e41c7f5787656b2ce26968575d8763f2c4b17f9cd711f"
|
||||
},
|
||||
"kernelspec": {
|
||||
"display_name": "Python 3.8.12 ('py38')",
|
||||
"language": "python",
|
||||
"name": "python3"
|
||||
},
|
||||
"language_info": {
|
||||
"codemirror_mode": {
|
||||
"name": "ipython",
|
||||
"version": 3
|
||||
},
|
||||
"file_extension": ".py",
|
||||
"mimetype": "text/x-python",
|
||||
"name": "python",
|
||||
"nbconvert_exporter": "python",
|
||||
"pygments_lexer": "ipython3",
|
||||
"version": "3.8.12"
|
||||
},
|
||||
"coopTranslator": {
|
||||
"original_hash": "7b9040985e748e4e2d4c689892456ad7",
|
||||
"translation_date": "2025-08-31T15:29:10+00:00",
|
||||
"source_file": "lessons/5-NLP/13-TextRep/TextRepresentationPyTorch.ipynb",
|
||||
"language_code": "fr"
|
||||
}
|
||||
},
|
||||
"nbformat": 4,
|
||||
"nbformat_minor": 2
|
||||
}
|
||||
|
|
@ -0,0 +1,647 @@
|
|||
{
|
||||
"cells": [
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"# Tâche de classification de texte\n",
|
||||
"\n",
|
||||
"Dans ce module, nous allons commencer par une tâche simple de classification de texte basée sur le jeu de données **[AG_NEWS](http://www.di.unipi.it/~gulli/AG_corpus_of_news_articles.html)** : nous allons classer des titres d'actualités en l'une des 4 catégories suivantes : Monde, Sports, Économie et Sci/Tech.\n",
|
||||
"\n",
|
||||
"## Le jeu de données\n",
|
||||
"\n",
|
||||
"Pour charger le jeu de données, nous utiliserons l'API **[TensorFlow Datasets](https://www.tensorflow.org/datasets)**.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 1,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"import tensorflow as tf\n",
|
||||
"from tensorflow import keras\n",
|
||||
"import tensorflow_datasets as tfds\n",
|
||||
"\n",
|
||||
"# In this tutorial, we will be training a lot of models. In order to use GPU memory cautiously,\n",
|
||||
"# we will set tensorflow option to grow GPU memory allocation when required.\n",
|
||||
"physical_devices = tf.config.list_physical_devices('GPU') \n",
|
||||
"if len(physical_devices)>0:\n",
|
||||
" tf.config.experimental.set_memory_growth(physical_devices[0], True)\n",
|
||||
"\n",
|
||||
"dataset = tfds.load('ag_news_subset')"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Nous pouvons maintenant accéder aux parties d'entraînement et de test du jeu de données en utilisant `dataset['train']` et `dataset['test']` respectivement :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 3,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Length of train dataset = 120000\n",
|
||||
"Length of test dataset = 7600\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"ds_train = dataset['train']\n",
|
||||
"ds_test = dataset['test']\n",
|
||||
"\n",
|
||||
"print(f\"Length of train dataset = {len(ds_train)}\")\n",
|
||||
"print(f\"Length of test dataset = {len(ds_test)}\")"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Imprimons les 10 premiers nouveaux titres de notre ensemble de données :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"3 (Sci/Tech) -> b'AMD Debuts Dual-Core Opteron Processor' b'AMD #39;s new dual-core Opteron chip is designed mainly for corporate computing applications, including databases, Web services, and financial transactions.'\n",
|
||||
"1 (Sports) -> b\"Wood's Suspension Upheld (Reuters)\" b'Reuters - Major League Baseball\\\\Monday announced a decision on the appeal filed by Chicago Cubs\\\\pitcher Kerry Wood regarding a suspension stemming from an\\\\incident earlier this season.'\n",
|
||||
"2 (Business) -> b'Bush reform may have blue states seeing red' b'President Bush #39;s quot;revenue-neutral quot; tax reform needs losers to balance its winners, and people claiming the federal deduction for state and local taxes may be in administration planners #39; sights, news reports say.'\n",
|
||||
"3 (Sci/Tech) -> b\"'Halt science decline in schools'\" b'Britain will run out of leading scientists unless science education is improved, says Professor Colin Pillinger.'\n",
|
||||
"1 (Sports) -> b'Gerrard leaves practice' b'London, England (Sports Network) - England midfielder Steven Gerrard injured his groin late in Thursday #39;s training session, but is hopeful he will be ready for Saturday #39;s World Cup qualifier against Austria.'\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"classes = ['World', 'Sports', 'Business', 'Sci/Tech']\n",
|
||||
"\n",
|
||||
"for i,x in zip(range(5),ds_train):\n",
|
||||
" print(f\"{x['label']} ({classes[x['label']]}) -> {x['title']} {x['description']}\")"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Vectorisation du texte\n",
|
||||
"\n",
|
||||
"Nous devons maintenant convertir le texte en **nombres** pouvant être représentés sous forme de tenseurs. Si nous souhaitons une représentation au niveau des mots, nous devons effectuer deux étapes :\n",
|
||||
"\n",
|
||||
"* Utiliser un **tokeniseur** pour diviser le texte en **tokens**.\n",
|
||||
"* Construire un **vocabulaire** à partir de ces tokens.\n",
|
||||
"\n",
|
||||
"### Limitation de la taille du vocabulaire\n",
|
||||
"\n",
|
||||
"Dans l'exemple du jeu de données AG News, la taille du vocabulaire est assez grande, avec plus de 100 000 mots. De manière générale, nous n'avons pas besoin des mots qui apparaissent rarement dans le texte — seuls quelques phrases les contiendront, et le modèle ne pourra pas en tirer d'apprentissage. Par conséquent, il est logique de limiter la taille du vocabulaire à un nombre plus restreint en passant un argument au constructeur du vectoriseur :\n",
|
||||
"\n",
|
||||
"Ces deux étapes peuvent être gérées à l'aide de la couche **TextVectorization**. Instancions l'objet vectoriseur, puis appelons la méthode `adapt` pour parcourir tout le texte et construire un vocabulaire :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 5,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"vocab_size = 50000\n",
|
||||
"vectorizer = keras.layers.experimental.preprocessing.TextVectorization(max_tokens=vocab_size)\n",
|
||||
"vectorizer.adapt(ds_train.take(500).map(lambda x: x['title']+' '+x['description']))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"> **Note** que nous utilisons uniquement un sous-ensemble de l'ensemble de données complet pour construire un vocabulaire. Nous faisons cela pour accélérer le temps d'exécution et éviter de vous faire attendre. Cependant, nous prenons le risque que certains mots de l'ensemble de données complet ne soient pas inclus dans le vocabulaire et soient ignorés pendant l'entraînement. Ainsi, utiliser la taille complète du vocabulaire et parcourir l'ensemble des données pendant `adapt` devrait augmenter la précision finale, mais pas de manière significative.\n",
|
||||
"\n",
|
||||
"Nous pouvons maintenant accéder au vocabulaire réel :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 6,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"['', '[UNK]', 'the', 'to', 'a', 'in', 'of', 'and', 'on', 'for']\n",
|
||||
"Length of vocabulary: 5335\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"vocab = vectorizer.get_vocabulary()\n",
|
||||
"vocab_size = len(vocab)\n",
|
||||
"print(vocab[:10])\n",
|
||||
"print(f\"Length of vocabulary: {vocab_size}\")"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"En utilisant le vectoriseur, nous pouvons facilement encoder n'importe quel texte en un ensemble de chiffres :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 7,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"<tf.Tensor: shape=(7,), dtype=int64, numpy=array([ 112, 3695, 3, 304, 11, 1041, 1], dtype=int64)>"
|
||||
]
|
||||
},
|
||||
"execution_count": 7,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"vectorizer('I love to play with my words')"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Représentation textuelle par sac de mots\n",
|
||||
"\n",
|
||||
"Parce que les mots véhiculent du sens, il est parfois possible de comprendre le sens d'un texte simplement en regardant les mots individuels, indépendamment de leur ordre dans la phrase. Par exemple, pour classifier des articles de presse, des mots comme *météo* et *neige* sont susceptibles d'indiquer une *prévision météorologique*, tandis que des mots comme *actions* et *dollar* seraient associés à des *nouvelles financières*.\n",
|
||||
"\n",
|
||||
"La représentation vectorielle par **sac de mots** (BoW) est la méthode traditionnelle la plus simple à comprendre. Chaque mot est associé à un indice de vecteur, et un élément du vecteur contient le nombre d'occurrences de chaque mot dans un document donné.\n",
|
||||
"\n",
|
||||
" \n",
|
||||
"\n",
|
||||
"> **Note** : Vous pouvez également considérer le BoW comme la somme de tous les vecteurs encodés en one-hot pour les mots individuels du texte.\n",
|
||||
"\n",
|
||||
"Voici un exemple de génération d'une représentation par sac de mots en utilisant la bibliothèque Python Scikit Learn :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 8,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"array([[1, 1, 0, 2, 0, 0, 0, 0, 0]], dtype=int64)"
|
||||
]
|
||||
},
|
||||
"execution_count": 8,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"from sklearn.feature_extraction.text import CountVectorizer\n",
|
||||
"sc_vectorizer = CountVectorizer()\n",
|
||||
"corpus = [\n",
|
||||
" 'I like hot dogs.',\n",
|
||||
" 'The dog ran fast.',\n",
|
||||
" 'Its hot outside.',\n",
|
||||
" ]\n",
|
||||
"sc_vectorizer.fit_transform(corpus)\n",
|
||||
"sc_vectorizer.transform(['My dog likes hot dogs on a hot day.']).toarray()"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Nous pouvons également utiliser le vectoriseur Keras que nous avons défini ci-dessus, en convertissant chaque numéro de mot en un encodage one-hot et en additionnant tous ces vecteurs.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 9,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"array([0., 5., 0., ..., 0., 0., 0.], dtype=float32)"
|
||||
]
|
||||
},
|
||||
"execution_count": 9,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"def to_bow(text):\n",
|
||||
" return tf.reduce_sum(tf.one_hot(vectorizer(text),vocab_size),axis=0)\n",
|
||||
"\n",
|
||||
"to_bow('My dog likes hot dogs on a hot day.').numpy()"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"> **Remarque** : Vous pourriez être surpris que le résultat diffère de l'exemple précédent. La raison en est que, dans l'exemple avec Keras, la longueur du vecteur correspond à la taille du vocabulaire, qui a été construit à partir de l'ensemble complet du jeu de données AG News, tandis que dans l'exemple avec Scikit Learn, nous avons construit le vocabulaire à partir du texte d'exemple à la volée.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Entraîner le classificateur BoW\n",
|
||||
"\n",
|
||||
"Maintenant que nous avons appris à construire la représentation sac de mots (bag-of-words) de notre texte, entraînons un classificateur qui l'utilise. Tout d'abord, nous devons convertir notre jeu de données en une représentation sac de mots. Cela peut être réalisé en utilisant la fonction `map` de la manière suivante :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 11,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"batch_size = 128\n",
|
||||
"\n",
|
||||
"ds_train_bow = ds_train.map(lambda x: (to_bow(x['title']+x['description']),x['label'])).batch(batch_size)\n",
|
||||
"ds_test_bow = ds_test.map(lambda x: (to_bow(x['title']+x['description']),x['label'])).batch(batch_size)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Définissons maintenant un réseau de neurones classificateur simple qui contient une couche linéaire. La taille de l'entrée est `vocab_size`, et la taille de la sortie correspond au nombre de classes (4). Étant donné que nous résolvons une tâche de classification, la fonction d'activation finale est **softmax** :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 12,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"938/938 [==============================] - 66s 70ms/step - loss: 0.6144 - acc: 0.8427 - val_loss: 0.4416 - val_acc: 0.8697\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"<keras.callbacks.History at 0x20c70a947f0>"
|
||||
]
|
||||
},
|
||||
"execution_count": 12,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"model = keras.models.Sequential([\n",
|
||||
" keras.layers.Dense(4,activation='softmax',input_shape=(vocab_size,))\n",
|
||||
"])\n",
|
||||
"model.compile(loss='sparse_categorical_crossentropy',optimizer='adam',metrics=['acc'])\n",
|
||||
"model.fit(ds_train_bow,validation_data=ds_test_bow)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Puisque nous avons 4 classes, une précision supérieure à 80 % est un bon résultat.\n",
|
||||
"\n",
|
||||
"## Entraîner un classificateur comme un réseau unique\n",
|
||||
"\n",
|
||||
"Étant donné que le vectoriseur est également une couche Keras, nous pouvons définir un réseau qui l'inclut et l'entraîner de bout en bout. De cette manière, nous n'avons pas besoin de vectoriser le jeu de données en utilisant `map`, nous pouvons simplement passer le jeu de données original à l'entrée du réseau.\n",
|
||||
"\n",
|
||||
"> **Note** : Nous devrons tout de même appliquer des maps à notre jeu de données pour convertir les champs des dictionnaires (comme `title`, `description` et `label`) en tuples. Cependant, lors du chargement des données depuis le disque, nous pouvons construire un jeu de données avec la structure requise dès le départ.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 13,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Model: \"model\"\n",
|
||||
"_________________________________________________________________\n",
|
||||
" Layer (type) Output Shape Param # \n",
|
||||
"=================================================================\n",
|
||||
" input_1 (InputLayer) [(None, 1)] 0 \n",
|
||||
" \n",
|
||||
" text_vectorization (TextVec (None, None) 0 \n",
|
||||
" torization) \n",
|
||||
" \n",
|
||||
" tf.one_hot (TFOpLambda) (None, None, 5335) 0 \n",
|
||||
" \n",
|
||||
" tf.math.reduce_sum (TFOpLam (None, 5335) 0 \n",
|
||||
" bda) \n",
|
||||
" \n",
|
||||
" dense_2 (Dense) (None, 4) 21344 \n",
|
||||
" \n",
|
||||
"=================================================================\n",
|
||||
"Total params: 21,344\n",
|
||||
"Trainable params: 21,344\n",
|
||||
"Non-trainable params: 0\n",
|
||||
"_________________________________________________________________\n",
|
||||
"938/938 [==============================] - 73s 77ms/step - loss: 0.6057 - acc: 0.8414 - val_loss: 0.4202 - val_acc: 0.8736\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"<keras.callbacks.History at 0x20c721521f0>"
|
||||
]
|
||||
},
|
||||
"execution_count": 13,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"def extract_text(x):\n",
|
||||
" return x['title']+' '+x['description']\n",
|
||||
"\n",
|
||||
"def tupelize(x):\n",
|
||||
" return (extract_text(x),x['label'])\n",
|
||||
"\n",
|
||||
"inp = keras.Input(shape=(1,),dtype=tf.string)\n",
|
||||
"x = vectorizer(inp)\n",
|
||||
"x = tf.reduce_sum(tf.one_hot(x,vocab_size),axis=1)\n",
|
||||
"out = keras.layers.Dense(4,activation='softmax')(x)\n",
|
||||
"model = keras.models.Model(inp,out)\n",
|
||||
"model.summary()\n",
|
||||
"\n",
|
||||
"model.compile(loss='sparse_categorical_crossentropy',optimizer='adam',metrics=['acc'])\n",
|
||||
"model.fit(ds_train.map(tupelize).batch(batch_size),validation_data=ds_test.map(tupelize).batch(batch_size))\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Bigrams, trigrams et n-grams\n",
|
||||
"\n",
|
||||
"Une des limites de l'approche bag-of-words est que certains mots font partie d'expressions composées de plusieurs mots. Par exemple, le terme « hot dog » a une signification complètement différente des mots « hot » et « dog » pris séparément dans d'autres contextes. Si nous représentons toujours les mots « hot » et « dog » avec les mêmes vecteurs, cela peut induire notre modèle en erreur.\n",
|
||||
"\n",
|
||||
"Pour résoudre ce problème, les **représentations n-gram** sont souvent utilisées dans les méthodes de classification de documents, où la fréquence de chaque mot, bi-mot ou tri-mot constitue une caractéristique utile pour entraîner des classificateurs. Dans les représentations bigram, par exemple, nous ajoutons toutes les paires de mots au vocabulaire, en plus des mots originaux.\n",
|
||||
"\n",
|
||||
"Voici un exemple de génération d'une représentation bag-of-words bigram en utilisant Scikit Learn :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 14,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Vocabulary:\n",
|
||||
" {'i': 7, 'like': 11, 'hot': 4, 'dogs': 2, 'i like': 8, 'like hot': 12, 'hot dogs': 5, 'the': 16, 'dog': 0, 'ran': 14, 'fast': 3, 'the dog': 17, 'dog ran': 1, 'ran fast': 15, 'its': 9, 'outside': 13, 'its hot': 10, 'hot outside': 6}\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"array([[1, 0, 1, 0, 2, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]],\n",
|
||||
" dtype=int64)"
|
||||
]
|
||||
},
|
||||
"execution_count": 14,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"bigram_vectorizer = CountVectorizer(ngram_range=(1, 2), token_pattern=r'\\b\\w+\\b', min_df=1)\n",
|
||||
"corpus = [\n",
|
||||
" 'I like hot dogs.',\n",
|
||||
" 'The dog ran fast.',\n",
|
||||
" 'Its hot outside.',\n",
|
||||
" ]\n",
|
||||
"bigram_vectorizer.fit_transform(corpus)\n",
|
||||
"print(\"Vocabulary:\\n\",bigram_vectorizer.vocabulary_)\n",
|
||||
"bigram_vectorizer.transform(['My dog likes hot dogs on a hot day.']).toarray()\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Le principal inconvénient de l'approche n-gram est que la taille du vocabulaire commence à croître extrêmement rapidement. En pratique, nous devons combiner la représentation n-gram avec une technique de réduction de dimensionnalité, comme les *embeddings*, que nous aborderons dans la prochaine unité.\n",
|
||||
"\n",
|
||||
"Pour utiliser une représentation n-gram dans notre ensemble de données **AG News**, nous devons passer le paramètre `ngrams` au constructeur de `TextVectorization`. La taille d'un vocabulaire de bigrammes est **significativement plus grande**, dans notre cas, elle dépasse 1,3 million de tokens ! Il est donc logique de limiter également les tokens de bigrammes à un nombre raisonnable.\n",
|
||||
"\n",
|
||||
"Nous pourrions utiliser le même code que précédemment pour entraîner le classificateur, mais cela serait très inefficace en termes de mémoire. Dans la prochaine unité, nous entraînerons le classificateur de bigrammes en utilisant des embeddings. En attendant, vous pouvez expérimenter avec l'entraînement du classificateur de bigrammes dans ce notebook et voir si vous pouvez obtenir une meilleure précision.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Calcul des vecteurs BoW automatiquement\n",
|
||||
"\n",
|
||||
"Dans l'exemple ci-dessus, nous avons calculé les vecteurs BoW manuellement en additionnant les encodages one-hot des mots individuels. Cependant, la dernière version de TensorFlow nous permet de calculer les vecteurs BoW automatiquement en passant le paramètre `output_mode='count'` au constructeur du vectoriseur. Cela rend la définition et l'entraînement de notre modèle beaucoup plus simples :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 15,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Training vectorizer\n",
|
||||
"938/938 [==============================] - 7s 7ms/step - loss: 0.5929 - acc: 0.8486 - val_loss: 0.4168 - val_acc: 0.8772\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"<keras.callbacks.History at 0x20c725217c0>"
|
||||
]
|
||||
},
|
||||
"execution_count": 15,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"model = keras.models.Sequential([\n",
|
||||
" keras.layers.experimental.preprocessing.TextVectorization(max_tokens=vocab_size,output_mode='count'),\n",
|
||||
" keras.layers.Dense(4,input_shape=(vocab_size,), activation='softmax')\n",
|
||||
"])\n",
|
||||
"print(\"Training vectorizer\")\n",
|
||||
"model.layers[0].adapt(ds_train.take(500).map(extract_text))\n",
|
||||
"model.compile(loss='sparse_categorical_crossentropy',optimizer='adam',metrics=['acc'])\n",
|
||||
"model.fit(ds_train.map(tupelize).batch(batch_size),validation_data=ds_test.map(tupelize).batch(batch_size))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Fréquence de terme - fréquence inverse de document (TF-IDF)\n",
|
||||
"\n",
|
||||
"Dans la représentation BoW, les occurrences des mots sont pondérées en utilisant la même technique, quel que soit le mot lui-même. Cependant, il est évident que les mots fréquents comme *un* et *dans* sont beaucoup moins importants pour la classification que les termes spécialisés. Dans la plupart des tâches de NLP, certains mots sont plus pertinents que d'autres.\n",
|
||||
"\n",
|
||||
"**TF-IDF** signifie **fréquence de terme - fréquence inverse de document**. C'est une variation du sac de mots, où au lieu d'une valeur binaire 0/1 indiquant la présence d'un mot dans un document, une valeur en virgule flottante est utilisée, qui est liée à la fréquence d'apparition du mot dans le corpus.\n",
|
||||
"\n",
|
||||
"Plus formellement, le poids $w_{ij}$ d'un mot $i$ dans le document $j$ est défini comme suit :\n",
|
||||
"$$\n",
|
||||
"w_{ij} = tf_{ij}\\times\\log({N\\over df_i})\n",
|
||||
"$$\n",
|
||||
"où\n",
|
||||
"* $tf_{ij}$ est le nombre d'occurrences de $i$ dans $j$, c'est-à-dire la valeur BoW que nous avons vue précédemment\n",
|
||||
"* $N$ est le nombre de documents dans la collection\n",
|
||||
"* $df_i$ est le nombre de documents contenant le mot $i$ dans l'ensemble de la collection\n",
|
||||
"\n",
|
||||
"La valeur TF-IDF $w_{ij}$ augmente proportionnellement au nombre de fois qu'un mot apparaît dans un document et est ajustée par le nombre de documents dans le corpus contenant ce mot, ce qui permet de compenser le fait que certains mots apparaissent plus fréquemment que d'autres. Par exemple, si le mot apparaît dans *chaque* document de la collection, $df_i=N$, et $w_{ij}=0$, ces termes seraient complètement ignorés.\n",
|
||||
"\n",
|
||||
"Vous pouvez facilement créer une vectorisation TF-IDF de texte en utilisant Scikit Learn :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 16,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"array([[0.43381609, 0. , 0.43381609, 0. , 0.65985664,\n",
|
||||
" 0.43381609, 0. , 0. , 0. , 0. ,\n",
|
||||
" 0. , 0. , 0. , 0. , 0. ,\n",
|
||||
" 0. ]])"
|
||||
]
|
||||
},
|
||||
"execution_count": 16,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"from sklearn.feature_extraction.text import TfidfVectorizer\n",
|
||||
"vectorizer = TfidfVectorizer(ngram_range=(1,2))\n",
|
||||
"vectorizer.fit_transform(corpus)\n",
|
||||
"vectorizer.transform(['My dog likes hot dogs on a hot day.']).toarray()"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Dans Keras, la couche `TextVectorization` peut calculer automatiquement les fréquences TF-IDF en passant le paramètre `output_mode='tf-idf'`. Répétons le code que nous avons utilisé ci-dessus pour voir si l'utilisation de TF-IDF augmente la précision :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 17,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Training vectorizer\n",
|
||||
"938/938 [==============================] - 12s 12ms/step - loss: 0.4197 - acc: 0.8662 - val_loss: 0.3432 - val_acc: 0.8849\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"<keras.callbacks.History at 0x20c729dfd30>"
|
||||
]
|
||||
},
|
||||
"execution_count": 17,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"model = keras.models.Sequential([\n",
|
||||
" keras.layers.experimental.preprocessing.TextVectorization(max_tokens=vocab_size,output_mode='tf-idf'),\n",
|
||||
" keras.layers.Dense(4,input_shape=(vocab_size,), activation='softmax')\n",
|
||||
"])\n",
|
||||
"print(\"Training vectorizer\")\n",
|
||||
"model.layers[0].adapt(ds_train.take(500).map(extract_text))\n",
|
||||
"model.compile(loss='sparse_categorical_crossentropy',optimizer='adam',metrics=['acc'])\n",
|
||||
"model.fit(ds_train.map(tupelize).batch(batch_size),validation_data=ds_test.map(tupelize).batch(batch_size))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Conclusion \n",
|
||||
"\n",
|
||||
"Bien que les représentations TF-IDF attribuent des poids de fréquence à différents mots, elles ne parviennent pas à représenter le sens ou l'ordre. Comme l'a dit le célèbre linguiste J. R. Firth en 1935 : \"Le sens complet d'un mot est toujours contextuel, et aucune étude du sens en dehors du contexte ne peut être prise au sérieux.\" Nous apprendrons plus tard dans le cours comment capturer les informations contextuelles à partir du texte en utilisant la modélisation du langage.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"\n---\n\n**Avertissement** : \nCe document a été traduit à l'aide du service de traduction automatique [Co-op Translator](https://github.com/Azure/co-op-translator). Bien que nous nous efforcions d'assurer l'exactitude, veuillez noter que les traductions automatisées peuvent contenir des erreurs ou des inexactitudes. Le document original dans sa langue d'origine doit être considéré comme la source faisant autorité. Pour des informations critiques, il est recommandé de recourir à une traduction professionnelle réalisée par un humain. Nous déclinons toute responsabilité en cas de malentendus ou d'interprétations erronées résultant de l'utilisation de cette traduction.\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
"interpreter": {
|
||||
"hash": "0cb620c6d4b9f7a635928804c26cf22403d89d98d79684e4529119355ee6d5a5"
|
||||
},
|
||||
"kernel_info": {
|
||||
"name": "conda-env-py37_tensorflow-py"
|
||||
},
|
||||
"kernelspec": {
|
||||
"display_name": "py37_tensorflow",
|
||||
"language": "python",
|
||||
"name": "python3"
|
||||
},
|
||||
"language_info": {
|
||||
"codemirror_mode": {
|
||||
"name": "ipython",
|
||||
"version": 3
|
||||
},
|
||||
"file_extension": ".py",
|
||||
"mimetype": "text/x-python",
|
||||
"name": "python",
|
||||
"nbconvert_exporter": "python",
|
||||
"pygments_lexer": "ipython3",
|
||||
"version": "3.8.12"
|
||||
},
|
||||
"nteract": {
|
||||
"version": "nteract-front-end@1.0.0"
|
||||
},
|
||||
"coopTranslator": {
|
||||
"original_hash": "19b43951d55b377a76209c24c1f017e4",
|
||||
"translation_date": "2025-08-31T15:30:49+00:00",
|
||||
"source_file": "lessons/5-NLP/13-TextRep/TextRepresentationTF.ipynb",
|
||||
"language_code": "fr"
|
||||
}
|
||||
},
|
||||
"nbformat": 4,
|
||||
"nbformat_minor": 4
|
||||
}
|
||||
|
|
@ -0,0 +1,724 @@
|
|||
{
|
||||
"cells": [
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Intégrations\n",
|
||||
"\n",
|
||||
"Dans notre exemple précédent, nous avons travaillé avec des vecteurs bag-of-words de haute dimension de longueur `vocab_size`, et nous convertissions explicitement des vecteurs de représentation positionnelle de basse dimension en représentation clairsemée one-hot. Cette représentation one-hot n'est pas efficace en termes de mémoire, de plus, chaque mot est traité indépendamment des autres, c'est-à-dire que les vecteurs encodés en one-hot n'expriment aucune similarité sémantique entre les mots.\n",
|
||||
"\n",
|
||||
"Dans cette unité, nous continuerons à explorer le jeu de données **News AG**. Pour commencer, chargeons les données et récupérons quelques définitions du notebook précédent.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 1,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Loading dataset...\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"name": "stderr",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"d:\\WORK\\ai-for-beginners\\5-NLP\\14-Embeddings\\data\\train.csv: 29.5MB [00:01, 18.8MB/s] \n",
|
||||
"d:\\WORK\\ai-for-beginners\\5-NLP\\14-Embeddings\\data\\test.csv: 1.86MB [00:00, 11.2MB/s] \n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Building vocab...\n",
|
||||
"Vocab size = 95812\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"import torch\n",
|
||||
"import torchtext\n",
|
||||
"import numpy as np\n",
|
||||
"from torchnlp import *\n",
|
||||
"train_dataset, test_dataset, classes, vocab = load_dataset()\n",
|
||||
"vocab_size = len(vocab)\n",
|
||||
"print(\"Vocab size = \",vocab_size)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Qu'est-ce qu'un embedding ?\n",
|
||||
"\n",
|
||||
"L'idée de l'**embedding** est de représenter les mots par des vecteurs denses de dimension inférieure, qui reflètent d'une certaine manière le sens sémantique d'un mot. Nous discuterons plus tard de la manière de construire des embeddings de mots significatifs, mais pour l'instant, considérons simplement les embeddings comme un moyen de réduire la dimensionnalité d'un vecteur de mots.\n",
|
||||
"\n",
|
||||
"Ainsi, une couche d'embedding prendrait un mot en entrée et produirait un vecteur de sortie de taille `embedding_size` spécifiée. En un sens, cela ressemble beaucoup à une couche `Linear`, mais au lieu de prendre un vecteur encodé en one-hot, elle pourra prendre un numéro de mot en entrée.\n",
|
||||
"\n",
|
||||
"En utilisant une couche d'embedding comme première couche de notre réseau, nous pouvons passer du modèle bag-of-words au modèle **embedding bag**, où nous convertissons d'abord chaque mot de notre texte en son embedding correspondant, puis nous calculons une fonction d'agrégation sur tous ces embeddings, comme `sum`, `average` ou `max`.\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"Notre réseau de neurones classificateur commencera par une couche d'embedding, suivie d'une couche d'agrégation, puis d'un classificateur linéaire au-dessus :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 2,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"class EmbedClassifier(torch.nn.Module):\n",
|
||||
" def __init__(self, vocab_size, embed_dim, num_class):\n",
|
||||
" super().__init__()\n",
|
||||
" self.embedding = torch.nn.Embedding(vocab_size, embed_dim)\n",
|
||||
" self.fc = torch.nn.Linear(embed_dim, num_class)\n",
|
||||
"\n",
|
||||
" def forward(self, x):\n",
|
||||
" x = self.embedding(x)\n",
|
||||
" x = torch.mean(x,dim=1)\n",
|
||||
" return self.fc(x)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"### Gérer la taille variable des séquences\n",
|
||||
"\n",
|
||||
"En raison de cette architecture, les minibatches pour notre réseau devront être créés d'une certaine manière. Dans l'unité précédente, en utilisant le sac de mots (BoW), tous les tenseurs BoW dans un minibatch avaient une taille égale à `vocab_size`, indépendamment de la longueur réelle de notre séquence de texte. Une fois que nous passons aux embeddings de mots, nous nous retrouvons avec un nombre variable de mots dans chaque échantillon de texte, et lors de la combinaison de ces échantillons en minibatches, nous devrons appliquer un certain remplissage.\n",
|
||||
"\n",
|
||||
"Cela peut être fait en utilisant la même technique qui consiste à fournir une fonction `collate_fn` à la source de données :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 3,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"def padify(b):\n",
|
||||
" # b is the list of tuples of length batch_size\n",
|
||||
" # - first element of a tuple = label, \n",
|
||||
" # - second = feature (text sequence)\n",
|
||||
" # build vectorized sequence\n",
|
||||
" v = [encode(x[1]) for x in b]\n",
|
||||
" # first, compute max length of a sequence in this minibatch\n",
|
||||
" l = max(map(len,v))\n",
|
||||
" return ( # tuple of two tensors - labels and features\n",
|
||||
" torch.LongTensor([t[0]-1 for t in b]),\n",
|
||||
" torch.stack([torch.nn.functional.pad(torch.tensor(t),(0,l-len(t)),mode='constant',value=0) for t in v])\n",
|
||||
" )\n",
|
||||
"\n",
|
||||
"train_loader = torch.utils.data.DataLoader(train_dataset, batch_size=16, collate_fn=padify, shuffle=True)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"### Entraîner le classificateur d'embedding\n",
|
||||
"\n",
|
||||
"Maintenant que nous avons défini un dataloader approprié, nous pouvons entraîner le modèle en utilisant la fonction d'entraînement que nous avons définie dans l'unité précédente :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"3200: acc=0.6415625\n",
|
||||
"6400: acc=0.6865625\n",
|
||||
"9600: acc=0.7103125\n",
|
||||
"12800: acc=0.726953125\n",
|
||||
"16000: acc=0.739375\n",
|
||||
"19200: acc=0.75046875\n",
|
||||
"22400: acc=0.7572321428571429\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"(0.889799795315499, 0.7623160588611644)"
|
||||
]
|
||||
},
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"net = EmbedClassifier(vocab_size,32,len(classes)).to(device)\n",
|
||||
"train_epoch(net,train_loader, lr=1, epoch_size=25000)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"> **Note** : Nous n'entraînons ici que sur 25 000 enregistrements (moins d'une époque complète) pour gagner du temps, mais vous pouvez continuer l'entraînement, écrire une fonction pour entraîner sur plusieurs époques, et expérimenter avec le paramètre de taux d'apprentissage pour atteindre une meilleure précision. Vous devriez pouvoir atteindre une précision d'environ 90 %.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"### Couche EmbeddingBag et Représentation de Séquences de Longueur Variable\n",
|
||||
"\n",
|
||||
"Dans l'architecture précédente, nous devions compléter toutes les séquences pour qu'elles aient la même longueur afin de les intégrer dans un minibatch. Ce n'est pas la manière la plus efficace de représenter des séquences de longueur variable - une autre approche consiste à utiliser un vecteur **offset**, qui contient les décalages de toutes les séquences stockées dans un grand vecteur unique.\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"> **Note** : Sur l'image ci-dessus, nous montrons une séquence de caractères, mais dans notre exemple, nous travaillons avec des séquences de mots. Cependant, le principe général de représentation des séquences avec un vecteur de décalage reste le même.\n",
|
||||
"\n",
|
||||
"Pour travailler avec la représentation par décalage, nous utilisons la couche [`EmbeddingBag`](https://pytorch.org/docs/stable/generated/torch.nn.EmbeddingBag.html). Elle est similaire à `Embedding`, mais elle prend un vecteur de contenu et un vecteur de décalage en entrée, et inclut également une couche de moyennage, qui peut être `mean`, `sum` ou `max`.\n",
|
||||
"\n",
|
||||
"Voici un réseau modifié qui utilise `EmbeddingBag` :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 5,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"class EmbedClassifier(torch.nn.Module):\n",
|
||||
" def __init__(self, vocab_size, embed_dim, num_class):\n",
|
||||
" super().__init__()\n",
|
||||
" self.embedding = torch.nn.EmbeddingBag(vocab_size, embed_dim)\n",
|
||||
" self.fc = torch.nn.Linear(embed_dim, num_class)\n",
|
||||
"\n",
|
||||
" def forward(self, text, off):\n",
|
||||
" x = self.embedding(text, off)\n",
|
||||
" return self.fc(x)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Pour préparer le jeu de données pour l'entraînement, nous devons fournir une fonction de conversion qui préparera le vecteur de décalage :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 6,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"def offsetify(b):\n",
|
||||
" # first, compute data tensor from all sequences\n",
|
||||
" x = [torch.tensor(encode(t[1])) for t in b]\n",
|
||||
" # now, compute the offsets by accumulating the tensor of sequence lengths\n",
|
||||
" o = [0] + [len(t) for t in x]\n",
|
||||
" o = torch.tensor(o[:-1]).cumsum(dim=0)\n",
|
||||
" return ( \n",
|
||||
" torch.LongTensor([t[0]-1 for t in b]), # labels\n",
|
||||
" torch.cat(x), # text \n",
|
||||
" o\n",
|
||||
" )\n",
|
||||
"\n",
|
||||
"train_loader = torch.utils.data.DataLoader(train_dataset, batch_size=16, collate_fn=offsetify, shuffle=True)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Notez que, contrairement à tous les exemples précédents, notre réseau accepte désormais deux paramètres : le vecteur de données et le vecteur de décalage, qui sont de tailles différentes. De même, notre chargeur de données nous fournit également 3 valeurs au lieu de 2 : les vecteurs de texte et de décalage sont fournis comme caractéristiques. Par conséquent, nous devons légèrement ajuster notre fonction d'entraînement pour en tenir compte :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 7,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"3200: acc=0.6153125\n",
|
||||
"6400: acc=0.6615625\n",
|
||||
"9600: acc=0.6932291666666667\n",
|
||||
"12800: acc=0.715078125\n",
|
||||
"16000: acc=0.7270625\n",
|
||||
"19200: acc=0.7382291666666667\n",
|
||||
"22400: acc=0.7486160714285715\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"(22.771553103007037, 0.7551983365323096)"
|
||||
]
|
||||
},
|
||||
"execution_count": 7,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"net = EmbedClassifier(vocab_size,32,len(classes)).to(device)\n",
|
||||
"\n",
|
||||
"def train_epoch_emb(net,dataloader,lr=0.01,optimizer=None,loss_fn = torch.nn.CrossEntropyLoss(),epoch_size=None, report_freq=200):\n",
|
||||
" optimizer = optimizer or torch.optim.Adam(net.parameters(),lr=lr)\n",
|
||||
" loss_fn = loss_fn.to(device)\n",
|
||||
" net.train()\n",
|
||||
" total_loss,acc,count,i = 0,0,0,0\n",
|
||||
" for labels,text,off in dataloader:\n",
|
||||
" optimizer.zero_grad()\n",
|
||||
" labels,text,off = labels.to(device), text.to(device), off.to(device)\n",
|
||||
" out = net(text, off)\n",
|
||||
" loss = loss_fn(out,labels) #cross_entropy(out,labels)\n",
|
||||
" loss.backward()\n",
|
||||
" optimizer.step()\n",
|
||||
" total_loss+=loss\n",
|
||||
" _,predicted = torch.max(out,1)\n",
|
||||
" acc+=(predicted==labels).sum()\n",
|
||||
" count+=len(labels)\n",
|
||||
" i+=1\n",
|
||||
" if i%report_freq==0:\n",
|
||||
" print(f\"{count}: acc={acc.item()/count}\")\n",
|
||||
" if epoch_size and count>epoch_size:\n",
|
||||
" break\n",
|
||||
" return total_loss.item()/count, acc.item()/count\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"train_epoch_emb(net,train_loader, lr=4, epoch_size=25000)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Intégrations Sémantiques : Word2Vec\n",
|
||||
"\n",
|
||||
"Dans notre exemple précédent, la couche d'intégration du modèle a appris à mapper des mots à une représentation vectorielle, mais cette représentation n'avait pas beaucoup de signification sémantique. Ce serait intéressant d'apprendre une telle représentation vectorielle où des mots similaires ou des synonymes correspondraient à des vecteurs proches les uns des autres en termes de distance vectorielle (par exemple, distance euclidienne).\n",
|
||||
"\n",
|
||||
"Pour cela, nous devons pré-entraîner notre modèle d'intégration sur une grande collection de textes d'une manière spécifique. L'une des premières méthodes pour entraîner des intégrations sémantiques s'appelle [Word2Vec](https://en.wikipedia.org/wiki/Word2vec). Elle repose sur deux principales architectures utilisées pour produire une représentation distribuée des mots :\n",
|
||||
"\n",
|
||||
" - **Sac de mots continu** (CBoW) — dans cette architecture, nous entraînons le modèle à prédire un mot à partir du contexte environnant. Étant donné le ngram $(W_{-2},W_{-1},W_0,W_1,W_2)$, l'objectif du modèle est de prédire $W_0$ à partir de $(W_{-2},W_{-1},W_1,W_2)$.\n",
|
||||
" - **Skip-gram continu** est l'opposé du CBoW. Le modèle utilise une fenêtre de mots contextuels environnants pour prédire le mot actuel.\n",
|
||||
"\n",
|
||||
"CBoW est plus rapide, tandis que skip-gram est plus lent, mais il représente mieux les mots peu fréquents.\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"Pour expérimenter avec l'intégration Word2Vec pré-entraînée sur le jeu de données Google News, nous pouvons utiliser la bibliothèque **gensim**. Ci-dessous, nous trouvons les mots les plus similaires à 'neural'.\n",
|
||||
"\n",
|
||||
"> **Note :** Lorsque vous créez des vecteurs de mots pour la première fois, leur téléchargement peut prendre un certain temps !\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 8,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"import gensim.downloader as api\n",
|
||||
"w2v = api.load('word2vec-google-news-300')"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 9,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"neuronal -> 0.7804799675941467\n",
|
||||
"neurons -> 0.7326500415802002\n",
|
||||
"neural_circuits -> 0.7252851724624634\n",
|
||||
"neuron -> 0.7174385190010071\n",
|
||||
"cortical -> 0.6941086649894714\n",
|
||||
"brain_circuitry -> 0.6923246383666992\n",
|
||||
"synaptic -> 0.6699118614196777\n",
|
||||
"neural_circuitry -> 0.6638563275337219\n",
|
||||
"neurochemical -> 0.6555314064025879\n",
|
||||
"neuronal_activity -> 0.6531826257705688\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"for w,p in w2v.most_similar('neural'):\n",
|
||||
" print(f\"{w} -> {p}\")"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Nous pouvons également calculer des embeddings de vecteurs à partir du mot, à utiliser dans l'entraînement du modèle de classification (nous montrons uniquement les 20 premiers composants du vecteur pour plus de clarté) :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 10,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"array([ 0.01226807, 0.06225586, 0.10693359, 0.05810547, 0.23828125,\n",
|
||||
" 0.03686523, 0.05151367, -0.20703125, 0.01989746, 0.10058594,\n",
|
||||
" -0.03759766, -0.1015625 , -0.15820312, -0.08105469, -0.0390625 ,\n",
|
||||
" -0.05053711, 0.16015625, 0.2578125 , 0.10058594, -0.25976562],\n",
|
||||
" dtype=float32)"
|
||||
]
|
||||
},
|
||||
"execution_count": 10,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"w2v.word_vec('play')[:20]"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"La grande chose à propos des embeddings sémantiques est que vous pouvez manipuler l'encodage vectoriel pour changer la sémantique. Par exemple, nous pouvons demander de trouver un mot dont la représentation vectorielle serait aussi proche que possible des mots *roi* et *femme*, et aussi éloignée que possible du mot *homme* :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 10,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"('queen', 0.7118192911148071)"
|
||||
]
|
||||
},
|
||||
"execution_count": 10,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"w2v.most_similar(positive=['king','woman'],negative=['man'])[0]"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Les modèles CBoW et Skip-Grams sont des embeddings dits \"prédictifs\", car ils ne prennent en compte que les contextes locaux. Word2Vec ne tire pas parti du contexte global.\n",
|
||||
"\n",
|
||||
"**FastText** s'appuie sur Word2Vec en apprenant des représentations vectorielles pour chaque mot ainsi que pour les n-grammes de caractères présents dans chaque mot. Les valeurs de ces représentations sont ensuite moyennées en un seul vecteur à chaque étape d'entraînement. Bien que cela ajoute beaucoup de calculs supplémentaires lors de la pré-formation, cela permet aux embeddings de mots d'intégrer des informations sur les sous-mots.\n",
|
||||
"\n",
|
||||
"Une autre méthode, **GloVe**, exploite l'idée de matrice de cooccurrence et utilise des méthodes neuronales pour décomposer cette matrice en vecteurs de mots plus expressifs et non linéaires.\n",
|
||||
"\n",
|
||||
"Vous pouvez expérimenter avec cet exemple en changeant les embeddings pour FastText et GloVe, car gensim prend en charge plusieurs modèles d'embeddings de mots différents.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Utilisation des embeddings pré-entraînés dans PyTorch\n",
|
||||
"\n",
|
||||
"Nous pouvons modifier l'exemple ci-dessus pour pré-remplir la matrice de notre couche d'embedding avec des embeddings sémantiques, comme Word2Vec. Il faut tenir compte du fait que les vocabulaires des embeddings pré-entraînés et de notre corpus de texte ne correspondront probablement pas, donc nous initialiserons les poids des mots manquants avec des valeurs aléatoires :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 11,
|
||||
"metadata": {
|
||||
"tags": []
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Embedding size: 300\n",
|
||||
"Populating matrix, this will take some time...Done, found 41080 words, 54732 words missing\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"embed_size = len(w2v.get_vector('hello'))\n",
|
||||
"print(f'Embedding size: {embed_size}')\n",
|
||||
"\n",
|
||||
"net = EmbedClassifier(vocab_size,embed_size,len(classes))\n",
|
||||
"\n",
|
||||
"print('Populating matrix, this will take some time...',end='')\n",
|
||||
"found, not_found = 0,0\n",
|
||||
"for i,w in enumerate(vocab.get_itos()):\n",
|
||||
" try:\n",
|
||||
" net.embedding.weight[i].data = torch.tensor(w2v.get_vector(w))\n",
|
||||
" found+=1\n",
|
||||
" except:\n",
|
||||
" net.embedding.weight[i].data = torch.normal(0.0,1.0,(embed_size,))\n",
|
||||
" not_found+=1\n",
|
||||
"\n",
|
||||
"print(f\"Done, found {found} words, {not_found} words missing\")\n",
|
||||
"net = net.to(device)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Maintenant, entraînons notre modèle. Notez que le temps nécessaire pour entraîner le modèle est significativement plus long que dans l'exemple précédent, en raison de la taille plus importante de la couche d'embedding, et donc d'un nombre de paramètres beaucoup plus élevé. De plus, à cause de cela, nous pourrions avoir besoin d'entraîner notre modèle sur davantage d'exemples si nous voulons éviter le surapprentissage.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 12,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"3200: acc=0.6359375\n",
|
||||
"6400: acc=0.68109375\n",
|
||||
"9600: acc=0.7067708333333333\n",
|
||||
"12800: acc=0.723671875\n",
|
||||
"16000: acc=0.73625\n",
|
||||
"19200: acc=0.7463541666666667\n",
|
||||
"22400: acc=0.7560714285714286\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"(214.1013875559821, 0.7626759436980166)"
|
||||
]
|
||||
},
|
||||
"execution_count": 12,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"train_epoch_emb(net,train_loader, lr=4, epoch_size=25000)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Dans notre cas, nous ne constatons pas une augmentation significative de la précision, ce qui est probablement dû à des vocabulaires très différents. \n",
|
||||
"Pour surmonter le problème des vocabulaires différents, nous pouvons utiliser l'une des solutions suivantes : \n",
|
||||
"* Réentraîner le modèle word2vec sur notre vocabulaire \n",
|
||||
"* Charger notre jeu de données avec le vocabulaire du modèle word2vec pré-entraîné. Le vocabulaire utilisé pour charger le jeu de données peut être spécifié lors du chargement. \n",
|
||||
"\n",
|
||||
"La deuxième approche semble plus simple, surtout parce que le framework `torchtext` de PyTorch contient un support intégré pour les embeddings. Nous pouvons, par exemple, instancier un vocabulaire basé sur GloVe de la manière suivante : \n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 14,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stderr",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"100%|█████████▉| 399999/400000 [00:15<00:00, 25411.14it/s]\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"vocab = torchtext.vocab.GloVe(name='6B', dim=50)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Le vocabulaire chargé propose les opérations de base suivantes : \n",
|
||||
"* Le dictionnaire `vocab.stoi` nous permet de convertir un mot en son index dans le dictionnaire. \n",
|
||||
"* `vocab.itos` fait l'inverse - il convertit un numéro en mot. \n",
|
||||
"* `vocab.vectors` est le tableau des vecteurs d'embedding, donc pour obtenir l'embedding d'un mot `s`, nous devons utiliser `vocab.vectors[vocab.stoi[s]]`. \n",
|
||||
"\n",
|
||||
"Voici un exemple de manipulation des embeddings pour démontrer l'équation **kind-man+woman = queen** (j'ai dû ajuster légèrement le coefficient pour que cela fonctionne) : \n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 15,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"'queen'"
|
||||
]
|
||||
},
|
||||
"execution_count": 15,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"# get the vector corresponding to kind-man+woman\n",
|
||||
"qvec = vocab.vectors[vocab.stoi['king']]-vocab.vectors[vocab.stoi['man']]+1.3*vocab.vectors[vocab.stoi['woman']]\n",
|
||||
"# find the index of the closest embedding vector \n",
|
||||
"d = torch.sum((vocab.vectors-qvec)**2,dim=1)\n",
|
||||
"min_idx = torch.argmin(d)\n",
|
||||
"# find the corresponding word\n",
|
||||
"vocab.itos[min_idx]"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Pour entraîner le classificateur en utilisant ces embeddings, nous devons d'abord encoder notre ensemble de données en utilisant le vocabulaire GloVe :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 16,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"def offsetify(b):\n",
|
||||
" # first, compute data tensor from all sequences\n",
|
||||
" x = [torch.tensor(encode(t[1],voc=vocab)) for t in b] # pass the instance of vocab to encode function!\n",
|
||||
" # now, compute the offsets by accumulating the tensor of sequence lengths\n",
|
||||
" o = [0] + [len(t) for t in x]\n",
|
||||
" o = torch.tensor(o[:-1]).cumsum(dim=0)\n",
|
||||
" return ( \n",
|
||||
" torch.LongTensor([t[0]-1 for t in b]), # labels\n",
|
||||
" torch.cat(x), # text \n",
|
||||
" o\n",
|
||||
" )"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Comme nous l'avons vu ci-dessus, toutes les représentations vectorielles sont stockées dans la matrice `vocab.vectors`. Cela rend extrêmement facile de charger ces poids dans les poids de la couche d'embedding en utilisant une simple copie :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 17,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"net = EmbedClassifier(len(vocab),len(vocab.vectors[0]),len(classes))\n",
|
||||
"net.embedding.weight.data = vocab.vectors\n",
|
||||
"net = net.to(device)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": []
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 18,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"3200: acc=0.6271875\n",
|
||||
"6400: acc=0.68078125\n",
|
||||
"9600: acc=0.7030208333333333\n",
|
||||
"12800: acc=0.71984375\n",
|
||||
"16000: acc=0.7346875\n",
|
||||
"19200: acc=0.7455729166666667\n",
|
||||
"22400: acc=0.7529464285714286\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"(35.53972978646833, 0.7575175943698017)"
|
||||
]
|
||||
},
|
||||
"execution_count": 18,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"train_loader = torch.utils.data.DataLoader(train_dataset, batch_size=16, collate_fn=offsetify, shuffle=True)\n",
|
||||
"train_epoch_emb(net,train_loader, lr=4, epoch_size=25000)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Une des raisons pour lesquelles nous ne constatons pas d'augmentation significative de la précision est le fait que certains mots de notre ensemble de données sont absents du vocabulaire pré-entraîné de GloVe, et sont donc essentiellement ignorés. Pour surmonter ce problème, nous pouvons entraîner nos propres embeddings sur notre ensemble de données.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Contextual Embeddings\n",
|
||||
"\n",
|
||||
"Une des principales limites des représentations d'embeddings préentraînés traditionnels comme Word2Vec est le problème de la désambiguïsation des sens des mots. Bien que les embeddings préentraînés puissent capturer une partie du sens des mots dans un contexte donné, tous les sens possibles d'un mot sont encodés dans le même embedding. Cela peut poser des problèmes dans les modèles en aval, car de nombreux mots, comme le mot \"play\", ont des significations différentes selon le contexte dans lequel ils sont utilisés.\n",
|
||||
"\n",
|
||||
"Par exemple, le mot \"play\" dans ces deux phrases a des significations très différentes :\n",
|
||||
"- Je suis allé voir une **pièce** au théâtre.\n",
|
||||
"- John veut **jouer** avec ses amis.\n",
|
||||
"\n",
|
||||
"Les embeddings préentraînés ci-dessus représentent ces deux significations du mot \"play\" dans le même embedding. Pour surmonter cette limitation, nous devons construire des embeddings basés sur le **modèle de langage**, qui est entraîné sur un large corpus de texte et *comprend* comment les mots peuvent être assemblés dans différents contextes. La discussion sur les embeddings contextuels dépasse le cadre de ce tutoriel, mais nous y reviendrons en abordant les modèles de langage dans la prochaine unité.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"\n---\n\n**Avertissement** : \nCe document a été traduit à l'aide du service de traduction automatique [Co-op Translator](https://github.com/Azure/co-op-translator). Bien que nous nous efforcions d'assurer l'exactitude, veuillez noter que les traductions automatisées peuvent contenir des erreurs ou des inexactitudes. Le document original dans sa langue d'origine doit être considéré comme la source faisant autorité. Pour des informations critiques, il est recommandé de recourir à une traduction professionnelle réalisée par un humain. Nous déclinons toute responsabilité en cas de malentendus ou d'interprétations erronées résultant de l'utilisation de cette traduction.\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
"interpreter": {
|
||||
"hash": "0cb620c6d4b9f7a635928804c26cf22403d89d98d79684e4529119355ee6d5a5"
|
||||
},
|
||||
"kernelspec": {
|
||||
"display_name": "py37_pytorch",
|
||||
"language": "python",
|
||||
"name": "python3"
|
||||
},
|
||||
"language_info": {
|
||||
"codemirror_mode": {
|
||||
"name": "ipython",
|
||||
"version": 3
|
||||
},
|
||||
"file_extension": ".py",
|
||||
"mimetype": "text/x-python",
|
||||
"name": "python",
|
||||
"nbconvert_exporter": "python",
|
||||
"pygments_lexer": "ipython3",
|
||||
"version": "3.8.12"
|
||||
},
|
||||
"coopTranslator": {
|
||||
"original_hash": "f50b026abce5cf36783a560ea72cb9b1",
|
||||
"translation_date": "2025-08-31T15:27:23+00:00",
|
||||
"source_file": "lessons/5-NLP/14-Embeddings/EmbeddingsPyTorch.ipynb",
|
||||
"language_code": "fr"
|
||||
}
|
||||
},
|
||||
"nbformat": 4,
|
||||
"nbformat_minor": 2
|
||||
}
|
||||
|
|
@ -0,0 +1,695 @@
|
|||
{
|
||||
"cells": [
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Intégrations\n",
|
||||
"\n",
|
||||
"Dans notre exemple précédent, nous avons travaillé avec des vecteurs bag-of-words de haute dimension de longueur `vocab_size`, et nous avons explicitement converti des vecteurs de représentation positionnelle de basse dimension en une représentation clairsemée à un seul bit actif. Cette représentation à un seul bit actif n'est pas efficace en termes de mémoire. De plus, chaque mot est traité indépendamment des autres, ce qui fait que les vecteurs encodés de cette manière ne reflètent pas les similitudes sémantiques entre les mots.\n",
|
||||
"\n",
|
||||
"Dans cette unité, nous continuerons à explorer le dataset **News AG**. Pour commencer, chargeons les données et récupérons quelques définitions de l'unité précédente.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 2,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"import tensorflow as tf\n",
|
||||
"from tensorflow import keras\n",
|
||||
"import tensorflow_datasets as tfds\n",
|
||||
"import numpy as np\n",
|
||||
"\n",
|
||||
"ds_train, ds_test = tfds.load('ag_news_subset').values()"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"### Qu'est-ce qu'un embedding ?\n",
|
||||
"\n",
|
||||
"L'idée d'un **embedding** est de représenter les mots à l'aide de vecteurs denses de dimension inférieure qui reflètent le sens sémantique du mot. Nous verrons plus tard comment construire des embeddings de mots significatifs, mais pour l'instant, considérons simplement les embeddings comme un moyen de réduire la dimensionnalité d'un vecteur de mots.\n",
|
||||
"\n",
|
||||
"Ainsi, une couche d'embedding prend un mot en entrée et produit un vecteur de sortie de taille `embedding_size`. En un sens, cela ressemble beaucoup à une couche `Dense`, mais au lieu de prendre un vecteur one-hot encodé en entrée, elle peut prendre un numéro de mot.\n",
|
||||
"\n",
|
||||
"En utilisant une couche d'embedding comme première couche de notre réseau, nous pouvons passer d'un modèle bag-of-words à un modèle **embedding bag**, où nous convertissons d'abord chaque mot de notre texte en l'embedding correspondant, puis calculons une fonction d'agrégation sur tous ces embeddings, comme `sum`, `average` ou `max`.\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"Notre réseau de neurones classificateur se compose des couches suivantes :\n",
|
||||
"\n",
|
||||
"* Une couche `TextVectorization`, qui prend une chaîne de caractères en entrée et produit un tenseur de numéros de tokens. Nous spécifierons une taille de vocabulaire raisonnable `vocab_size` et ignorerons les mots moins fréquemment utilisés. La forme d'entrée sera 1, et la forme de sortie sera $n$, car nous obtiendrons $n$ tokens en résultat, chacun contenant des numéros allant de 0 à `vocab_size`.\n",
|
||||
"* Une couche `Embedding`, qui prend $n$ numéros et réduit chaque numéro à un vecteur dense d'une longueur donnée (100 dans notre exemple). Ainsi, le tenseur d'entrée de forme $n$ sera transformé en un tenseur de forme $n\\times 100$.\n",
|
||||
"* Une couche d'agrégation, qui calcule la moyenne de ce tenseur le long du premier axe, c'est-à-dire qu'elle calculera la moyenne de tous les $n$ tenseurs d'entrée correspondant à différents mots. Pour implémenter cette couche, nous utiliserons une couche `Lambda` et lui passerons la fonction pour calculer la moyenne. La sortie aura une forme de 100, et ce sera la représentation numérique de toute la séquence d'entrée.\n",
|
||||
"* Enfin, un classificateur linéaire `Dense`.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 3,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Model: \"sequential\"\n",
|
||||
"_________________________________________________________________\n",
|
||||
" Layer (type) Output Shape Param # \n",
|
||||
"=================================================================\n",
|
||||
" text_vectorization (TextVec (None, None) 0 \n",
|
||||
" torization) \n",
|
||||
" \n",
|
||||
" embedding (Embedding) (None, None, 100) 3000000 \n",
|
||||
" \n",
|
||||
" lambda (Lambda) (None, 100) 0 \n",
|
||||
" \n",
|
||||
" dense (Dense) (None, 4) 404 \n",
|
||||
" \n",
|
||||
"=================================================================\n",
|
||||
"Total params: 3,000,404\n",
|
||||
"Trainable params: 3,000,404\n",
|
||||
"Non-trainable params: 0\n",
|
||||
"_________________________________________________________________\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"vocab_size = 30000\n",
|
||||
"batch_size = 128\n",
|
||||
"\n",
|
||||
"vectorizer = keras.layers.experimental.preprocessing.TextVectorization(max_tokens=vocab_size,input_shape=(1,))\n",
|
||||
"\n",
|
||||
"model = keras.models.Sequential([\n",
|
||||
" vectorizer, \n",
|
||||
" keras.layers.Embedding(vocab_size,100),\n",
|
||||
" keras.layers.Lambda(lambda x: tf.reduce_mean(x,axis=1)),\n",
|
||||
" keras.layers.Dense(4, activation='softmax')\n",
|
||||
"])\n",
|
||||
"model.summary()"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Dans le résumé, dans la colonne **forme de sortie**, la première dimension du tenseur `None` correspond à la taille du lot (minibatch), et la seconde correspond à la longueur de la séquence de tokens. Toutes les séquences de tokens dans le lot ont des longueurs différentes. Nous verrons comment gérer cela dans la section suivante.\n",
|
||||
"\n",
|
||||
"Passons maintenant à l'entraînement du réseau :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Training vectorizer\n",
|
||||
"938/938 [==============================] - 20s 20ms/step - loss: 0.7891 - acc: 0.8155 - val_loss: 0.4470 - val_acc: 0.8642\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"<keras.callbacks.History at 0x22255515100>"
|
||||
]
|
||||
},
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"def extract_text(x):\n",
|
||||
" return x['title']+' '+x['description']\n",
|
||||
"\n",
|
||||
"def tupelize(x):\n",
|
||||
" return (extract_text(x),x['label'])\n",
|
||||
"\n",
|
||||
"print(\"Training vectorizer\")\n",
|
||||
"vectorizer.adapt(ds_train.take(500).map(extract_text))\n",
|
||||
"\n",
|
||||
"model.compile(loss='sparse_categorical_crossentropy',metrics=['acc'])\n",
|
||||
"model.fit(ds_train.map(tupelize).batch(batch_size),validation_data=ds_test.map(tupelize).batch(batch_size))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {
|
||||
"nteract": {
|
||||
"transient": {
|
||||
"deleting": false
|
||||
}
|
||||
}
|
||||
},
|
||||
"source": [
|
||||
"> **Note** que nous construisons un vectoriseur basé sur un sous-ensemble des données. Cela est fait afin d'accélérer le processus, et cela pourrait entraîner une situation où tous les tokens de notre texte ne sont pas présents dans le vocabulaire. Dans ce cas, ces tokens seraient ignorés, ce qui pourrait entraîner une précision légèrement inférieure. Cependant, dans la réalité, un sous-ensemble de texte donne souvent une bonne estimation du vocabulaire.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"### Gestion des tailles de séquences de variables\n",
|
||||
"\n",
|
||||
"Comprenons comment l'entraînement se déroule dans les mini-lots. Dans l'exemple ci-dessus, le tenseur d'entrée a une dimension de 1, et nous utilisons des mini-lots de taille 128, ce qui donne une taille réelle du tenseur de $128 \\times 1$. Cependant, le nombre de tokens dans chaque phrase est différent. Si nous appliquons la couche `TextVectorization` à une seule entrée, le nombre de tokens retournés varie en fonction de la manière dont le texte est tokenisé :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 5,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"tf.Tensor([ 1 45], shape=(2,), dtype=int64)\n",
|
||||
"tf.Tensor([ 112 1271 1 3 1747 158], shape=(6,), dtype=int64)\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"print(vectorizer('Hello, world!'))\n",
|
||||
"print(vectorizer('I am glad to meet you!'))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Cependant, lorsque nous appliquons le vectoriseur à plusieurs séquences, il doit produire un tenseur de forme rectangulaire, donc il remplit les éléments inutilisés avec le jeton PAD (qui dans notre cas est zéro) :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 6,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"<tf.Tensor: shape=(2, 6), dtype=int64, numpy=\n",
|
||||
"array([[ 1, 45, 0, 0, 0, 0],\n",
|
||||
" [ 112, 1271, 1, 3, 1747, 158]], dtype=int64)>"
|
||||
]
|
||||
},
|
||||
"execution_count": 6,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"vectorizer(['Hello, world!','I am glad to meet you!'])"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Voici les incorporations :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 7,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"array([[[ 1.53059261e-02, 6.80514947e-02, 3.14026810e-02, ...,\n",
|
||||
" -8.92002955e-02, 1.52911525e-04, -5.65562584e-02],\n",
|
||||
" [ 2.57456154e-01, 2.79364467e-01, -2.03605562e-01, ...,\n",
|
||||
" -2.07474351e-01, 8.31158683e-02, -2.03911960e-01],\n",
|
||||
" [ 3.98201384e-02, -8.03454965e-03, 2.39790026e-02, ...,\n",
|
||||
" -7.18549127e-04, 2.66963355e-02, -4.30646613e-02],\n",
|
||||
" [ 3.98201384e-02, -8.03454965e-03, 2.39790026e-02, ...,\n",
|
||||
" -7.18549127e-04, 2.66963355e-02, -4.30646613e-02],\n",
|
||||
" [ 3.98201384e-02, -8.03454965e-03, 2.39790026e-02, ...,\n",
|
||||
" -7.18549127e-04, 2.66963355e-02, -4.30646613e-02],\n",
|
||||
" [ 3.98201384e-02, -8.03454965e-03, 2.39790026e-02, ...,\n",
|
||||
" -7.18549127e-04, 2.66963355e-02, -4.30646613e-02]],\n",
|
||||
"\n",
|
||||
" [[ 1.89674050e-01, 2.61548996e-01, -3.67433839e-02, ...,\n",
|
||||
" -2.07366899e-01, -1.05442435e-01, -2.36952081e-01],\n",
|
||||
" [ 6.16133213e-02, 1.80511594e-01, 9.77298319e-02, ...,\n",
|
||||
" -5.46628237e-02, -1.07340455e-01, -1.06589928e-01],\n",
|
||||
" [ 1.53059261e-02, 6.80514947e-02, 3.14026810e-02, ...,\n",
|
||||
" -8.92002955e-02, 1.52911525e-04, -5.65562584e-02],\n",
|
||||
" [-4.84890305e-02, -8.41715634e-02, 1.51529670e-01, ...,\n",
|
||||
" 1.28192469e-01, -7.77286515e-02, 1.26041949e-01],\n",
|
||||
" [-4.17212099e-02, -5.60694858e-02, 4.08860669e-02, ...,\n",
|
||||
" 8.70475471e-02, 8.92383084e-02, 1.67974353e-01],\n",
|
||||
" [ 2.85779923e-01, 4.57767487e-01, 4.52292450e-02, ...,\n",
|
||||
" -1.97419018e-01, -2.04659685e-01, -2.79758364e-01]]],\n",
|
||||
" dtype=float32)"
|
||||
]
|
||||
},
|
||||
"execution_count": 7,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"model.layers[1](vectorizer(['Hello, world!','I am glad to meet you!'])).numpy()"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"> **Remarque** : Pour minimiser la quantité de remplissage, il peut être judicieux dans certains cas de trier toutes les séquences du jeu de données par ordre croissant de longueur (ou, plus précisément, par nombre de tokens). Cela garantira que chaque minibatch contient des séquences de longueur similaire.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Incrustations sémantiques : Word2Vec\n",
|
||||
"\n",
|
||||
"Dans notre exemple précédent, la couche d'incrustation a appris à mapper des mots à des représentations vectorielles, mais ces représentations n'avaient pas de signification sémantique. Il serait intéressant d'apprendre une représentation vectorielle où des mots similaires ou des synonymes correspondent à des vecteurs proches les uns des autres selon une certaine distance vectorielle (par exemple, la distance euclidienne).\n",
|
||||
"\n",
|
||||
"Pour cela, nous devons préentraîner notre modèle d'incrustation sur une grande collection de textes en utilisant une technique telle que [Word2Vec](https://en.wikipedia.org/wiki/Word2vec). Cette méthode repose sur deux architectures principales utilisées pour produire une représentation distribuée des mots :\n",
|
||||
"\n",
|
||||
" - **Sac de mots continu** (CBoW), où l'on entraîne le modèle à prédire un mot à partir du contexte environnant. Étant donné le ngram $(W_{-2},W_{-1},W_0,W_1,W_2)$, l'objectif du modèle est de prédire $W_0$ à partir de $(W_{-2},W_{-1},W_1,W_2)$.\n",
|
||||
" - **Skip-gram continu**, qui est l'opposé du CBoW. Le modèle utilise la fenêtre de mots du contexte environnant pour prédire le mot actuel.\n",
|
||||
"\n",
|
||||
"CBoW est plus rapide, tandis que skip-gram, bien que plus lent, représente mieux les mots peu fréquents.\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"Pour expérimenter avec l'incrustation Word2Vec préentraînée sur le dataset Google News, nous pouvons utiliser la bibliothèque **gensim**. Ci-dessous, nous trouvons les mots les plus similaires à 'neural'.\n",
|
||||
"\n",
|
||||
"> **Note:** Lorsque vous créez des vecteurs de mots pour la première fois, leur téléchargement peut prendre un certain temps !\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 8,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"import gensim.downloader as api\n",
|
||||
"w2v = api.load('word2vec-google-news-300')"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 12,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"neuronal -> 0.7804799675941467\n",
|
||||
"neurons -> 0.7326500415802002\n",
|
||||
"neural_circuits -> 0.7252851724624634\n",
|
||||
"neuron -> 0.7174385190010071\n",
|
||||
"cortical -> 0.6941086649894714\n",
|
||||
"brain_circuitry -> 0.6923246383666992\n",
|
||||
"synaptic -> 0.6699118614196777\n",
|
||||
"neural_circuitry -> 0.6638563275337219\n",
|
||||
"neurochemical -> 0.6555314064025879\n",
|
||||
"neuronal_activity -> 0.6531826257705688\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"for w,p in w2v.most_similar('neural'):\n",
|
||||
" print(f\"{w} -> {p}\")"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Nous pouvons également extraire l'incorporation vectorielle du mot, à utiliser dans l'entraînement du modèle de classification. L'incorporation comporte 300 composantes, mais ici nous montrons seulement les 20 premières composantes du vecteur pour plus de clarté :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 13,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"array([ 0.01226807, 0.06225586, 0.10693359, 0.05810547, 0.23828125,\n",
|
||||
" 0.03686523, 0.05151367, -0.20703125, 0.01989746, 0.10058594,\n",
|
||||
" -0.03759766, -0.1015625 , -0.15820312, -0.08105469, -0.0390625 ,\n",
|
||||
" -0.05053711, 0.16015625, 0.2578125 , 0.10058594, -0.25976562],\n",
|
||||
" dtype=float32)"
|
||||
]
|
||||
},
|
||||
"execution_count": 13,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"w2v['play'][:20]"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"La grande particularité des embeddings sémantiques est que vous pouvez manipuler l'encodage vectoriel en fonction des sémantiques. Par exemple, nous pouvons demander de trouver un mot dont la représentation vectorielle est aussi proche que possible des mots *roi* et *femme*, et aussi éloignée que possible du mot *homme* :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 14,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"('queen', 0.7118192911148071)"
|
||||
]
|
||||
},
|
||||
"execution_count": 14,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"w2v.most_similar(positive=['king','woman'],negative=['man'])[0]"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {
|
||||
"tags": []
|
||||
},
|
||||
"source": [
|
||||
"Un exemple ci-dessus utilise une certaine magie interne de GenSym, mais la logique sous-jacente est en réalité assez simple. Une chose intéressante à propos des embeddings est que vous pouvez effectuer des opérations vectorielles normales sur les vecteurs d'embedding, et cela refléterait des opérations sur les **significations** des mots. L'exemple ci-dessus peut être exprimé en termes d'opérations vectorielles : nous calculons le vecteur correspondant à **ROI-HOMME+FEMME** (les opérations `+` et `-` sont effectuées sur les représentations vectorielles des mots correspondants), puis nous trouvons le mot le plus proche dans le dictionnaire de ce vecteur :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 15,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"'queen'"
|
||||
]
|
||||
},
|
||||
"execution_count": 15,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"# get the vector corresponding to kind-man+woman\n",
|
||||
"qvec = w2v['king']-1.7*w2v['man']+1.7*w2v['woman']\n",
|
||||
"# find the index of the closest embedding vector \n",
|
||||
"d = np.sum((w2v.vectors-qvec)**2,axis=1)\n",
|
||||
"min_idx = np.argmin(d)\n",
|
||||
"# find the corresponding word\n",
|
||||
"w2v.index_to_key[min_idx]"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"> **NOTE** : Nous avons dû ajouter de petits coefficients aux vecteurs *homme* et *femme* - essayez de les supprimer pour voir ce qui se passe.\n",
|
||||
"\n",
|
||||
"Pour trouver le vecteur le plus proche, nous utilisons les outils de TensorFlow pour calculer un vecteur de distances entre notre vecteur et tous les vecteurs du vocabulaire, puis nous trouvons l'index du mot minimal en utilisant `argmin`.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Bien que Word2Vec semble être un excellent moyen d'exprimer la sémantique des mots, il présente de nombreux inconvénients, notamment les suivants :\n",
|
||||
"\n",
|
||||
"* Les modèles CBoW et skip-gram sont des **représentations prédictives**, et ils ne prennent en compte que le contexte local. Word2Vec ne profite pas du contexte global.\n",
|
||||
"* Word2Vec ne prend pas en compte la **morphologie** des mots, c'est-à-dire le fait que le sens d'un mot peut dépendre de différentes parties du mot, comme la racine.\n",
|
||||
"\n",
|
||||
"**FastText** tente de surmonter cette deuxième limitation et s'appuie sur Word2Vec en apprenant des représentations vectorielles pour chaque mot ainsi que pour les n-grammes de caractères trouvés dans chaque mot. Les valeurs des représentations sont ensuite moyennées en un seul vecteur à chaque étape d'entraînement. Bien que cela ajoute beaucoup de calculs supplémentaires lors de la pré-formation, cela permet aux représentations vectorielles d'intégrer des informations sur les sous-mots.\n",
|
||||
"\n",
|
||||
"Une autre méthode, **GloVe**, utilise une approche différente pour les représentations vectorielles, basée sur la factorisation de la matrice mot-contexte. Tout d'abord, elle construit une grande matrice qui compte le nombre d'occurrences des mots dans différents contextes, puis elle tente de représenter cette matrice dans des dimensions inférieures de manière à minimiser la perte de reconstruction.\n",
|
||||
"\n",
|
||||
"La bibliothèque gensim prend en charge ces représentations vectorielles, et vous pouvez les expérimenter en modifiant le code de chargement du modèle ci-dessus.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Utiliser des embeddings préentraînés dans Keras\n",
|
||||
"\n",
|
||||
"Nous pouvons modifier l'exemple ci-dessus pour préremplir la matrice de notre couche d'embedding avec des embeddings sémantiques, tels que Word2Vec. Les vocabulaires de l'embedding préentraîné et du corpus de texte ne correspondront probablement pas, donc nous devons en choisir un. Ici, nous explorons les deux options possibles : utiliser le vocabulaire du tokenizer et utiliser le vocabulaire des embeddings Word2Vec.\n",
|
||||
"\n",
|
||||
"### Utiliser le vocabulaire du tokenizer\n",
|
||||
"\n",
|
||||
"En utilisant le vocabulaire du tokenizer, certains mots du vocabulaire auront des embeddings Word2Vec correspondants, tandis que d'autres seront absents. Étant donné que la taille de notre vocabulaire est `vocab_size`, et que la longueur du vecteur d'embedding Word2Vec est `embed_size`, la couche d'embedding sera représentée par une matrice de poids de forme `vocab_size`$\\times$`embed_size`. Nous remplirons cette matrice en parcourant le vocabulaire :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 9,
|
||||
"metadata": {
|
||||
"tags": []
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Embedding size: 300\n",
|
||||
"Populating matrix, this will take some time...Done, found 4551 words, 784 words missing\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"embed_size = len(w2v.get_vector('hello'))\n",
|
||||
"print(f'Embedding size: {embed_size}')\n",
|
||||
"\n",
|
||||
"vocab = vectorizer.get_vocabulary()\n",
|
||||
"W = np.zeros((vocab_size,embed_size))\n",
|
||||
"print('Populating matrix, this will take some time...',end='')\n",
|
||||
"found, not_found = 0,0\n",
|
||||
"for i,w in enumerate(vocab):\n",
|
||||
" try:\n",
|
||||
" W[i] = w2v.get_vector(w)\n",
|
||||
" found+=1\n",
|
||||
" except:\n",
|
||||
" # W[i] = np.random.normal(0.0,0.3,size=(embed_size,))\n",
|
||||
" not_found+=1\n",
|
||||
"\n",
|
||||
"print(f\"Done, found {found} words, {not_found} words missing\")"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Pour les mots qui ne sont pas présents dans le vocabulaire de Word2Vec, nous pouvons soit les laisser comme des zéros, soit générer un vecteur aléatoire.\n",
|
||||
"\n",
|
||||
"Nous pouvons maintenant définir une couche d'embedding avec des poids préentraînés :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 10,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"emb = keras.layers.Embedding(vocab_size,embed_size,weights=[W],trainable=False)\n",
|
||||
"model = keras.models.Sequential([\n",
|
||||
" vectorizer, emb,\n",
|
||||
" keras.layers.Lambda(lambda x: tf.reduce_mean(x,axis=1)),\n",
|
||||
" keras.layers.Dense(4, activation='softmax')\n",
|
||||
"])"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": []
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 11,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"938/938 [==============================] - 10s 10ms/step - loss: 1.1075 - acc: 0.7822 - val_loss: 0.9134 - val_acc: 0.8175\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"<keras.callbacks.History at 0x2220226ef10>"
|
||||
]
|
||||
},
|
||||
"execution_count": 11,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"model.compile(loss='sparse_categorical_crossentropy',metrics=['acc'])\n",
|
||||
"model.fit(ds_train.map(tupelize).batch(batch_size),\n",
|
||||
" validation_data=ds_test.map(tupelize).batch(batch_size))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"> **Remarque** : Notez que nous avons défini `trainable=False` lors de la création de `Embedding`, ce qui signifie que nous ne réentraînons pas la couche Embedding. Cela peut entraîner une légère baisse de précision, mais cela accélère l'entraînement.\n",
|
||||
"\n",
|
||||
"### Utilisation du vocabulaire d'embedding\n",
|
||||
"\n",
|
||||
"Un problème avec l'approche précédente est que les vocabulaires utilisés dans TextVectorization et Embedding sont différents. Pour résoudre ce problème, nous pouvons utiliser l'une des solutions suivantes :\n",
|
||||
"* Réentraîner le modèle Word2Vec sur notre vocabulaire.\n",
|
||||
"* Charger notre jeu de données avec le vocabulaire du modèle Word2Vec préentraîné. Les vocabulaires utilisés pour charger le jeu de données peuvent être spécifiés lors du chargement.\n",
|
||||
"\n",
|
||||
"La deuxième approche semble plus simple, alors mettons-la en œuvre. Tout d'abord, nous allons créer une couche `TextVectorization` avec le vocabulaire spécifié, tiré des embeddings Word2Vec :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 12,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"vocab = list(w2v.vocab.keys())\n",
|
||||
"vectorizer = keras.layers.experimental.preprocessing.TextVectorization(input_shape=(1,))\n",
|
||||
"vectorizer.set_vocabulary(vocab)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"La bibliothèque d'embeddings de mots gensim contient une fonction pratique, `get_keras_embeddings`, qui créera automatiquement la couche d'embeddings Keras correspondante pour vous.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 13,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Epoch 1/5\n",
|
||||
"938/938 [==============================] - 20s 14ms/step - loss: 1.3377 - acc: 0.4978 - val_loss: 1.2995 - val_acc: 0.5647\n",
|
||||
"Epoch 2/5\n",
|
||||
"938/938 [==============================] - 10s 10ms/step - loss: 1.2587 - acc: 0.5722 - val_loss: 1.2339 - val_acc: 0.5842\n",
|
||||
"Epoch 3/5\n",
|
||||
"938/938 [==============================] - 10s 10ms/step - loss: 1.1980 - acc: 0.5884 - val_loss: 1.1826 - val_acc: 0.5954\n",
|
||||
"Epoch 4/5\n",
|
||||
"938/938 [==============================] - 12s 13ms/step - loss: 1.1503 - acc: 0.6002 - val_loss: 1.1417 - val_acc: 0.6018\n",
|
||||
"Epoch 5/5\n",
|
||||
"938/938 [==============================] - 11s 12ms/step - loss: 1.1120 - acc: 0.6097 - val_loss: 1.1083 - val_acc: 0.6104\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"<keras.callbacks.History at 0x2220ccb81c0>"
|
||||
]
|
||||
},
|
||||
"execution_count": 13,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"model = keras.models.Sequential([\n",
|
||||
" vectorizer, \n",
|
||||
" w2v.get_keras_embedding(train_embeddings=False),\n",
|
||||
" keras.layers.Lambda(lambda x: tf.reduce_mean(x,axis=1)),\n",
|
||||
" keras.layers.Dense(4, activation='softmax')\n",
|
||||
"])\n",
|
||||
"model.compile(loss='sparse_categorical_crossentropy',metrics=['acc'])\n",
|
||||
"model.fit(ds_train.map(tupelize).batch(128),validation_data=ds_test.map(tupelize).batch(128),epochs=5)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Une des raisons pour lesquelles nous n'observons pas une précision plus élevée est que certains mots de notre ensemble de données sont absents du vocabulaire préentraîné de GloVe, et sont donc essentiellement ignorés. Pour surmonter cela, nous pouvons entraîner nos propres embeddings basés sur notre ensemble de données.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Les embeddings contextuels\n",
|
||||
"\n",
|
||||
"Une des principales limites des représentations d'embeddings préentraînés traditionnels comme Word2Vec est qu'ils ne peuvent pas différencier les différents sens d'un mot, même s'ils peuvent en capturer une partie de la signification. Cela peut poser des problèmes dans les modèles en aval.\n",
|
||||
"\n",
|
||||
"Par exemple, le mot \"play\" a des significations différentes dans ces deux phrases :\n",
|
||||
"- Je suis allé voir une **pièce** au théâtre.\n",
|
||||
"- John veut **jouer** avec ses amis.\n",
|
||||
"\n",
|
||||
"Les embeddings préentraînés dont nous avons parlé représentent les deux sens du mot \"play\" dans le même embedding. Pour surmonter cette limitation, nous devons construire des embeddings basés sur le **modèle de langage**, qui est entraîné sur un large corpus de texte et *sait* comment les mots peuvent être assemblés dans différents contextes. Discuter des embeddings contextuels dépasse le cadre de ce tutoriel, mais nous y reviendrons en parlant des modèles de langage dans la prochaine unité.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"\n---\n\n**Avertissement** : \nCe document a été traduit à l'aide du service de traduction automatique [Co-op Translator](https://github.com/Azure/co-op-translator). Bien que nous nous efforcions d'assurer l'exactitude, veuillez noter que les traductions automatisées peuvent contenir des erreurs ou des inexactitudes. Le document original dans sa langue d'origine doit être considéré comme la source faisant autorité. Pour des informations critiques, il est recommandé de recourir à une traduction professionnelle réalisée par un humain. Nous déclinons toute responsabilité en cas de malentendus ou d'interprétations erronées résultant de l'utilisation de cette traduction.\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
"interpreter": {
|
||||
"hash": "0cb620c6d4b9f7a635928804c26cf22403d89d98d79684e4529119355ee6d5a5"
|
||||
},
|
||||
"kernel_info": {
|
||||
"name": "conda-env-py37_tensorflow-py"
|
||||
},
|
||||
"kernelspec": {
|
||||
"display_name": "py37_tensorflow",
|
||||
"language": "python",
|
||||
"name": "python3"
|
||||
},
|
||||
"language_info": {
|
||||
"codemirror_mode": {
|
||||
"name": "ipython",
|
||||
"version": 3
|
||||
},
|
||||
"file_extension": ".py",
|
||||
"mimetype": "text/x-python",
|
||||
"name": "python",
|
||||
"nbconvert_exporter": "python",
|
||||
"pygments_lexer": "ipython3",
|
||||
"version": "3.8.12"
|
||||
},
|
||||
"nteract": {
|
||||
"version": "nteract-front-end@1.0.0"
|
||||
},
|
||||
"coopTranslator": {
|
||||
"original_hash": "b859482be7f61d1eadc2c6a2720a37e4",
|
||||
"translation_date": "2025-08-31T15:25:15+00:00",
|
||||
"source_file": "lessons/5-NLP/14-Embeddings/EmbeddingsTF.ipynb",
|
||||
"language_code": "fr"
|
||||
}
|
||||
},
|
||||
"nbformat": 4,
|
||||
"nbformat_minor": 4
|
||||
}
|
||||
|
|
@ -0,0 +1,576 @@
|
|||
{
|
||||
"cells": [
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {
|
||||
"id": "NXTSugt6ieXh"
|
||||
},
|
||||
"source": [
|
||||
"## Entraîner un modèle CBoW\n",
|
||||
"\n",
|
||||
"Ce notebook fait partie du [Curriculum AI pour Débutants](http://aka.ms/ai-beginners)\n",
|
||||
"\n",
|
||||
"Dans cet exemple, nous allons apprendre à entraîner un modèle de langage CBoW pour obtenir notre propre espace d'embedding Word2Vec. Nous utiliserons le jeu de données AG News comme source de texte.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"source": [
|
||||
"import torch\n",
|
||||
"import torchtext\n",
|
||||
"import os\n",
|
||||
"import collections\n",
|
||||
"import builtins\n",
|
||||
"import random\n",
|
||||
"import numpy as np"
|
||||
],
|
||||
"metadata": {
|
||||
"id": "q-UiiJUKaxHj"
|
||||
},
|
||||
"execution_count": null,
|
||||
"outputs": []
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"source": [
|
||||
"device = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")"
|
||||
],
|
||||
"metadata": {
|
||||
"id": "TFbR8CZaTZ1q"
|
||||
},
|
||||
"execution_count": null,
|
||||
"outputs": []
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"source": [
|
||||
"Tout d'abord, chargeons notre ensemble de données et définissons le tokenizer et le vocabulaire. Nous allons définir `vocab_size` à 5000 pour limiter un peu les calculs.\n"
|
||||
],
|
||||
"metadata": {
|
||||
"id": "HIwC7lI5T-ov"
|
||||
}
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"source": [
|
||||
"def load_dataset(ngrams = 1, min_freq = 1, vocab_size = 5000 , lines_cnt = 500):\n",
|
||||
" tokenizer = torchtext.data.utils.get_tokenizer('basic_english')\n",
|
||||
" print(\"Loading dataset...\")\n",
|
||||
" test_dataset, train_dataset = torchtext.datasets.AG_NEWS(root='./data')\n",
|
||||
" train_dataset = list(train_dataset)\n",
|
||||
" test_dataset = list(test_dataset)\n",
|
||||
" classes = ['World', 'Sports', 'Business', 'Sci/Tech']\n",
|
||||
" print('Building vocab...')\n",
|
||||
" counter = collections.Counter()\n",
|
||||
" for i, (_, line) in enumerate(train_dataset):\n",
|
||||
" counter.update(torchtext.data.utils.ngrams_iterator(tokenizer(line),ngrams=ngrams))\n",
|
||||
" if i == lines_cnt:\n",
|
||||
" break\n",
|
||||
" vocab = torchtext.vocab.Vocab(collections.Counter(dict(counter.most_common(vocab_size))), min_freq=min_freq)\n",
|
||||
" return train_dataset, test_dataset, classes, vocab, tokenizer"
|
||||
],
|
||||
"metadata": {
|
||||
"id": "wdZuygtgiuLG"
|
||||
},
|
||||
"execution_count": null,
|
||||
"outputs": []
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"source": [
|
||||
"train_dataset, test_dataset, _, vocab, tokenizer = load_dataset()"
|
||||
],
|
||||
"metadata": {
|
||||
"colab": {
|
||||
"base_uri": "https://localhost:8080/"
|
||||
},
|
||||
"id": "4d1nU1gsivGu",
|
||||
"outputId": "949fe272-ae0e-49f5-c373-6703458b3a74"
|
||||
},
|
||||
"execution_count": null,
|
||||
"outputs": [
|
||||
{
|
||||
"output_type": "stream",
|
||||
"name": "stdout",
|
||||
"text": [
|
||||
"Loading dataset...\n",
|
||||
"Building vocab...\n"
|
||||
]
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"source": [
|
||||
"def encode(x, vocabulary, tokenizer = tokenizer):\n",
|
||||
" return [vocabulary[s] for s in tokenizer(x)]"
|
||||
],
|
||||
"metadata": {
|
||||
"id": "1XDYNhG8ToFV"
|
||||
},
|
||||
"execution_count": null,
|
||||
"outputs": []
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {
|
||||
"id": "LIlQk6_PaHVY"
|
||||
},
|
||||
"source": [
|
||||
"## Modèle CBoW\n",
|
||||
"\n",
|
||||
"CBoW apprend à prédire un mot en se basant sur les $2N$ mots voisins. Par exemple, lorsque $N=1$, nous obtiendrons les paires suivantes à partir de la phrase *I like to train networks* : (like, I), (I, like), (to, like), (like, to), (train, to), (to, train), (networks, train), (train, networks). Ici, le premier mot est le mot voisin utilisé comme entrée, et le second mot est celui que nous cherchons à prédire.\n",
|
||||
"\n",
|
||||
"Pour construire un réseau capable de prédire le mot suivant, nous devrons fournir le mot voisin comme entrée et obtenir le numéro du mot en sortie. L'architecture du réseau CBoW est la suivante :\n",
|
||||
"\n",
|
||||
"* Le mot d'entrée est passé à travers la couche d'embedding. Cette même couche d'embedding sera notre embedding Word2Vec, nous la définirons donc séparément comme variable `embedder`. Dans cet exemple, nous utiliserons une taille d'embedding de 30, bien que vous puissiez expérimenter avec des dimensions plus élevées (le Word2Vec réel utilise 300).\n",
|
||||
"* Le vecteur d'embedding sera ensuite passé à une couche linéaire qui prédira le mot en sortie. Cette couche contient donc `vocab_size` neurones.\n",
|
||||
"\n",
|
||||
"Pour la sortie, si nous utilisons `CrossEntropyLoss` comme fonction de perte, nous devrons également fournir uniquement les numéros des mots comme résultats attendus, sans encodage one-hot.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"source": [
|
||||
"vocab_size = len(vocab)\n",
|
||||
"\n",
|
||||
"embedder = torch.nn.Embedding(num_embeddings = vocab_size, embedding_dim = 30)\n",
|
||||
"model = torch.nn.Sequential(\n",
|
||||
" embedder,\n",
|
||||
" torch.nn.Linear(in_features = 30, out_features = vocab_size),\n",
|
||||
")\n",
|
||||
"\n",
|
||||
"print(model)"
|
||||
],
|
||||
"metadata": {
|
||||
"colab": {
|
||||
"base_uri": "https://localhost:8080/"
|
||||
},
|
||||
"id": "akKTcKQKkfl2",
|
||||
"outputId": "da687e3e-a8ec-4c1a-e456-ab8cd6ac7dad"
|
||||
},
|
||||
"execution_count": null,
|
||||
"outputs": [
|
||||
{
|
||||
"output_type": "stream",
|
||||
"name": "stdout",
|
||||
"text": [
|
||||
"Sequential(\n",
|
||||
" (0): Embedding(5002, 30)\n",
|
||||
" (1): Linear(in_features=30, out_features=5002, bias=True)\n",
|
||||
")\n"
|
||||
]
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {
|
||||
"id": "Nud6jgGPaHVa"
|
||||
},
|
||||
"source": [
|
||||
"## Préparation des données d'entraînement\n",
|
||||
"\n",
|
||||
"Programmons maintenant la fonction principale qui calculera les paires de mots CBoW à partir du texte. Cette fonction nous permettra de spécifier la taille de la fenêtre et renverra un ensemble de paires - mot d'entrée et mot de sortie. Notez que cette fonction peut être utilisée sur des mots, ainsi que sur des vecteurs/tenseurs - ce qui nous permettra d'encoder le texte avant de le transmettre à la fonction `to_cbow`.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {
|
||||
"colab": {
|
||||
"base_uri": "https://localhost:8080/"
|
||||
},
|
||||
"id": "x-dsXygOieXn",
|
||||
"outputId": "c2218280-e540-40ba-9546-efe48d0d714f"
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"output_type": "stream",
|
||||
"name": "stdout",
|
||||
"text": [
|
||||
"[['like', 'I'], ['to', 'I'], ['I', 'like'], ['to', 'like'], ['train', 'like'], ['I', 'to'], ['like', 'to'], ['train', 'to'], ['networks', 'to'], ['like', 'train'], ['to', 'train'], ['networks', 'train'], ['to', 'networks'], ['train', 'networks']]\n",
|
||||
"[[232, 172], [5, 172], [172, 232], [5, 232], [0, 232], [172, 5], [232, 5], [0, 5], [1202, 5], [232, 0], [5, 0], [1202, 0], [5, 1202], [0, 1202]]\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"def to_cbow(sent,window_size=2):\n",
|
||||
" res = []\n",
|
||||
" for i,x in enumerate(sent):\n",
|
||||
" for j in range(max(0,i-window_size),min(i+window_size+1,len(sent))):\n",
|
||||
" if i!=j:\n",
|
||||
" res.append([sent[j],x])\n",
|
||||
" return res\n",
|
||||
"\n",
|
||||
"print(to_cbow(['I','like','to','train','networks']))\n",
|
||||
"print(to_cbow(encode('I like to train networks', vocab)))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {
|
||||
"id": "XVaaDLjaaHVb"
|
||||
},
|
||||
"source": [
|
||||
"Préparons le jeu de données d'entraînement. Nous allons parcourir toutes les actualités, appeler `to_cbow` pour obtenir la liste des paires de mots, et ajouter ces paires à `X` et `Y`. Par souci de temps, nous ne considérerons que les 10 000 premiers articles - vous pouvez facilement supprimer cette limitation si vous avez plus de temps à attendre et souhaitez obtenir de meilleures représentations :)\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {
|
||||
"id": "54b-Gd9TieXo"
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"X = []\n",
|
||||
"Y = []\n",
|
||||
"for i, x in zip(range(10000), train_dataset):\n",
|
||||
" for w1, w2 in to_cbow(encode(x[1], vocab), window_size = 5):\n",
|
||||
" X.append(w1)\n",
|
||||
" Y.append(w2)\n",
|
||||
"\n",
|
||||
"X = torch.tensor(X)\n",
|
||||
"Y = torch.tensor(Y)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"source": [
|
||||
"Nous convertirons également ces données en un seul ensemble de données et créerons un dataloader :\n"
|
||||
],
|
||||
"metadata": {
|
||||
"id": "cwWy0PzXWhN5"
|
||||
}
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"source": [
|
||||
"class SimpleIterableDataset(torch.utils.data.IterableDataset):\n",
|
||||
" def __init__(self, X, Y):\n",
|
||||
" super(SimpleIterableDataset).__init__()\n",
|
||||
" self.data = []\n",
|
||||
" for i in range(len(X)):\n",
|
||||
" self.data.append( (Y[i], X[i]) )\n",
|
||||
" random.shuffle(self.data)\n",
|
||||
"\n",
|
||||
" def __iter__(self):\n",
|
||||
" return iter(self.data)"
|
||||
],
|
||||
"metadata": {
|
||||
"id": "mfoAcGPFZU8p"
|
||||
},
|
||||
"execution_count": null,
|
||||
"outputs": []
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {
|
||||
"id": "e4NQ_-5waHVc"
|
||||
},
|
||||
"source": [
|
||||
"Nous convertirons également ces données en un seul ensemble de données et créerons un dataloader :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {
|
||||
"id": "AbLUcojlieXo"
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"ds = SimpleIterableDataset(X, Y)\n",
|
||||
"dl = torch.utils.data.DataLoader(ds, batch_size = 256)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {
|
||||
"id": "pKQr7sXeaHVc"
|
||||
},
|
||||
"source": [
|
||||
"Maintenant, passons à l'entraînement proprement dit. Nous utiliserons l'optimiseur `SGD` avec un taux d'apprentissage assez élevé. Vous pouvez également essayer d'utiliser d'autres optimiseurs, comme `Adam`. Nous allons entraîner pendant 10 époques pour commencer - et vous pouvez relancer cette cellule si vous souhaitez une perte encore plus faible.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"source": [
|
||||
"def train_epoch(net, dataloader, lr = 0.01, optimizer = None, loss_fn = torch.nn.CrossEntropyLoss(), epochs = None, report_freq = 1):\n",
|
||||
" optimizer = optimizer or torch.optim.Adam(net.parameters(), lr = lr)\n",
|
||||
" loss_fn = loss_fn.to(device)\n",
|
||||
" net.train()\n",
|
||||
"\n",
|
||||
" for i in range(epochs):\n",
|
||||
" total_loss, j = 0, 0, \n",
|
||||
" for labels, features in dataloader:\n",
|
||||
" optimizer.zero_grad()\n",
|
||||
" features, labels = features.to(device), labels.to(device)\n",
|
||||
" out = net(features)\n",
|
||||
" loss = loss_fn(out, labels)\n",
|
||||
" loss.backward()\n",
|
||||
" optimizer.step()\n",
|
||||
" total_loss += loss\n",
|
||||
" j += 1\n",
|
||||
" if i % report_freq == 0:\n",
|
||||
" print(f\"Epoch: {i+1}: loss={total_loss.item()/j}\")\n",
|
||||
"\n",
|
||||
" return total_loss.item()/j"
|
||||
],
|
||||
"metadata": {
|
||||
"id": "HeeCYKr_KF1w"
|
||||
},
|
||||
"execution_count": null,
|
||||
"outputs": []
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"source": [
|
||||
"train_epoch(net = model, dataloader = dl, optimizer = torch.optim.SGD(model.parameters(), lr = 0.1), loss_fn = torch.nn.CrossEntropyLoss(), epochs = 10)"
|
||||
],
|
||||
"metadata": {
|
||||
"colab": {
|
||||
"base_uri": "https://localhost:8080/"
|
||||
},
|
||||
"id": "KVgwGtDHgDlT",
|
||||
"outputId": "2447833f-f0e3-4566-c33d-addbfe2f451d"
|
||||
},
|
||||
"execution_count": null,
|
||||
"outputs": [
|
||||
{
|
||||
"output_type": "stream",
|
||||
"name": "stdout",
|
||||
"text": [
|
||||
"Epoch: 1: loss=5.664632366860172\n",
|
||||
"Epoch: 2: loss=5.632101973960962\n",
|
||||
"Epoch: 3: loss=5.610399051405015\n",
|
||||
"Epoch: 4: loss=5.594621561080262\n",
|
||||
"Epoch: 5: loss=5.582538017415446\n",
|
||||
"Epoch: 6: loss=5.572900234519603\n",
|
||||
"Epoch: 7: loss=5.564951676341915\n",
|
||||
"Epoch: 8: loss=5.558288112064614\n",
|
||||
"Epoch: 9: loss=5.552576955031129\n",
|
||||
"Epoch: 10: loss=5.547634165194347\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"output_type": "execute_result",
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"5.547634165194347"
|
||||
]
|
||||
},
|
||||
"metadata": {},
|
||||
"execution_count": 16
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {
|
||||
"id": "W8u2qXZmaHVd"
|
||||
},
|
||||
"source": [
|
||||
"## Essayer Word2Vec\n",
|
||||
"\n",
|
||||
"Pour utiliser Word2Vec, extrayons les vecteurs correspondant à tous les mots de notre vocabulaire :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {
|
||||
"id": "r8TatcXjkU_t"
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"vectors = torch.stack([embedder(torch.tensor(vocab[s])) for s in vocab.itos], 0)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {
|
||||
"id": "3OcX21UOaHVd"
|
||||
},
|
||||
"source": [
|
||||
"Voyons, par exemple, comment le mot **Paris** est encodé en un vecteur :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {
|
||||
"colab": {
|
||||
"base_uri": "https://localhost:8080/"
|
||||
},
|
||||
"id": "bz6tAeLzieXp",
|
||||
"outputId": "5b20850e-4342-45e9-f840-cfac2b4d61d8"
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"output_type": "stream",
|
||||
"name": "stdout",
|
||||
"text": [
|
||||
"tensor([-0.0915, 2.1224, -0.0281, -0.6819, 1.1219, 0.6458, -1.3704, -1.3314,\n",
|
||||
" -1.1437, 0.4496, 0.2301, -0.3515, -0.8485, 1.0481, 0.4386, -0.8949,\n",
|
||||
" 0.5644, 1.0939, -2.5096, 3.2949, -0.2601, -0.8640, 0.1421, -0.0804,\n",
|
||||
" -0.5083, -1.0560, 0.9753, -0.5949, -1.6046, 0.5774],\n",
|
||||
" grad_fn=<EmbeddingBackward>)\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"paris_vec = embedder(torch.tensor(vocab['paris']))\n",
|
||||
"print(paris_vec)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {
|
||||
"id": "pHTJlaeYaHVd"
|
||||
},
|
||||
"source": [
|
||||
"Il est intéressant d'utiliser Word2Vec pour rechercher des synonymes. La fonction suivante retournera les `n` mots les plus proches d'une entrée donnée. Pour les trouver, nous calculons la norme de $|w_i - v|$, où $v$ est le vecteur correspondant à notre mot d'entrée, et $w_i$ est l'encodage du $i$-ème mot dans le vocabulaire. Nous trions ensuite le tableau et retournons les indices correspondants en utilisant `argsort`, puis prenons les premiers `n` éléments de la liste, qui encodent les positions des mots les plus proches dans le vocabulaire.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {
|
||||
"colab": {
|
||||
"base_uri": "https://localhost:8080/"
|
||||
},
|
||||
"id": "NlZyi-_olFar",
|
||||
"outputId": "b5dbb163-88c4-4d5a-eaf2-6751f700e98c"
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"output_type": "execute_result",
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"['microsoft', 'quoted', 'lp', 'rate', 'top']"
|
||||
]
|
||||
},
|
||||
"metadata": {},
|
||||
"execution_count": 56
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"def close_words(x, n = 5):\n",
|
||||
" vec = embedder(torch.tensor(vocab[x]))\n",
|
||||
" top5 = np.linalg.norm(vectors.detach().numpy() - vec.detach().numpy(), axis = 1).argsort()[:n]\n",
|
||||
" return [ vocab.itos[x] for x in top5 ]\n",
|
||||
"\n",
|
||||
"close_words('microsoft')"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {
|
||||
"colab": {
|
||||
"base_uri": "https://localhost:8080/"
|
||||
},
|
||||
"id": "-dQq7xeAln0U",
|
||||
"outputId": "66f768c3-c248-4bfd-ce4f-c8ffc6d0dd0d"
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"output_type": "execute_result",
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"['basketball', 'lot', 'sinai', 'states', 'healthdaynews']"
|
||||
]
|
||||
},
|
||||
"metadata": {},
|
||||
"execution_count": 51
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"close_words('basketball')"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {
|
||||
"colab": {
|
||||
"base_uri": "https://localhost:8080/"
|
||||
},
|
||||
"id": "fJXqK26b29sa",
|
||||
"outputId": "78f0baba-ffd0-485a-dd87-0a12bedfd7fa"
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"output_type": "execute_result",
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"['funds', 'travel', 'sydney', 'japan', 'business']"
|
||||
]
|
||||
},
|
||||
"metadata": {},
|
||||
"execution_count": 77
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"close_words('funds')"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {
|
||||
"id": "My0VeTDd3Ji8"
|
||||
},
|
||||
"source": [
|
||||
"## À retenir\n",
|
||||
"\n",
|
||||
"En utilisant des techniques astucieuses comme CBoW, nous pouvons entraîner un modèle Word2Vec. Vous pouvez également essayer d'entraîner un modèle skip-gram, conçu pour prédire les mots voisins à partir du mot central, et observer ses performances.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"\n---\n\n**Avertissement** : \nCe document a été traduit à l'aide du service de traduction automatique [Co-op Translator](https://github.com/Azure/co-op-translator). Bien que nous nous efforcions d'assurer l'exactitude, veuillez noter que les traductions automatisées peuvent contenir des erreurs ou des inexactitudes. Le document original dans sa langue d'origine doit être considéré comme la source faisant autorité. Pour des informations critiques, il est recommandé de recourir à une traduction professionnelle effectuée par un humain. Nous déclinons toute responsabilité en cas de malentendus ou d'interprétations erronées résultant de l'utilisation de cette traduction.\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
"colab": {
|
||||
"collapsed_sections": [],
|
||||
"name": "CBoW-PyTorch.ipynb",
|
||||
"provenance": []
|
||||
},
|
||||
"interpreter": {
|
||||
"hash": "16af2a8bbb083ea23e5e41c7f5787656b2ce26968575d8763f2c4b17f9cd711f"
|
||||
},
|
||||
"kernelspec": {
|
||||
"display_name": "Python 3.8.12 ('py38')",
|
||||
"language": "python",
|
||||
"name": "python3"
|
||||
},
|
||||
"language_info": {
|
||||
"codemirror_mode": {
|
||||
"name": "ipython",
|
||||
"version": 3
|
||||
},
|
||||
"file_extension": ".py",
|
||||
"mimetype": "text/x-python",
|
||||
"name": "python",
|
||||
"nbconvert_exporter": "python",
|
||||
"pygments_lexer": "ipython3",
|
||||
"version": "3.8.12"
|
||||
},
|
||||
"orig_nbformat": 4,
|
||||
"gpuClass": "standard",
|
||||
"coopTranslator": {
|
||||
"original_hash": "36df28efe3fe40b6fb0a7fa48fe3ea82",
|
||||
"translation_date": "2025-08-31T15:14:16+00:00",
|
||||
"source_file": "lessons/5-NLP/15-LanguageModeling/CBoW-PyTorch.ipynb",
|
||||
"language_code": "fr"
|
||||
}
|
||||
},
|
||||
"nbformat": 4,
|
||||
"nbformat_minor": 0
|
||||
}
|
||||
File diff suppressed because it is too large
Load Diff
|
|
@ -0,0 +1,479 @@
|
|||
{
|
||||
"cells": [
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"# Réseaux de neurones récurrents\n",
|
||||
"\n",
|
||||
"Dans le module précédent, nous avons utilisé des représentations sémantiques riches de texte, associées à un simple classificateur linéaire au-dessus des embeddings. Cette architecture permet de capturer le sens global des mots dans une phrase, mais elle ne prend pas en compte l'**ordre** des mots, car l'opération d'agrégation appliquée aux embeddings supprime cette information issue du texte original. Étant donné que ces modèles ne peuvent pas modéliser l'ordre des mots, ils ne sont pas capables de résoudre des tâches plus complexes ou ambiguës, comme la génération de texte ou la réponse à des questions.\n",
|
||||
"\n",
|
||||
"Pour capturer le sens d'une séquence de texte, nous devons utiliser une autre architecture de réseau de neurones, appelée **réseau de neurones récurrent**, ou RNN. Dans un RNN, nous faisons passer notre phrase à travers le réseau, un symbole à la fois, et le réseau produit un certain **état**, que nous transmettons ensuite au réseau avec le symbole suivant.\n",
|
||||
"\n",
|
||||
"Étant donné la séquence d'entrée de tokens $X_0,\\dots,X_n$, le RNN crée une séquence de blocs de réseau de neurones et entraîne cette séquence de bout en bout à l'aide de la rétropropagation. Chaque bloc de réseau prend une paire $(X_i,S_i)$ en entrée et produit $S_{i+1}$ en sortie. L'état final $S_n$ ou la sortie $X_n$ est ensuite transmis à un classificateur linéaire pour produire le résultat. Tous les blocs de réseau partagent les mêmes poids et sont entraînés de bout en bout en une seule passe de rétropropagation.\n",
|
||||
"\n",
|
||||
"Grâce aux vecteurs d'état $S_0,\\dots,S_n$ qui sont transmis à travers le réseau, celui-ci est capable d'apprendre les dépendances séquentielles entre les mots. Par exemple, lorsque le mot *pas* apparaît quelque part dans la séquence, le réseau peut apprendre à inverser certains éléments du vecteur d'état, ce qui entraîne une négation.\n",
|
||||
"\n",
|
||||
"> Étant donné que les poids de tous les blocs RNN sur l'image sont partagés, la même image peut être représentée par un seul bloc (à droite) avec une boucle de rétroaction récurrente, qui renvoie l'état de sortie du réseau à l'entrée.\n",
|
||||
"\n",
|
||||
"Voyons comment les réseaux de neurones récurrents peuvent nous aider à classifier notre ensemble de données de nouvelles.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 1,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Loading dataset...\n",
|
||||
"Building vocab...\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"import torch\n",
|
||||
"import torchtext\n",
|
||||
"from torchnlp import *\n",
|
||||
"train_dataset, test_dataset, classes, vocab = load_dataset()\n",
|
||||
"vocab_size = len(vocab)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Classificateur RNN simple\n",
|
||||
"\n",
|
||||
"Dans le cas d'un RNN simple, chaque unité récurrente est un réseau linéaire simple, qui prend un vecteur d'entrée concaténé et un vecteur d'état, et produit un nouveau vecteur d'état. PyTorch représente cette unité avec la classe `RNNCell`, et un réseau de telles cellules - comme une couche `RNN`.\n",
|
||||
"\n",
|
||||
"Pour définir un classificateur RNN, nous appliquerons d'abord une couche d'embedding pour réduire la dimensionnalité du vocabulaire d'entrée, puis ajouterons une couche RNN par-dessus :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 2,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"class RNNClassifier(torch.nn.Module):\n",
|
||||
" def __init__(self, vocab_size, embed_dim, hidden_dim, num_class):\n",
|
||||
" super().__init__()\n",
|
||||
" self.hidden_dim = hidden_dim\n",
|
||||
" self.embedding = torch.nn.Embedding(vocab_size, embed_dim)\n",
|
||||
" self.rnn = torch.nn.RNN(embed_dim,hidden_dim,batch_first=True)\n",
|
||||
" self.fc = torch.nn.Linear(hidden_dim, num_class)\n",
|
||||
"\n",
|
||||
" def forward(self, x):\n",
|
||||
" batch_size = x.size(0)\n",
|
||||
" x = self.embedding(x)\n",
|
||||
" x,h = self.rnn(x)\n",
|
||||
" return self.fc(x.mean(dim=1))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"> **Note:** Nous utilisons ici une couche d'embedding non entraînée pour simplifier, mais pour obtenir de meilleurs résultats, nous pouvons utiliser une couche d'embedding pré-entraînée avec des embeddings Word2Vec ou GloVe, comme décrit dans l'unité précédente. Pour mieux comprendre, vous pourriez adapter ce code pour qu'il fonctionne avec des embeddings pré-entraînés.\n",
|
||||
"\n",
|
||||
"Dans notre cas, nous utiliserons un chargeur de données avec padding, de sorte que chaque lot contiendra un certain nombre de séquences remplies pour avoir la même longueur. La couche RNN prendra la séquence de tenseurs d'embedding et produira deux sorties : \n",
|
||||
"* $x$ est une séquence des sorties des cellules RNN à chaque étape \n",
|
||||
"* $h$ est l'état caché final pour le dernier élément de la séquence \n",
|
||||
"\n",
|
||||
"Nous appliquons ensuite un classificateur linéaire entièrement connecté pour obtenir le nombre de classes.\n",
|
||||
"\n",
|
||||
"> **Note:** Les RNN sont assez difficiles à entraîner, car une fois que les cellules RNN sont déroulées sur la longueur de la séquence, le nombre de couches impliquées dans la rétropropagation devient assez important. Par conséquent, nous devons sélectionner un faible taux d'apprentissage et entraîner le réseau sur un ensemble de données plus large pour obtenir de bons résultats. Cela peut prendre beaucoup de temps, donc l'utilisation d'un GPU est préférable.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 3,
|
||||
"metadata": {
|
||||
"scrolled": true
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"3200: acc=0.3090625\n",
|
||||
"6400: acc=0.38921875\n",
|
||||
"9600: acc=0.4590625\n",
|
||||
"12800: acc=0.511953125\n",
|
||||
"16000: acc=0.5506875\n",
|
||||
"19200: acc=0.57921875\n",
|
||||
"22400: acc=0.6070089285714285\n",
|
||||
"25600: acc=0.6304296875\n",
|
||||
"28800: acc=0.6484027777777778\n",
|
||||
"32000: acc=0.66509375\n",
|
||||
"35200: acc=0.6790056818181818\n",
|
||||
"38400: acc=0.6929166666666666\n",
|
||||
"41600: acc=0.7035817307692308\n",
|
||||
"44800: acc=0.7137276785714286\n",
|
||||
"48000: acc=0.72225\n",
|
||||
"51200: acc=0.73001953125\n",
|
||||
"54400: acc=0.7372794117647059\n",
|
||||
"57600: acc=0.7436631944444444\n",
|
||||
"60800: acc=0.7503947368421052\n",
|
||||
"64000: acc=0.75634375\n",
|
||||
"67200: acc=0.7615773809523809\n",
|
||||
"70400: acc=0.7662642045454545\n",
|
||||
"73600: acc=0.7708423913043478\n",
|
||||
"76800: acc=0.7751822916666666\n",
|
||||
"80000: acc=0.7790625\n",
|
||||
"83200: acc=0.7825\n",
|
||||
"86400: acc=0.7858564814814815\n",
|
||||
"89600: acc=0.7890513392857142\n",
|
||||
"92800: acc=0.7920474137931034\n",
|
||||
"96000: acc=0.7952708333333334\n",
|
||||
"99200: acc=0.7982258064516129\n",
|
||||
"102400: acc=0.80099609375\n",
|
||||
"105600: acc=0.8037594696969697\n",
|
||||
"108800: acc=0.8060569852941176\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"train_loader = torch.utils.data.DataLoader(train_dataset, batch_size=16, collate_fn=padify, shuffle=True)\n",
|
||||
"net = RNNClassifier(vocab_size,64,32,len(classes)).to(device)\n",
|
||||
"train_epoch(net,train_loader, lr=0.001)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Mémoire à Long et Court Terme (LSTM)\n",
|
||||
"\n",
|
||||
"L'un des principaux problèmes des RNN classiques est le problème des **gradients qui disparaissent**. Étant donné que les RNN sont entraînés de bout en bout en une seule passe de rétropropagation, il est difficile de propager l'erreur jusqu'aux premières couches du réseau, ce qui empêche le réseau d'apprendre les relations entre des tokens éloignés. Une des façons de contourner ce problème est d'introduire une **gestion explicite de l'état** en utilisant ce qu'on appelle des **portes**. Les deux architectures les plus connues de ce type sont : **Mémoire à Long et Court Terme** (LSTM) et **Unité de Relais Gâtée** (GRU).\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"Le réseau LSTM est organisé de manière similaire au RNN, mais il y a deux états qui sont transmis d'une couche à l'autre : l'état actuel $c$, et le vecteur caché $h$. À chaque unité, le vecteur caché $h_i$ est concaténé avec l'entrée $x_i$, et ils contrôlent ce qui arrive à l'état $c$ via des **portes**. Chaque porte est un réseau neuronal avec une activation sigmoïde (sortie dans la plage $[0,1]$), qui peut être considérée comme un masque binaire lorsqu'elle est multipliée par le vecteur d'état. Les portes suivantes existent (de gauche à droite sur l'image ci-dessus) :\n",
|
||||
"* **Porte d'oubli** : prend le vecteur caché et détermine quelles composantes du vecteur $c$ doivent être oubliées et lesquelles doivent être conservées.\n",
|
||||
"* **Porte d'entrée** : prend certaines informations de l'entrée et du vecteur caché, et les insère dans l'état.\n",
|
||||
"* **Porte de sortie** : transforme l'état via une couche linéaire avec activation $\\tanh$, puis sélectionne certaines de ses composantes en utilisant le vecteur caché $h_i$ pour produire le nouvel état $c_{i+1}$.\n",
|
||||
"\n",
|
||||
"Les composantes de l'état $c$ peuvent être considérées comme des indicateurs qui peuvent être activés ou désactivés. Par exemple, lorsque nous rencontrons un nom comme *Alice* dans une séquence, nous pouvons supposer qu'il fait référence à un personnage féminin et activer l'indicateur dans l'état indiquant que nous avons un nom féminin dans la phrase. Lorsque nous rencontrons ensuite des expressions comme *et Tom*, nous activons l'indicateur indiquant que nous avons un nom au pluriel. Ainsi, en manipulant l'état, nous pouvons théoriquement suivre les propriétés grammaticales des parties de la phrase.\n",
|
||||
"\n",
|
||||
"> **Note** : Une excellente ressource pour comprendre les détails des LSTM est cet article [Understanding LSTM Networks](https://colah.github.io/posts/2015-08-Understanding-LSTMs/) de Christopher Olah.\n",
|
||||
"\n",
|
||||
"Bien que la structure interne d'une cellule LSTM puisse sembler complexe, PyTorch cache cette implémentation dans la classe `LSTMCell` et fournit l'objet `LSTM` pour représenter toute la couche LSTM. Ainsi, l'implémentation d'un classificateur LSTM sera assez similaire à celle du RNN simple que nous avons vu précédemment :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"class LSTMClassifier(torch.nn.Module):\n",
|
||||
" def __init__(self, vocab_size, embed_dim, hidden_dim, num_class):\n",
|
||||
" super().__init__()\n",
|
||||
" self.hidden_dim = hidden_dim\n",
|
||||
" self.embedding = torch.nn.Embedding(vocab_size, embed_dim)\n",
|
||||
" self.embedding.weight.data = torch.randn_like(self.embedding.weight.data)-0.5\n",
|
||||
" self.rnn = torch.nn.LSTM(embed_dim,hidden_dim,batch_first=True)\n",
|
||||
" self.fc = torch.nn.Linear(hidden_dim, num_class)\n",
|
||||
"\n",
|
||||
" def forward(self, x):\n",
|
||||
" batch_size = x.size(0)\n",
|
||||
" x = self.embedding(x)\n",
|
||||
" x,(h,c) = self.rnn(x)\n",
|
||||
" return self.fc(h[-1])"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": []
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 5,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"3200: acc=0.259375\n",
|
||||
"6400: acc=0.25859375\n",
|
||||
"9600: acc=0.26177083333333334\n",
|
||||
"12800: acc=0.2784375\n",
|
||||
"16000: acc=0.313\n",
|
||||
"19200: acc=0.3528645833333333\n",
|
||||
"22400: acc=0.3965625\n",
|
||||
"25600: acc=0.4385546875\n",
|
||||
"28800: acc=0.4752777777777778\n",
|
||||
"32000: acc=0.505375\n",
|
||||
"35200: acc=0.5326704545454546\n",
|
||||
"38400: acc=0.5557552083333334\n",
|
||||
"41600: acc=0.5760817307692307\n",
|
||||
"44800: acc=0.5954910714285714\n",
|
||||
"48000: acc=0.6118333333333333\n",
|
||||
"51200: acc=0.62681640625\n",
|
||||
"54400: acc=0.6404779411764706\n",
|
||||
"57600: acc=0.6520138888888889\n",
|
||||
"60800: acc=0.662828947368421\n",
|
||||
"64000: acc=0.673546875\n",
|
||||
"67200: acc=0.6831547619047619\n",
|
||||
"70400: acc=0.6917897727272727\n",
|
||||
"73600: acc=0.6997146739130434\n",
|
||||
"76800: acc=0.707109375\n",
|
||||
"80000: acc=0.714075\n",
|
||||
"83200: acc=0.7209134615384616\n",
|
||||
"86400: acc=0.727037037037037\n",
|
||||
"89600: acc=0.7326674107142858\n",
|
||||
"92800: acc=0.7379633620689655\n",
|
||||
"96000: acc=0.7433645833333333\n",
|
||||
"99200: acc=0.7479032258064516\n",
|
||||
"102400: acc=0.752119140625\n",
|
||||
"105600: acc=0.7562405303030303\n",
|
||||
"108800: acc=0.76015625\n",
|
||||
"112000: acc=0.7641339285714286\n",
|
||||
"115200: acc=0.7677777777777778\n",
|
||||
"118400: acc=0.7711233108108108\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"(0.03487814127604167, 0.7728)"
|
||||
]
|
||||
},
|
||||
"execution_count": 5,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"net = LSTMClassifier(vocab_size,64,32,len(classes)).to(device)\n",
|
||||
"train_epoch(net,train_loader, lr=0.001)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Séquences compactées\n",
|
||||
"\n",
|
||||
"Dans notre exemple, nous avons dû compléter toutes les séquences du minibatch avec des vecteurs de zéros. Bien que cela entraîne un certain gaspillage de mémoire, avec les RNN, il est encore plus problématique que des cellules RNN supplémentaires soient créées pour les éléments d'entrée complétés, qui participent à l'entraînement mais ne contiennent aucune information d'entrée importante. Il serait bien mieux d'entraîner le RNN uniquement sur la taille réelle des séquences.\n",
|
||||
"\n",
|
||||
"Pour cela, un format spécial de stockage des séquences complétées est introduit dans PyTorch. Supposons que nous ayons un minibatch complété qui ressemble à ceci : \n",
|
||||
"```\n",
|
||||
"[[1,2,3,4,5],\n",
|
||||
" [6,7,8,0,0],\n",
|
||||
" [9,0,0,0,0]]\n",
|
||||
"``` \n",
|
||||
"Ici, 0 représente les valeurs complétées, et le vecteur des longueurs réelles des séquences d'entrée est `[5,3,1]`.\n",
|
||||
"\n",
|
||||
"Pour entraîner efficacement un RNN avec des séquences complétées, nous souhaitons commencer l'entraînement du premier groupe de cellules RNN avec un grand minibatch (`[1,6,9]`), mais ensuite arrêter le traitement de la troisième séquence et continuer l'entraînement avec des minibatches réduits (`[2,7]`, `[3,8]`), et ainsi de suite. Ainsi, une séquence compactée est représentée comme un seul vecteur - dans notre cas `[1,6,9,2,7,3,8,4,5]`, et un vecteur de longueurs (`[5,3,1]`), à partir duquel nous pouvons facilement reconstruire le minibatch complété d'origine.\n",
|
||||
"\n",
|
||||
"Pour produire une séquence compactée, nous pouvons utiliser la fonction `torch.nn.utils.rnn.pack_padded_sequence`. Toutes les couches récurrentes, y compris RNN, LSTM et GRU, prennent en charge les séquences compactées en tant qu'entrée et produisent une sortie compactée, qui peut être décodée à l'aide de `torch.nn.utils.rnn.pad_packed_sequence`.\n",
|
||||
"\n",
|
||||
"Pour pouvoir produire une séquence compactée, nous devons transmettre le vecteur des longueurs au réseau, et donc nous avons besoin d'une fonction différente pour préparer les minibatches :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 6,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"def pad_length(b):\n",
|
||||
" # build vectorized sequence\n",
|
||||
" v = [encode(x[1]) for x in b]\n",
|
||||
" # compute max length of a sequence in this minibatch and length sequence itself\n",
|
||||
" len_seq = list(map(len,v))\n",
|
||||
" l = max(len_seq)\n",
|
||||
" return ( # tuple of three tensors - labels, padded features, length sequence\n",
|
||||
" torch.LongTensor([t[0]-1 for t in b]),\n",
|
||||
" torch.stack([torch.nn.functional.pad(torch.tensor(t),(0,l-len(t)),mode='constant',value=0) for t in v]),\n",
|
||||
" torch.tensor(len_seq)\n",
|
||||
" )\n",
|
||||
"\n",
|
||||
"train_loader_len = torch.utils.data.DataLoader(train_dataset, batch_size=16, collate_fn=pad_length, shuffle=True)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Le réseau réel serait très similaire à `LSTMClassifier` ci-dessus, mais le passage `forward` recevra à la fois le mini-lot avec remplissage et le vecteur des longueurs de séquence. Après avoir calculé l'embedding, nous calculons la séquence empaquetée, la passons à la couche LSTM, puis dépaquetons le résultat.\n",
|
||||
"\n",
|
||||
"> **Note** : En réalité, nous n'utilisons pas le résultat dépaqueté `x`, car nous utilisons la sortie des couches cachées dans les calculs suivants. Ainsi, nous pouvons supprimer complètement le dépaquetage de ce code. La raison pour laquelle nous le plaçons ici est de vous permettre de modifier ce code facilement, au cas où vous auriez besoin d'utiliser la sortie du réseau dans des calculs ultérieurs.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 7,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"class LSTMPackClassifier(torch.nn.Module):\n",
|
||||
" def __init__(self, vocab_size, embed_dim, hidden_dim, num_class):\n",
|
||||
" super().__init__()\n",
|
||||
" self.hidden_dim = hidden_dim\n",
|
||||
" self.embedding = torch.nn.Embedding(vocab_size, embed_dim)\n",
|
||||
" self.embedding.weight.data = torch.randn_like(self.embedding.weight.data)-0.5\n",
|
||||
" self.rnn = torch.nn.LSTM(embed_dim,hidden_dim,batch_first=True)\n",
|
||||
" self.fc = torch.nn.Linear(hidden_dim, num_class)\n",
|
||||
"\n",
|
||||
" def forward(self, x, lengths):\n",
|
||||
" batch_size = x.size(0)\n",
|
||||
" x = self.embedding(x)\n",
|
||||
" pad_x = torch.nn.utils.rnn.pack_padded_sequence(x,lengths,batch_first=True,enforce_sorted=False)\n",
|
||||
" pad_x,(h,c) = self.rnn(pad_x)\n",
|
||||
" x, _ = torch.nn.utils.rnn.pad_packed_sequence(pad_x,batch_first=True)\n",
|
||||
" return self.fc(h[-1])"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": []
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 8,
|
||||
"metadata": {
|
||||
"scrolled": true
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"3200: acc=0.285625\n",
|
||||
"6400: acc=0.33359375\n",
|
||||
"9600: acc=0.3876041666666667\n",
|
||||
"12800: acc=0.44078125\n",
|
||||
"16000: acc=0.4825\n",
|
||||
"19200: acc=0.5235416666666667\n",
|
||||
"22400: acc=0.5559821428571429\n",
|
||||
"25600: acc=0.58609375\n",
|
||||
"28800: acc=0.6116666666666667\n",
|
||||
"32000: acc=0.63340625\n",
|
||||
"35200: acc=0.6525284090909091\n",
|
||||
"38400: acc=0.668515625\n",
|
||||
"41600: acc=0.6822596153846154\n",
|
||||
"44800: acc=0.6948214285714286\n",
|
||||
"48000: acc=0.7052708333333333\n",
|
||||
"51200: acc=0.71521484375\n",
|
||||
"54400: acc=0.7239889705882353\n",
|
||||
"57600: acc=0.7315277777777778\n",
|
||||
"60800: acc=0.7388486842105263\n",
|
||||
"64000: acc=0.74571875\n",
|
||||
"67200: acc=0.7518303571428572\n",
|
||||
"70400: acc=0.7576988636363636\n",
|
||||
"73600: acc=0.7628940217391305\n",
|
||||
"76800: acc=0.7681510416666667\n",
|
||||
"80000: acc=0.7728125\n",
|
||||
"83200: acc=0.7772235576923077\n",
|
||||
"86400: acc=0.7815393518518519\n",
|
||||
"89600: acc=0.7857700892857142\n",
|
||||
"92800: acc=0.7895043103448276\n",
|
||||
"96000: acc=0.7930520833333333\n",
|
||||
"99200: acc=0.7959072580645161\n",
|
||||
"102400: acc=0.798994140625\n",
|
||||
"105600: acc=0.802064393939394\n",
|
||||
"108800: acc=0.8051378676470589\n",
|
||||
"112000: acc=0.8077857142857143\n",
|
||||
"115200: acc=0.8104600694444445\n",
|
||||
"118400: acc=0.8128293918918919\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"(0.029785829671223958, 0.8138166666666666)"
|
||||
]
|
||||
},
|
||||
"execution_count": 8,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"net = LSTMPackClassifier(vocab_size,64,32,len(classes)).to(device)\n",
|
||||
"train_epoch_emb(net,train_loader_len, lr=0.001,use_pack_sequence=True)\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"> **Remarque :** Vous avez peut-être remarqué le paramètre `use_pack_sequence` que nous passons à la fonction d'entraînement. Actuellement, la fonction `pack_padded_sequence` nécessite que le tenseur de séquence de longueur soit sur le périphérique CPU, et donc la fonction d'entraînement doit éviter de déplacer les données de séquence de longueur vers le GPU lors de l'entraînement. Vous pouvez consulter l'implémentation de la fonction `train_emb` dans le fichier [`torchnlp.py`](../../../../../lessons/5-NLP/16-RNN/torchnlp.py).\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## RNN bidirectionnels et multicouches\n",
|
||||
"\n",
|
||||
"Dans nos exemples, tous les réseaux récurrents fonctionnaient dans une seule direction, de l'origine d'une séquence jusqu'à sa fin. Cela semble naturel, car cela ressemble à la manière dont nous lisons et écoutons un discours. Cependant, dans de nombreux cas pratiques où nous avons un accès aléatoire à la séquence d'entrée, il peut être pertinent d'exécuter le calcul récurrent dans les deux directions. Ces réseaux sont appelés **RNN bidirectionnels**, et ils peuvent être créés en passant le paramètre `bidirectional=True` au constructeur RNN/LSTM/GRU.\n",
|
||||
"\n",
|
||||
"Lorsqu'on travaille avec un réseau bidirectionnel, il nous faut deux vecteurs d'état caché, un pour chaque direction. PyTorch encode ces vecteurs en un seul vecteur de taille double, ce qui est assez pratique, car on passe généralement l'état caché résultant à une couche linéaire entièrement connectée, et il suffit de prendre en compte cette augmentation de taille lors de la création de la couche.\n",
|
||||
"\n",
|
||||
"Un réseau récurrent, qu'il soit unidirectionnel ou bidirectionnel, capture certains motifs au sein d'une séquence et peut les stocker dans un vecteur d'état ou les transmettre en sortie. Comme pour les réseaux convolutionnels, on peut construire une autre couche récurrente au-dessus de la première pour capturer des motifs de niveau supérieur, construits à partir des motifs de bas niveau extraits par la première couche. Cela nous amène à la notion de **RNN multicouche**, qui consiste en deux ou plusieurs réseaux récurrents, où la sortie de la couche précédente est transmise à la couche suivante comme entrée.\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"*Image tirée de [cet excellent article](https://towardsdatascience.com/from-a-lstm-cell-to-a-multilayer-lstm-network-with-pytorch-2899eb5696f3) par Fernando López*\n",
|
||||
"\n",
|
||||
"PyTorch simplifie la construction de tels réseaux, car il suffit de passer le paramètre `num_layers` au constructeur RNN/LSTM/GRU pour créer automatiquement plusieurs couches de récurrence. Cela signifie également que la taille du vecteur caché/d'état augmente proportionnellement, et il faut en tenir compte lors de la gestion de la sortie des couches récurrentes.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## RNNs pour d'autres tâches\n",
|
||||
"\n",
|
||||
"Dans cette unité, nous avons vu que les RNNs peuvent être utilisés pour la classification de séquences, mais en réalité, ils peuvent gérer bien d'autres tâches, comme la génération de texte, la traduction automatique, et bien plus encore. Nous aborderons ces tâches dans la prochaine unité.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"\n---\n\n**Avertissement** : \nCe document a été traduit à l'aide du service de traduction automatique [Co-op Translator](https://github.com/Azure/co-op-translator). Bien que nous nous efforcions d'assurer l'exactitude, veuillez noter que les traductions automatisées peuvent contenir des erreurs ou des inexactitudes. Le document original dans sa langue d'origine doit être considéré comme la source faisant autorité. Pour des informations critiques, il est recommandé de recourir à une traduction professionnelle réalisée par un humain. Nous déclinons toute responsabilité en cas de malentendus ou d'interprétations erronées résultant de l'utilisation de cette traduction.\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
"interpreter": {
|
||||
"hash": "16af2a8bbb083ea23e5e41c7f5787656b2ce26968575d8763f2c4b17f9cd711f"
|
||||
},
|
||||
"kernelspec": {
|
||||
"display_name": "Python 3.8.12 ('py38')",
|
||||
"language": "python",
|
||||
"name": "python3"
|
||||
},
|
||||
"language_info": {
|
||||
"codemirror_mode": {
|
||||
"name": "ipython",
|
||||
"version": 3
|
||||
},
|
||||
"file_extension": ".py",
|
||||
"mimetype": "text/x-python",
|
||||
"name": "python",
|
||||
"nbconvert_exporter": "python",
|
||||
"pygments_lexer": "ipython3",
|
||||
"version": "3.8.12"
|
||||
},
|
||||
"coopTranslator": {
|
||||
"original_hash": "522ee52ae3d5ae933e283286254e9a55",
|
||||
"translation_date": "2025-08-31T15:23:20+00:00",
|
||||
"source_file": "lessons/5-NLP/16-RNN/RNNPyTorch.ipynb",
|
||||
"language_code": "fr"
|
||||
}
|
||||
},
|
||||
"nbformat": 4,
|
||||
"nbformat_minor": 2
|
||||
}
|
||||
|
|
@ -0,0 +1,460 @@
|
|||
{
|
||||
"cells": [
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"# Réseaux neuronaux récurrents\n",
|
||||
"\n",
|
||||
"Dans le module précédent, nous avons abordé les représentations sémantiques riches du texte. L'architecture que nous avons utilisée capture le sens global des mots dans une phrase, mais elle ne prend pas en compte l'**ordre** des mots, car l'opération d'agrégation qui suit les embeddings élimine cette information du texte original. Étant donné que ces modèles ne peuvent pas représenter l'ordre des mots, ils ne peuvent pas résoudre des tâches plus complexes ou ambiguës comme la génération de texte ou la réponse à des questions.\n",
|
||||
"\n",
|
||||
"Pour capturer le sens d'une séquence de texte, nous utiliserons une architecture de réseau neuronal appelée **réseau neuronal récurrent**, ou RNN. Lorsqu'on utilise un RNN, on fait passer notre phrase à travers le réseau un jeton à la fois, et le réseau produit un certain **état**, que l'on transmet ensuite au réseau avec le jeton suivant.\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"Étant donné la séquence d'entrée de jetons $X_0,\\dots,X_n$, le RNN crée une séquence de blocs de réseau neuronal et entraîne cette séquence de bout en bout en utilisant la rétropropagation. Chaque bloc de réseau prend une paire $(X_i,S_i)$ en entrée et produit $S_{i+1}$ en résultat. L'état final $S_n$ ou la sortie $Y_n$ est ensuite transmis à un classificateur linéaire pour produire le résultat. Tous les blocs de réseau partagent les mêmes poids et sont entraînés de bout en bout en une seule passe de rétropropagation.\n",
|
||||
"\n",
|
||||
"> La figure ci-dessus montre un réseau neuronal récurrent sous forme déroulée (à gauche) et sous une représentation récurrente plus compacte (à droite). Il est important de comprendre que toutes les cellules RNN partagent les mêmes **poids partageables**.\n",
|
||||
"\n",
|
||||
"Comme les vecteurs d'état $S_0,\\dots,S_n$ sont transmis à travers le réseau, le RNN est capable d'apprendre les dépendances séquentielles entre les mots. Par exemple, lorsque le mot *pas* apparaît quelque part dans la séquence, il peut apprendre à négativer certains éléments dans le vecteur d'état.\n",
|
||||
"\n",
|
||||
"À l'intérieur, chaque cellule RNN contient deux matrices de poids : $W_H$ et $W_I$, ainsi qu'un biais $b$. À chaque étape du RNN, étant donné l'entrée $X_i$ et l'état d'entrée $S_i$, l'état de sortie est calculé comme $S_{i+1} = f(W_H\\times S_i + W_I\\times X_i+b)$, où $f$ est une fonction d'activation (souvent $\\tanh$).\n",
|
||||
"\n",
|
||||
"> Pour des problèmes comme la génération de texte (que nous aborderons dans la prochaine unité) ou la traduction automatique, nous souhaitons également obtenir une valeur de sortie à chaque étape du RNN. Dans ce cas, il y a une autre matrice $W_O$, et la sortie est calculée comme $Y_i=f(W_O\\times S_i+b_O)$.\n",
|
||||
"\n",
|
||||
"Voyons comment les réseaux neuronaux récurrents peuvent nous aider à classifier notre ensemble de données de nouvelles.\n",
|
||||
"\n",
|
||||
"> Pour l'environnement sandbox, nous devons exécuter la cellule suivante pour nous assurer que la bibliothèque requise est installée et que les données sont préchargées. Si vous travaillez en local, vous pouvez ignorer la cellule suivante.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 1,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"import sys\n",
|
||||
"!{sys.executable} -m pip install --quiet tensorflow_datasets==4.4.0\n",
|
||||
"!cd ~ && wget -q -O - https://mslearntensorflowlp.blob.core.windows.net/data/tfds-ag-news.tgz | tar xz"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 2,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"import tensorflow as tf\n",
|
||||
"from tensorflow import keras\n",
|
||||
"import tensorflow_datasets as tfds\n",
|
||||
"import numpy as np\n",
|
||||
"\n",
|
||||
"# We are going to be training pretty large models. In order not to face errors, we need\n",
|
||||
"# to set tensorflow option to grow GPU memory allocation when required\n",
|
||||
"physical_devices = tf.config.list_physical_devices('GPU') \n",
|
||||
"if len(physical_devices)>0:\n",
|
||||
" tf.config.experimental.set_memory_growth(physical_devices[0], True)\n",
|
||||
"\n",
|
||||
"ds_train, ds_test = tfds.load('ag_news_subset').values()"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {
|
||||
"nteract": {
|
||||
"transient": {
|
||||
"deleting": false
|
||||
}
|
||||
}
|
||||
},
|
||||
"source": [
|
||||
"Lors de l'entraînement de modèles de grande taille, l'allocation de mémoire GPU peut poser problème. Nous pourrions également avoir besoin d'expérimenter avec différentes tailles de minibatch, afin que les données tiennent dans la mémoire GPU tout en garantissant un entraînement suffisamment rapide. Si vous exécutez ce code sur votre propre machine équipée d'un GPU, vous pouvez essayer d'ajuster la taille des minibatchs pour accélérer l'entraînement.\n",
|
||||
"\n",
|
||||
"> **Note** : Certaines versions des pilotes NVidia sont connues pour ne pas libérer la mémoire après l'entraînement du modèle. Nous exécutons plusieurs exemples dans ce notebook, ce qui pourrait entraîner une saturation de la mémoire dans certains cas, en particulier si vous réalisez vos propres expériences dans le même notebook. Si vous rencontrez des erreurs étranges au moment de commencer l'entraînement du modèle, il peut être utile de redémarrer le noyau du notebook.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 3,
|
||||
"metadata": {
|
||||
"collapsed": true,
|
||||
"jupyter": {
|
||||
"outputs_hidden": false,
|
||||
"source_hidden": false
|
||||
},
|
||||
"nteract": {
|
||||
"transient": {
|
||||
"deleting": false
|
||||
}
|
||||
}
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"batch_size = 16\n",
|
||||
"embed_size = 64"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Classificateur RNN simple\n",
|
||||
"\n",
|
||||
"Dans le cas d'un RNN simple, chaque unité récurrente est un réseau linéaire simple, qui prend en entrée un vecteur d'entrée et un vecteur d'état, et produit un nouveau vecteur d'état. Dans Keras, cela peut être représenté par la couche `SimpleRNN`.\n",
|
||||
"\n",
|
||||
"Bien que nous puissions transmettre directement des tokens encodés en one-hot à la couche RNN, ce n'est pas une bonne idée en raison de leur haute dimensionnalité. Par conséquent, nous utiliserons une couche d'embedding pour réduire la dimensionnalité des vecteurs de mots, suivie d'une couche RNN, et enfin d'un classificateur `Dense`.\n",
|
||||
"\n",
|
||||
"> **Note** : Dans les cas où la dimensionnalité n'est pas si élevée, par exemple lors de l'utilisation de la tokenisation au niveau des caractères, il peut être pertinent de transmettre directement les tokens encodés en one-hot dans la cellule RNN.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Model: \"sequential\"\n",
|
||||
"_________________________________________________________________\n",
|
||||
"Layer (type) Output Shape Param # \n",
|
||||
"=================================================================\n",
|
||||
"text_vectorization (TextVect (None, None) 0 \n",
|
||||
"_________________________________________________________________\n",
|
||||
"embedding (Embedding) (None, None, 64) 1280000 \n",
|
||||
"_________________________________________________________________\n",
|
||||
"simple_rnn (SimpleRNN) (None, 16) 1296 \n",
|
||||
"_________________________________________________________________\n",
|
||||
"dense (Dense) (None, 4) 68 \n",
|
||||
"=================================================================\n",
|
||||
"Total params: 1,281,364\n",
|
||||
"Trainable params: 1,281,364\n",
|
||||
"Non-trainable params: 0\n",
|
||||
"_________________________________________________________________\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"vocab_size = 20000\n",
|
||||
"\n",
|
||||
"vectorizer = keras.layers.experimental.preprocessing.TextVectorization(\n",
|
||||
" max_tokens=vocab_size,\n",
|
||||
" input_shape=(1,))\n",
|
||||
"\n",
|
||||
"model = keras.models.Sequential([\n",
|
||||
" vectorizer,\n",
|
||||
" keras.layers.Embedding(vocab_size, embed_size),\n",
|
||||
" keras.layers.SimpleRNN(16),\n",
|
||||
" keras.layers.Dense(4,activation='softmax')\n",
|
||||
"])\n",
|
||||
"\n",
|
||||
"model.summary()"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"> **Note:** Nous utilisons ici une couche d'embedding non entraînée pour simplifier, mais pour obtenir de meilleurs résultats, nous pouvons utiliser une couche d'embedding préentraînée avec Word2Vec, comme décrit dans l'unité précédente. Ce serait un bon exercice pour vous d'adapter ce code afin de fonctionner avec des embeddings préentraînés.\n",
|
||||
"\n",
|
||||
"Passons maintenant à l'entraînement de notre RNN. Les RNNs sont généralement assez difficiles à entraîner, car une fois que les cellules RNN sont déroulées sur la longueur de la séquence, le nombre de couches impliquées dans la rétropropagation devient très important. Par conséquent, nous devons sélectionner un taux d'apprentissage plus faible et entraîner le réseau sur un ensemble de données plus large pour obtenir de bons résultats. Cela peut prendre beaucoup de temps, donc l'utilisation d'un GPU est préférable.\n",
|
||||
"\n",
|
||||
"Pour accélérer les choses, nous allons entraîner le modèle RNN uniquement sur les titres des actualités, en omettant la description. Vous pouvez essayer d'entraîner avec la description et voir si vous parvenez à faire fonctionner le modèle.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 5,
|
||||
"metadata": {
|
||||
"scrolled": true
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Training vectorizer\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"def extract_title(x):\n",
|
||||
" return x['title']\n",
|
||||
"\n",
|
||||
"def tupelize_title(x):\n",
|
||||
" return (extract_title(x),x['label'])\n",
|
||||
"\n",
|
||||
"print('Training vectorizer')\n",
|
||||
"vectorizer.adapt(ds_train.take(2000).map(extract_title))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 6,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"7500/7500 [==============================] - 82s 11ms/step - loss: 0.6629 - acc: 0.7623 - val_loss: 0.5559 - val_acc: 0.7995\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"<tensorflow.python.keras.callbacks.History at 0x7f3e0030d350>"
|
||||
]
|
||||
},
|
||||
"execution_count": 6,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"model.compile(loss='sparse_categorical_crossentropy',metrics=['acc'], optimizer='adam')\n",
|
||||
"model.fit(ds_train.map(tupelize_title).batch(batch_size),validation_data=ds_test.map(tupelize_title).batch(batch_size))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {
|
||||
"nteract": {
|
||||
"transient": {
|
||||
"deleting": false
|
||||
}
|
||||
}
|
||||
},
|
||||
"source": [
|
||||
"> **Note** que la précision est probablement plus faible ici, car nous nous entraînons uniquement sur les titres des actualités.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Revoir les séquences de variables\n",
|
||||
"\n",
|
||||
"Rappelez-vous que la couche `TextVectorization` ajoutera automatiquement des tokens de remplissage aux séquences de longueur variable dans un minibatch. Il s'avère que ces tokens participent également à l'entraînement, ce qui peut compliquer la convergence du modèle.\n",
|
||||
"\n",
|
||||
"Il existe plusieurs approches pour minimiser la quantité de remplissage. L'une d'elles consiste à réorganiser le dataset par longueur de séquence et à regrouper toutes les séquences par taille. Cela peut être réalisé en utilisant la fonction `tf.data.experimental.bucket_by_sequence_length` (voir [documentation](https://www.tensorflow.org/api_docs/python/tf/data/experimental/bucket_by_sequence_length)).\n",
|
||||
"\n",
|
||||
"Une autre approche consiste à utiliser **le masquage**. Dans Keras, certaines couches prennent en charge des entrées supplémentaires qui indiquent quels tokens doivent être pris en compte lors de l'entraînement. Pour intégrer le masquage dans notre modèle, nous pouvons soit inclure une couche `Masking` séparée ([docs](https://keras.io/api/layers/core_layers/masking/)), soit spécifier le paramètre `mask_zero=True` dans notre couche `Embedding`.\n",
|
||||
"\n",
|
||||
"> **Note** : Cet entraînement prendra environ 5 minutes pour compléter une époque sur l'ensemble du dataset. N'hésitez pas à interrompre l'entraînement à tout moment si vous manquez de patience. Ce que vous pouvez également faire, c'est limiter la quantité de données utilisées pour l'entraînement, en ajoutant une clause `.take(...)` après les datasets `ds_train` et `ds_test`.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 7,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"7500/7500 [==============================] - 371s 49ms/step - loss: 0.5401 - acc: 0.8079 - val_loss: 0.3780 - val_acc: 0.8822\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"<tensorflow.python.keras.callbacks.History at 0x7f3dec118850>"
|
||||
]
|
||||
},
|
||||
"execution_count": 7,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"def extract_text(x):\n",
|
||||
" return x['title']+' '+x['description']\n",
|
||||
"\n",
|
||||
"def tupelize(x):\n",
|
||||
" return (extract_text(x),x['label'])\n",
|
||||
"\n",
|
||||
"model = keras.models.Sequential([\n",
|
||||
" vectorizer,\n",
|
||||
" keras.layers.Embedding(vocab_size,embed_size,mask_zero=True),\n",
|
||||
" keras.layers.SimpleRNN(16),\n",
|
||||
" keras.layers.Dense(4,activation='softmax')\n",
|
||||
"])\n",
|
||||
"\n",
|
||||
"model.compile(loss='sparse_categorical_crossentropy',metrics=['acc'], optimizer='adam')\n",
|
||||
"model.fit(ds_train.map(tupelize).batch(batch_size),validation_data=ds_test.map(tupelize).batch(batch_size))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Maintenant que nous utilisons le masquage, nous pouvons entraîner le modèle sur l'ensemble du jeu de données des titres et descriptions.\n",
|
||||
"\n",
|
||||
"> **Note** : Avez-vous remarqué que nous avons utilisé un vectoriseur entraîné sur les titres des actualités, et non sur l'intégralité du corps de l'article ? Cela peut potentiellement entraîner l'ignorance de certains tokens, il serait donc préférable de réentraîner le vectoriseur. Cependant, l'impact pourrait être très minime, donc nous continuerons à utiliser le vectoriseur pré-entraîné précédent pour des raisons de simplicité.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## LSTM : Mémoire à long et court terme\n",
|
||||
"\n",
|
||||
"L'un des principaux problèmes des RNN est le phénomène de **gradients évanescents**. Les RNN peuvent être assez longs et peuvent avoir du mal à propager les gradients jusqu'à la première couche du réseau lors de la rétropropagation. Lorsque cela se produit, le réseau ne peut pas apprendre les relations entre des tokens éloignés. Une façon d'éviter ce problème est d'introduire une **gestion explicite de l'état** en utilisant des **portes**. Les deux architectures les plus courantes qui introduisent des portes sont la **mémoire à long et court terme** (LSTM) et l'**unité de relais à portes** (GRU). Nous allons nous concentrer ici sur les LSTM.\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"Un réseau LSTM est organisé de manière similaire à un RNN, mais il y a deux états qui sont transmis de couche en couche : l'état réel $c$ et le vecteur caché $h$. À chaque unité, le vecteur caché $h_{t-1}$ est combiné avec l'entrée $x_t$, et ensemble, ils contrôlent ce qui arrive à l'état $c_t$ et à la sortie $h_{t}$ via des **portes**. Chaque porte utilise une activation sigmoïde (avec une sortie dans l'intervalle $[0,1]$), que l'on peut considérer comme un masque binaire lorsqu'elle est multipliée par le vecteur d'état. Les LSTM possèdent les portes suivantes (de gauche à droite sur l'image ci-dessus) :\n",
|
||||
"* **Porte d'oubli**, qui détermine quelles composantes du vecteur $c_{t-1}$ doivent être oubliées et lesquelles doivent être conservées.\n",
|
||||
"* **Porte d'entrée**, qui détermine la quantité d'informations provenant du vecteur d'entrée et du vecteur caché précédent à incorporer dans le vecteur d'état.\n",
|
||||
"* **Porte de sortie**, qui prend le nouveau vecteur d'état et décide quelles de ses composantes seront utilisées pour produire le nouveau vecteur caché $h_t$.\n",
|
||||
"\n",
|
||||
"Les composantes de l'état $c$ peuvent être considérées comme des indicateurs que l'on peut activer ou désactiver. Par exemple, lorsque nous rencontrons le nom *Alice* dans une séquence, nous supposons qu'il s'agit d'une femme et activons l'indicateur dans l'état qui signale la présence d'un nom féminin dans la phrase. Lorsque nous rencontrons ensuite les mots *et Tom*, nous activons l'indicateur signalant la présence d'un nom pluriel. Ainsi, en manipulant l'état, nous pouvons suivre les propriétés grammaticales de la phrase.\n",
|
||||
"\n",
|
||||
"> **Note** : Voici une excellente ressource pour comprendre les mécanismes internes des LSTM : [Understanding LSTM Networks](https://colah.github.io/posts/2015-08-Understanding-LSTMs/) par Christopher Olah.\n",
|
||||
"\n",
|
||||
"Bien que la structure interne d'une cellule LSTM puisse sembler complexe, Keras masque cette implémentation dans la couche `LSTM`, donc la seule chose que nous devons faire dans l'exemple ci-dessus est de remplacer la couche récurrente :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 8,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"15000/15000 [==============================] - 188s 13ms/step - loss: 0.5692 - acc: 0.7916 - val_loss: 0.3441 - val_acc: 0.8870\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"<tensorflow.python.keras.callbacks.History at 0x7f3d6af5c350>"
|
||||
]
|
||||
},
|
||||
"execution_count": 8,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"model = keras.models.Sequential([\n",
|
||||
" vectorizer,\n",
|
||||
" keras.layers.Embedding(vocab_size, embed_size),\n",
|
||||
" keras.layers.LSTM(8),\n",
|
||||
" keras.layers.Dense(4,activation='softmax')\n",
|
||||
"])\n",
|
||||
"\n",
|
||||
"model.compile(loss='sparse_categorical_crossentropy',metrics=['acc'], optimizer='adam')\n",
|
||||
"model.fit(ds_train.map(tupelize).batch(8),validation_data=ds_test.map(tupelize).batch(8))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": []
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## RNN bidirectionnels et multicouches\n",
|
||||
"\n",
|
||||
"Dans nos exemples jusqu'à présent, les réseaux récurrents fonctionnent du début d'une séquence jusqu'à la fin. Cela nous semble naturel car cela suit la même direction que celle dans laquelle nous lisons ou écoutons un discours. Cependant, pour des scénarios nécessitant un accès aléatoire à la séquence d'entrée, il est plus logique d'exécuter le calcul récurrent dans les deux directions. Les RNN qui permettent des calculs dans les deux directions sont appelés **RNN bidirectionnels**, et ils peuvent être créés en enveloppant la couche récurrente avec une couche spéciale `Bidirectional`.\n",
|
||||
"\n",
|
||||
"> **Note** : La couche `Bidirectional` crée deux copies de la couche qu'elle contient et définit la propriété `go_backwards` de l'une de ces copies sur `True`, ce qui lui permet de parcourir la séquence dans la direction opposée.\n",
|
||||
"\n",
|
||||
"Les réseaux récurrents, qu'ils soient unidirectionnels ou bidirectionnels, capturent des motifs au sein d'une séquence et les stockent dans des vecteurs d'état ou les renvoient comme sortie. Comme pour les réseaux convolutionnels, nous pouvons construire une autre couche récurrente après la première pour capturer des motifs de niveau supérieur, construits à partir des motifs de niveau inférieur extraits par la première couche. Cela nous amène à la notion de **RNN multicouche**, qui consiste en deux réseaux récurrents ou plus, où la sortie de la couche précédente est transmise à la couche suivante comme entrée.\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"*Image tirée de [cet excellent article](https://towardsdatascience.com/from-a-lstm-cell-to-a-multilayer-lstm-network-with-pytorch-2899eb5696f3) par Fernando López.*\n",
|
||||
"\n",
|
||||
"Keras facilite la construction de ces réseaux, car il suffit d'ajouter davantage de couches récurrentes au modèle. Pour toutes les couches sauf la dernière, nous devons spécifier le paramètre `return_sequences=True`, car nous avons besoin que la couche renvoie tous les états intermédiaires, et non seulement l'état final du calcul récurrent.\n",
|
||||
"\n",
|
||||
"Construisons un LSTM bidirectionnel à deux couches pour notre problème de classification.\n",
|
||||
"\n",
|
||||
"> **Note** : Ce code prend encore beaucoup de temps à s'exécuter, mais il nous donne la meilleure précision que nous ayons vue jusqu'à présent. Cela vaut peut-être la peine d'attendre pour voir le résultat.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 9,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"5044/7500 [===================>..........] - ETA: 2:33 - loss: 0.3709 - acc: 0.8706\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\r5045/7500 [===================>..........] - ETA: 2:33 - loss: 0.3709 - acc: 0.8706"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"model = keras.models.Sequential([\n",
|
||||
" vectorizer,\n",
|
||||
" keras.layers.Embedding(vocab_size, 128, mask_zero=True),\n",
|
||||
" keras.layers.Bidirectional(keras.layers.LSTM(64,return_sequences=True)),\n",
|
||||
" keras.layers.Bidirectional(keras.layers.LSTM(64)), \n",
|
||||
" keras.layers.Dense(4,activation='softmax')\n",
|
||||
"])\n",
|
||||
"\n",
|
||||
"model.compile(loss='sparse_categorical_crossentropy',metrics=['acc'], optimizer='adam')\n",
|
||||
"model.fit(ds_train.map(tupelize).batch(batch_size),\n",
|
||||
" validation_data=ds_test.map(tupelize).batch(batch_size))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## RNNs pour d'autres tâches\n",
|
||||
"\n",
|
||||
"Jusqu'à présent, nous nous sommes concentrés sur l'utilisation des RNNs pour classifier des séquences de texte. Mais ils peuvent gérer bien d'autres tâches, comme la génération de texte et la traduction automatique — nous aborderons ces tâches dans la prochaine unité.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"\n---\n\n**Avertissement** : \nCe document a été traduit à l'aide du service de traduction automatique [Co-op Translator](https://github.com/Azure/co-op-translator). Bien que nous nous efforcions d'assurer l'exactitude, veuillez noter que les traductions automatisées peuvent contenir des erreurs ou des inexactitudes. Le document original dans sa langue d'origine doit être considéré comme la source faisant autorité. Pour des informations critiques, il est recommandé de recourir à une traduction professionnelle réalisée par un humain. Nous déclinons toute responsabilité en cas de malentendus ou d'interprétations erronées résultant de l'utilisation de cette traduction.\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
"kernel_info": {
|
||||
"name": "conda-env-py37_tensorflow-py"
|
||||
},
|
||||
"kernelspec": {
|
||||
"display_name": "py37_tensorflow",
|
||||
"language": "python",
|
||||
"name": "conda-env-py37_tensorflow-py"
|
||||
},
|
||||
"language_info": {
|
||||
"codemirror_mode": {
|
||||
"name": "ipython",
|
||||
"version": 3
|
||||
},
|
||||
"file_extension": ".py",
|
||||
"mimetype": "text/x-python",
|
||||
"name": "python",
|
||||
"nbconvert_exporter": "python",
|
||||
"pygments_lexer": "ipython3",
|
||||
"version": "3.7.9"
|
||||
},
|
||||
"nteract": {
|
||||
"version": "nteract-front-end@1.0.0"
|
||||
},
|
||||
"coopTranslator": {
|
||||
"original_hash": "81351e61f619b432ff51010a4f993194",
|
||||
"translation_date": "2025-08-31T15:21:42+00:00",
|
||||
"source_file": "lessons/5-NLP/16-RNN/RNNTF.ipynb",
|
||||
"language_code": "fr"
|
||||
}
|
||||
},
|
||||
"nbformat": 4,
|
||||
"nbformat_minor": 4
|
||||
}
|
||||
|
|
@ -0,0 +1,414 @@
|
|||
{
|
||||
"cells": [
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"# Réseaux génératifs\n",
|
||||
"\n",
|
||||
"Les réseaux neuronaux récurrents (RNN) et leurs variantes à cellules à portes, comme les cellules de mémoire à long court terme (LSTM) et les unités récurrentes à portes (GRU), ont fourni un mécanisme pour la modélisation du langage, c'est-à-dire qu'ils peuvent apprendre l'ordre des mots et fournir des prédictions pour le mot suivant dans une séquence. Cela nous permet d'utiliser les RNN pour des **tâches génératives**, telles que la génération de texte ordinaire, la traduction automatique et même la génération de légendes pour des images.\n",
|
||||
"\n",
|
||||
"Dans l'architecture RNN que nous avons abordée dans l'unité précédente, chaque unité RNN produisait le prochain état caché comme sortie. Cependant, nous pouvons également ajouter une autre sortie à chaque unité récurrente, ce qui nous permettrait de produire une **séquence** (de même longueur que la séquence originale). De plus, nous pouvons utiliser des unités RNN qui n'acceptent pas d'entrée à chaque étape, mais qui prennent simplement un vecteur d'état initial, puis produisent une séquence de sorties.\n",
|
||||
"\n",
|
||||
"Dans ce notebook, nous allons nous concentrer sur des modèles génératifs simples qui nous aident à générer du texte. Pour simplifier, construisons un **réseau au niveau des caractères**, qui génère du texte lettre par lettre. Pendant l'entraînement, nous devons prendre un corpus de texte et le diviser en séquences de lettres.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 1,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Loading dataset...\n",
|
||||
"Building vocab...\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"import torch\n",
|
||||
"import torchtext\n",
|
||||
"import numpy as np\n",
|
||||
"from torchnlp import *\n",
|
||||
"train_dataset,test_dataset,classes,vocab = load_dataset()"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Construire un vocabulaire de caractères\n",
|
||||
"\n",
|
||||
"Pour créer un réseau génératif au niveau des caractères, il est nécessaire de diviser le texte en caractères individuels plutôt qu'en mots. Cela peut être réalisé en définissant un tokenizer différent :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 2,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Vocabulary size = 82\n",
|
||||
"Encoding of 'a' is 1\n",
|
||||
"Character with code 13 is c\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"def char_tokenizer(words):\n",
|
||||
" return list(words) #[word for word in words]\n",
|
||||
"\n",
|
||||
"counter = collections.Counter()\n",
|
||||
"for (label, line) in train_dataset:\n",
|
||||
" counter.update(char_tokenizer(line))\n",
|
||||
"vocab = torchtext.vocab.vocab(counter)\n",
|
||||
"\n",
|
||||
"vocab_size = len(vocab)\n",
|
||||
"print(f\"Vocabulary size = {vocab_size}\")\n",
|
||||
"print(f\"Encoding of 'a' is {vocab.get_stoi()['a']}\")\n",
|
||||
"print(f\"Character with code 13 is {vocab.get_itos()[13]}\")"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Voyons l'exemple de la façon dont nous pouvons encoder le texte de notre ensemble de données :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 3,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"tensor([ 0, 1, 2, 2, 3, 4, 5, 6, 3, 7, 8, 1, 9, 10, 3, 11, 2, 1,\n",
|
||||
" 12, 3, 7, 1, 13, 14, 3, 15, 16, 5, 17, 3, 5, 18, 8, 3, 7, 2,\n",
|
||||
" 1, 13, 14, 3, 19, 20, 8, 21, 5, 8, 9, 10, 22, 3, 20, 8, 21, 5,\n",
|
||||
" 8, 9, 10, 3, 23, 3, 4, 18, 17, 9, 5, 23, 10, 8, 2, 2, 8, 9,\n",
|
||||
" 10, 24, 3, 0, 1, 2, 2, 3, 4, 5, 9, 8, 8, 5, 25, 10, 3, 26,\n",
|
||||
" 12, 27, 16, 26, 2, 27, 16, 28, 29, 30, 1, 16, 26, 3, 17, 31, 3, 21,\n",
|
||||
" 2, 5, 9, 1, 23, 13, 32, 16, 27, 13, 10, 24, 3, 1, 9, 8, 3, 10,\n",
|
||||
" 8, 8, 27, 16, 28, 3, 28, 9, 8, 8, 16, 3, 1, 28, 1, 27, 16, 6])"
|
||||
]
|
||||
},
|
||||
"execution_count": 3,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"def enc(x):\n",
|
||||
" return torch.LongTensor(encode(x,voc=vocab,tokenizer=char_tokenizer))\n",
|
||||
"\n",
|
||||
"enc(train_dataset[0][1])"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Entraîner un RNN génératif\n",
|
||||
"\n",
|
||||
"La manière dont nous allons entraîner un RNN à générer du texte est la suivante. À chaque étape, nous prendrons une séquence de caractères de longueur `nchars` et demanderons au réseau de générer le caractère de sortie suivant pour chaque caractère d'entrée :\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"Selon le scénario spécifique, nous pourrions également vouloir inclure certains caractères spéciaux, tels que *fin de séquence* `<eos>`. Dans notre cas, nous souhaitons simplement entraîner le réseau pour une génération de texte infinie. Par conséquent, nous fixerons la taille de chaque séquence à `nchars` tokens. Ainsi, chaque exemple d'entraînement sera composé de `nchars` entrées et de `nchars` sorties (qui correspondent à la séquence d'entrée décalée d'un symbole vers la gauche). Un minibatch sera constitué de plusieurs de ces séquences.\n",
|
||||
"\n",
|
||||
"La manière dont nous générerons les minibatches consiste à prendre chaque texte d'actualité de longueur `l` et à en extraire toutes les combinaisons possibles entrée-sortie (il y aura `l-nchars` combinaisons de ce type). Ces combinaisons formeront un minibatch, et la taille des minibatches variera à chaque étape d'entraînement.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"(tensor([[ 0, 1, 2, ..., 28, 29, 30],\n",
|
||||
" [ 1, 2, 2, ..., 29, 30, 1],\n",
|
||||
" [ 2, 2, 3, ..., 30, 1, 16],\n",
|
||||
" ...,\n",
|
||||
" [20, 8, 21, ..., 1, 28, 1],\n",
|
||||
" [ 8, 21, 5, ..., 28, 1, 27],\n",
|
||||
" [21, 5, 8, ..., 1, 27, 16]]),\n",
|
||||
" tensor([[ 1, 2, 2, ..., 29, 30, 1],\n",
|
||||
" [ 2, 2, 3, ..., 30, 1, 16],\n",
|
||||
" [ 2, 3, 4, ..., 1, 16, 26],\n",
|
||||
" ...,\n",
|
||||
" [ 8, 21, 5, ..., 28, 1, 27],\n",
|
||||
" [21, 5, 8, ..., 1, 27, 16],\n",
|
||||
" [ 5, 8, 9, ..., 27, 16, 6]]))"
|
||||
]
|
||||
},
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"nchars = 100\n",
|
||||
"\n",
|
||||
"def get_batch(s,nchars=nchars):\n",
|
||||
" ins = torch.zeros(len(s)-nchars,nchars,dtype=torch.long,device=device)\n",
|
||||
" outs = torch.zeros(len(s)-nchars,nchars,dtype=torch.long,device=device)\n",
|
||||
" for i in range(len(s)-nchars):\n",
|
||||
" ins[i] = enc(s[i:i+nchars])\n",
|
||||
" outs[i] = enc(s[i+1:i+nchars+1])\n",
|
||||
" return ins,outs\n",
|
||||
"\n",
|
||||
"get_batch(train_dataset[0][1])"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Définissons maintenant le réseau générateur. Il peut être basé sur n'importe quelle cellule récurrente que nous avons abordée dans l'unité précédente (simple, LSTM ou GRU). Dans notre exemple, nous utiliserons un LSTM.\n",
|
||||
"\n",
|
||||
"Étant donné que le réseau prend des caractères en entrée et que la taille du vocabulaire est assez petite, nous n'avons pas besoin de couche d'embedding ; une entrée encodée en one-hot peut être directement transmise à la cellule LSTM. Cependant, comme nous passons des numéros de caractères en entrée, nous devons les encoder en one-hot avant de les transmettre au LSTM. Cela se fait en appelant la fonction `one_hot` pendant le passage `forward`. L'encodeur de sortie sera une couche linéaire qui convertira l'état caché en une sortie encodée en one-hot.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 5,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"class LSTMGenerator(torch.nn.Module):\n",
|
||||
" def __init__(self, vocab_size, hidden_dim):\n",
|
||||
" super().__init__()\n",
|
||||
" self.rnn = torch.nn.LSTM(vocab_size,hidden_dim,batch_first=True)\n",
|
||||
" self.fc = torch.nn.Linear(hidden_dim, vocab_size)\n",
|
||||
"\n",
|
||||
" def forward(self, x, s=None):\n",
|
||||
" x = torch.nn.functional.one_hot(x,vocab_size).to(torch.float32)\n",
|
||||
" x,s = self.rnn(x,s)\n",
|
||||
" return self.fc(x),s"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Pendant l'entraînement, nous voulons pouvoir échantillonner du texte généré. Pour cela, nous allons définir une fonction `generate` qui produira une chaîne de caractères de longueur `size`, en commençant par la chaîne initiale `start`.\n",
|
||||
"\n",
|
||||
"Voici comment cela fonctionne. Tout d'abord, nous passons la chaîne de départ complète à travers le réseau, et nous obtenons l'état de sortie `s` ainsi que le prochain caractère prédit `out`. Comme `out` est encodé en one-hot, nous utilisons `argmax` pour obtenir l'indice du caractère `nc` dans le vocabulaire, puis nous utilisons `itos` pour déterminer le caractère réel et l'ajouter à la liste résultante de caractères `chars`. Ce processus de génération d'un caractère est répété `size` fois pour générer le nombre requis de caractères.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 8,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"def generate(net,size=100,start='today '):\n",
|
||||
" chars = list(start)\n",
|
||||
" out, s = net(enc(chars).view(1,-1).to(device))\n",
|
||||
" for i in range(size):\n",
|
||||
" nc = torch.argmax(out[0][-1])\n",
|
||||
" chars.append(vocab.get_itos()[nc])\n",
|
||||
" out, s = net(nc.view(1,-1),s)\n",
|
||||
" return ''.join(chars)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Passons à l'entraînement ! La boucle d'entraînement est presque identique à celle de tous nos exemples précédents, mais au lieu d'afficher la précision, nous affichons un texte généré échantillonné tous les 1000 epochs.\n",
|
||||
"\n",
|
||||
"Une attention particulière doit être portée à la manière dont nous calculons la perte. Nous devons calculer la perte en utilisant une sortie encodée en one-hot `out` et le texte attendu `text_out`, qui est la liste des indices de caractères. Heureusement, la fonction `cross_entropy` attend en premier argument la sortie non normalisée du réseau, et en second le numéro de classe, ce qui correspond exactement à ce que nous avons. Elle effectue également une moyenne automatique sur la taille du minibatch.\n",
|
||||
"\n",
|
||||
"Nous limitons également l'entraînement à `samples_to_train` échantillons, afin de ne pas attendre trop longtemps. Nous vous encourageons à expérimenter et à essayer un entraînement plus long, éventuellement sur plusieurs epochs (dans ce cas, vous devrez créer une autre boucle autour de ce code).\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 9,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Current loss = 4.398899078369141\n",
|
||||
"today sr sr sr sr sr sr sr sr sr sr sr sr sr sr sr sr sr sr sr sr sr sr sr sr sr sr sr sr sr sr sr sr sr s\n",
|
||||
"Current loss = 2.161320447921753\n",
|
||||
"today and to the tor to to the tor to to the tor to to the tor to to the tor to to the tor to to the tor t\n",
|
||||
"Current loss = 1.6722588539123535\n",
|
||||
"today and the court to the could to the could to the could to the could to the could to the could to the c\n",
|
||||
"Current loss = 2.423795223236084\n",
|
||||
"today and a second to the conternation of the conternation of the conternation of the conternation of the \n",
|
||||
"Current loss = 1.702607274055481\n",
|
||||
"today and the company to the company to the company to the company to the company to the company to the co\n",
|
||||
"Current loss = 1.692358136177063\n",
|
||||
"today and the company to the company to the company to the company to the company to the company to the co\n",
|
||||
"Current loss = 1.9722288846969604\n",
|
||||
"today and the control the control the control the control the control the control the control the control \n",
|
||||
"Current loss = 1.8705692291259766\n",
|
||||
"today and the second to the second to the second to the second to the second to the second to the second t\n",
|
||||
"Current loss = 1.7626899480819702\n",
|
||||
"today and a security and a security and a security and a security and a security and a security and a secu\n",
|
||||
"Current loss = 1.5574463605880737\n",
|
||||
"today and the company and the company and the company and the company and the company and the company and \n",
|
||||
"Current loss = 1.5620026588439941\n",
|
||||
"today and the be that the be the be that the be the be that the be the be that the be the be that the be t\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"net = LSTMGenerator(vocab_size,64).to(device)\n",
|
||||
"\n",
|
||||
"samples_to_train = 10000\n",
|
||||
"optimizer = torch.optim.Adam(net.parameters(),0.01)\n",
|
||||
"loss_fn = torch.nn.CrossEntropyLoss()\n",
|
||||
"net.train()\n",
|
||||
"for i,x in enumerate(train_dataset):\n",
|
||||
" # x[0] is class label, x[1] is text\n",
|
||||
" if len(x[1])-nchars<10:\n",
|
||||
" continue\n",
|
||||
" samples_to_train-=1\n",
|
||||
" if not samples_to_train: break\n",
|
||||
" text_in, text_out = get_batch(x[1])\n",
|
||||
" optimizer.zero_grad()\n",
|
||||
" out,s = net(text_in)\n",
|
||||
" loss = torch.nn.functional.cross_entropy(out.view(-1,vocab_size),text_out.flatten()) #cross_entropy(out,labels)\n",
|
||||
" loss.backward()\n",
|
||||
" optimizer.step()\n",
|
||||
" if i%1000==0:\n",
|
||||
" print(f\"Current loss = {loss.item()}\")\n",
|
||||
" print(generate(net))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Cet exemple génère déjà un texte de bonne qualité, mais il peut être encore amélioré de plusieurs façons :\n",
|
||||
"\n",
|
||||
"* **Meilleure génération de minibatchs**. La manière dont nous avons préparé les données pour l'entraînement consistait à générer un minibatch à partir d'un seul échantillon. Ce n'est pas idéal, car les minibatchs ont tous des tailles différentes, et certains ne peuvent même pas être générés, car le texte est plus petit que `nchars`. De plus, les petits minibatchs n'exploitent pas suffisamment le GPU. Il serait plus judicieux de prendre un grand bloc de texte à partir de tous les échantillons, de générer ensuite toutes les paires entrée-sortie, de les mélanger, puis de créer des minibatchs de taille égale.\n",
|
||||
"\n",
|
||||
"* **LSTM multicouche**. Il est pertinent d'essayer 2 ou 3 couches de cellules LSTM. Comme nous l'avons mentionné dans l'unité précédente, chaque couche de LSTM extrait certains motifs du texte, et dans le cas d'un générateur au niveau des caractères, on peut s'attendre à ce que les couches inférieures du LSTM soient responsables de l'extraction des syllabes, tandis que les couches supérieures s'occupent des mots et des combinaisons de mots. Cela peut être simplement mis en œuvre en passant un paramètre pour le nombre de couches au constructeur LSTM.\n",
|
||||
"\n",
|
||||
"* Vous pouvez également expérimenter avec des **unités GRU** pour voir lesquelles donnent de meilleurs résultats, ainsi qu'avec **différentes tailles de couches cachées**. Une couche cachée trop grande peut entraîner un surapprentissage (par exemple, le réseau apprendra le texte exact), tandis qu'une taille trop petite pourrait ne pas produire de bons résultats.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Génération de texte souple et température\n",
|
||||
"\n",
|
||||
"Dans la définition précédente de `generate`, nous choisissions toujours le caractère avec la probabilité la plus élevée comme prochain caractère dans le texte généré. Cela avait pour conséquence que le texte \"tournait\" souvent en boucle entre les mêmes séquences de caractères, comme dans cet exemple :\n",
|
||||
"```\n",
|
||||
"today of the second the company and a second the company ...\n",
|
||||
"```\n",
|
||||
"\n",
|
||||
"Cependant, si nous examinons la distribution de probabilité pour le prochain caractère, il se peut que la différence entre quelques-unes des probabilités les plus élevées ne soit pas énorme, par exemple un caractère peut avoir une probabilité de 0,2, un autre de 0,19, etc. Par exemple, lorsqu'on cherche le prochain caractère dans la séquence '*play*', le caractère suivant pourrait tout aussi bien être un espace ou **e** (comme dans le mot *player*).\n",
|
||||
"\n",
|
||||
"Cela nous amène à la conclusion qu'il n'est pas toujours \"juste\" de sélectionner le caractère avec la probabilité la plus élevée, car choisir le deuxième plus probable pourrait également conduire à un texte cohérent. Il est plus judicieux de **prélever un échantillon** des caractères à partir de la distribution de probabilité donnée par la sortie du réseau.\n",
|
||||
"\n",
|
||||
"Ce prélèvement peut être effectué à l'aide de la fonction `multinomial`, qui met en œuvre ce qu'on appelle la **distribution multinomiale**. Une fonction qui implémente cette génération de texte **souple** est définie ci-dessous :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 10,
|
||||
"metadata": {
|
||||
"scrolled": true
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"--- Temperature = 0.3\n",
|
||||
"Today and a company and complete an all the land the restrational the as a security and has provers the pay to and a report and the computer in the stand has filities and working the law the stations for a company and with the company and the final the first company and refight of the state and and workin\n",
|
||||
"\n",
|
||||
"--- Temperature = 0.8\n",
|
||||
"Today he oniis its first to Aus bomblaties the marmation a to manan boogot that pirate assaid a relaid their that goverfin the the Cappets Ecrotional Assonia Cition targets it annight the w scyments Blamity #39;s TVeer Diercheg Reserals fran envyuil that of ster said access what succers of Dour-provelith\n",
|
||||
"\n",
|
||||
"--- Temperature = 1.0\n",
|
||||
"Today holy they a 11 will meda a toket subsuaties, engins for Chanos, they's has stainger past to opening orital his thempting new Nattona was al innerforder advan-than #36;s night year his religuled talitatian what the but with Wednesday to Justment will wemen of Mark CCC Camp as Timed Nae wome a leaders\n",
|
||||
"\n",
|
||||
"--- Temperature = 1.3\n",
|
||||
"Today gpone 2.5 fech atcusion poor cocles toparsdorM.cht Line Pamage put 43 his calt lowed to the book, that has authh-the silia rruch ailing to'ory andhes beutirsimi- Aefffive heading offil an auf eacklets is charged evis, Gunymy oy) Mony has it after-sloythyor loveId out filme, the Natabl -Najuntaxiggs \n",
|
||||
"\n",
|
||||
"--- Temperature = 1.8\n",
|
||||
"Today plary, P.slan chly\\401 mardregationly #39;t 8.1Mide) closes ,filtcon alfly playin roven!\\grea.-QFBEP: Iss onfarchQ/itilia CCf Zivesigntwasta orce.-Peul-aw.uicrin of fuglinfsut aftaningwo, MIEX awayew Aice Woiduar Corvagiugge oppo esig ThusBratourid canthly-RyI.co lagitems\\eexciaishes.conBabntusmor I\n",
|
||||
"\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"def generate_soft(net,size=100,start='today ',temperature=1.0):\n",
|
||||
" chars = list(start)\n",
|
||||
" out, s = net(enc(chars).view(1,-1).to(device))\n",
|
||||
" for i in range(size):\n",
|
||||
" #nc = torch.argmax(out[0][-1])\n",
|
||||
" out_dist = out[0][-1].div(temperature).exp()\n",
|
||||
" nc = torch.multinomial(out_dist,1)[0]\n",
|
||||
" chars.append(vocab.get_itos()[nc])\n",
|
||||
" out, s = net(nc.view(1,-1),s)\n",
|
||||
" return ''.join(chars)\n",
|
||||
" \n",
|
||||
"for i in [0.3,0.8,1.0,1.3,1.8]:\n",
|
||||
" print(f\"--- Temperature = {i}\\n{generate_soft(net,size=300,start='Today ',temperature=i)}\\n\")"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Nous avons introduit un paramètre supplémentaire appelé **température**, qui est utilisé pour indiquer à quel point nous devons nous en tenir à la probabilité la plus élevée. Si la température est de 1,0, nous effectuons un échantillonnage multinomial équitable, et lorsque la température tend vers l'infini - toutes les probabilités deviennent égales, et nous sélectionnons aléatoirement le prochain caractère. Dans l'exemple ci-dessous, nous pouvons observer que le texte devient dénué de sens lorsque nous augmentons trop la température, et qu'il ressemble à un texte \"cyclé\" généré de manière rigide lorsqu'il se rapproche de 0.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"\n---\n\n**Avertissement** : \nCe document a été traduit à l'aide du service de traduction automatique [Co-op Translator](https://github.com/Azure/co-op-translator). Bien que nous nous efforcions d'assurer l'exactitude, veuillez noter que les traductions automatisées peuvent contenir des erreurs ou des inexactitudes. Le document original dans sa langue d'origine doit être considéré comme la source faisant autorité. Pour des informations critiques, il est recommandé de recourir à une traduction professionnelle réalisée par un humain. Nous déclinons toute responsabilité en cas de malentendus ou d'interprétations erronées résultant de l'utilisation de cette traduction.\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
"interpreter": {
|
||||
"hash": "16af2a8bbb083ea23e5e41c7f5787656b2ce26968575d8763f2c4b17f9cd711f"
|
||||
},
|
||||
"kernelspec": {
|
||||
"display_name": "Python 3.8.12 ('py38')",
|
||||
"language": "python",
|
||||
"name": "python3"
|
||||
},
|
||||
"language_info": {
|
||||
"codemirror_mode": {
|
||||
"name": "ipython",
|
||||
"version": 3
|
||||
},
|
||||
"file_extension": ".py",
|
||||
"mimetype": "text/x-python",
|
||||
"name": "python",
|
||||
"nbconvert_exporter": "python",
|
||||
"pygments_lexer": "ipython3",
|
||||
"version": "3.8.12"
|
||||
},
|
||||
"coopTranslator": {
|
||||
"original_hash": "7673cd150d96c74c6d6011460094efb4",
|
||||
"translation_date": "2025-08-31T15:13:04+00:00",
|
||||
"source_file": "lessons/5-NLP/17-GenerativeNetworks/GenerativePyTorch.ipynb",
|
||||
"language_code": "fr"
|
||||
}
|
||||
},
|
||||
"nbformat": 4,
|
||||
"nbformat_minor": 4
|
||||
}
|
||||
|
|
@ -0,0 +1,495 @@
|
|||
{
|
||||
"cells": [
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"# Réseaux génératifs\n",
|
||||
"\n",
|
||||
"Les réseaux neuronaux récurrents (RNNs) et leurs variantes à cellules à portes, comme les cellules Long Short Term Memory (LSTMs) et les Gated Recurrent Units (GRUs), ont introduit un mécanisme pour la modélisation du langage, c'est-à-dire qu'ils peuvent apprendre l'ordre des mots et fournir des prédictions pour le mot suivant dans une séquence. Cela nous permet d'utiliser les RNNs pour des **tâches génératives**, telles que la génération de texte ordinaire, la traduction automatique et même la génération de légendes pour des images.\n",
|
||||
"\n",
|
||||
"Dans l'architecture RNN que nous avons abordée dans l'unité précédente, chaque unité RNN produisait le prochain état caché comme sortie. Cependant, nous pouvons également ajouter une autre sortie à chaque unité récurrente, ce qui nous permettrait de produire une **séquence** (de même longueur que la séquence d'origine). De plus, nous pouvons utiliser des unités RNN qui n'acceptent pas d'entrée à chaque étape, mais qui prennent simplement un vecteur d'état initial, puis produisent une séquence de sorties.\n",
|
||||
"\n",
|
||||
"Dans ce notebook, nous allons nous concentrer sur des modèles génératifs simples qui nous aident à générer du texte. Pour simplifier, construisons un **réseau au niveau des caractères**, qui génère du texte lettre par lettre. Pendant l'entraînement, nous devons prendre un corpus de texte et le diviser en séquences de lettres.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 1,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"import tensorflow as tf\n",
|
||||
"from tensorflow import keras\n",
|
||||
"import tensorflow_datasets as tfds\n",
|
||||
"import numpy as np\n",
|
||||
"\n",
|
||||
"ds_train, ds_test = tfds.load('ag_news_subset').values()"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Construire un vocabulaire de caractères\n",
|
||||
"\n",
|
||||
"Pour créer un réseau génératif au niveau des caractères, nous devons diviser le texte en caractères individuels plutôt qu'en mots. La couche `TextVectorization` que nous avons utilisée auparavant ne peut pas le faire, donc nous avons deux options :\n",
|
||||
"\n",
|
||||
"* Charger manuellement le texte et effectuer la tokenisation \"à la main\", comme dans [cet exemple officiel de Keras](https://keras.io/examples/generative/lstm_character_level_text_generation/)\n",
|
||||
"* Utiliser la classe `Tokenizer` pour la tokenisation au niveau des caractères.\n",
|
||||
"\n",
|
||||
"Nous allons opter pour la deuxième option. `Tokenizer` peut également être utilisé pour la tokenisation en mots, ce qui permet de passer facilement de la tokenisation au niveau des caractères à celle au niveau des mots.\n",
|
||||
"\n",
|
||||
"Pour effectuer une tokenisation au niveau des caractères, nous devons passer le paramètre `char_level=True` :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 2,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"def extract_text(x):\n",
|
||||
" return x['title']+' '+x['description']\n",
|
||||
"\n",
|
||||
"def tupelize(x):\n",
|
||||
" return (extract_text(x),x['label'])\n",
|
||||
"\n",
|
||||
"tokenizer = keras.preprocessing.text.Tokenizer(char_level=True,lower=False)\n",
|
||||
"tokenizer.fit_on_texts([x['title'].numpy().decode('utf-8') for x in ds_train])"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Nous voulons également utiliser un jeton spécial pour indiquer **fin de séquence**, que nous appellerons `<eos>`. Ajoutons-le manuellement au vocabulaire :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 3,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"eos_token = len(tokenizer.word_index)+1\n",
|
||||
"tokenizer.word_index['<eos>'] = eos_token\n",
|
||||
"\n",
|
||||
"vocab_size = eos_token + 1"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": []
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"[[48, 2, 10, 10, 5, 44, 1, 25, 5, 8, 10, 13, 78]]"
|
||||
]
|
||||
},
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"tokenizer.texts_to_sequences(['Hello, world!'])"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Entraîner un RNN génératif à créer des titres\n",
|
||||
"\n",
|
||||
"La manière dont nous allons entraîner un RNN à générer des titres d'actualités est la suivante. À chaque étape, nous prendrons un titre, qui sera introduit dans un RNN, et pour chaque caractère d'entrée, nous demanderons au réseau de générer le caractère de sortie suivant :\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"Pour le dernier caractère de notre séquence, nous demanderons au réseau de générer le token `<eos>`.\n",
|
||||
"\n",
|
||||
"La principale différence avec le RNN génératif que nous utilisons ici est que nous prendrons une sortie à chaque étape du RNN, et pas seulement à partir de la cellule finale. Cela peut être réalisé en spécifiant le paramètre `return_sequences` à la cellule RNN.\n",
|
||||
"\n",
|
||||
"Ainsi, pendant l'entraînement, une entrée pour le réseau sera une séquence de caractères encodés d'une certaine longueur, et une sortie sera une séquence de la même longueur, mais décalée d'un élément et terminée par `<eos>`. Un minibatch sera constitué de plusieurs de ces séquences, et nous devrons utiliser **padding** pour aligner toutes les séquences.\n",
|
||||
"\n",
|
||||
"Créons des fonctions qui transformeront le jeu de données pour nous. Comme nous voulons ajouter du padding aux séquences au niveau du minibatch, nous commencerons par regrouper le jeu de données en appelant `.batch()`, puis nous utiliserons `map` pour effectuer la transformation. Ainsi, la fonction de transformation prendra un minibatch entier comme paramètre :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 5,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"def title_batch(x):\n",
|
||||
" x = [t.numpy().decode('utf-8') for t in x]\n",
|
||||
" z = tokenizer.texts_to_sequences(x)\n",
|
||||
" z = tf.keras.preprocessing.sequence.pad_sequences(z)\n",
|
||||
" return tf.one_hot(z,vocab_size), tf.one_hot(tf.concat([z[:,1:],tf.constant(eos_token,shape=(len(z),1))],axis=1),vocab_size)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Quelques points importants que nous faisons ici :\n",
|
||||
"* Nous commençons par extraire le texte réel du tenseur de chaînes\n",
|
||||
"* `text_to_sequences` convertit la liste de chaînes en une liste de tenseurs d'entiers\n",
|
||||
"* `pad_sequences` remplit ces tenseurs jusqu'à leur longueur maximale\n",
|
||||
"* Enfin, nous encodons tous les caractères en one-hot, tout en effectuant le décalage et en ajoutant `<eos>`. Nous verrons bientôt pourquoi nous avons besoin de caractères encodés en one-hot.\n",
|
||||
"\n",
|
||||
"Cependant, cette fonction est **Pythonique**, c'est-à-dire qu'elle ne peut pas être automatiquement traduite en graphe computationnel Tensorflow. Nous obtiendrons des erreurs si nous essayons d'utiliser cette fonction directement dans la fonction `Dataset.map`. Nous devons encapsuler cet appel Pythonique en utilisant le wrapper `py_function` :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 6,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"def title_batch_fn(x):\n",
|
||||
" x = x['title']\n",
|
||||
" a,b = tf.py_function(title_batch,inp=[x],Tout=(tf.float32,tf.float32))\n",
|
||||
" return a,b"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"> **Note** : Différencier entre les fonctions de transformation Pythonic et Tensorflow peut sembler un peu trop complexe, et vous pourriez vous demander pourquoi nous ne transformons pas le dataset en utilisant des fonctions Python standard avant de le passer à `fit`. Bien que cela soit tout à fait possible, utiliser `Dataset.map` présente un énorme avantage, car le pipeline de transformation des données est exécuté via le graphe computationnel de Tensorflow, ce qui permet de tirer parti des calculs sur GPU et de minimiser le besoin de transférer les données entre le CPU et le GPU.\n",
|
||||
"\n",
|
||||
"Nous pouvons maintenant construire notre réseau générateur et commencer l'entraînement. Il peut être basé sur n'importe quelle cellule récurrente que nous avons abordée dans l'unité précédente (simple, LSTM ou GRU). Dans notre exemple, nous utiliserons LSTM.\n",
|
||||
"\n",
|
||||
"Étant donné que le réseau prend des caractères en entrée et que la taille du vocabulaire est assez petite, nous n'avons pas besoin de couche d'embedding ; une entrée encodée en one-hot peut directement être transmise à la cellule LSTM. La couche de sortie sera un classificateur `Dense` qui convertira la sortie du LSTM en numéros de tokens encodés en one-hot.\n",
|
||||
"\n",
|
||||
"De plus, comme nous travaillons avec des séquences de longueur variable, nous pouvons utiliser une couche `Masking` pour créer un masque qui ignorera la partie remplie de la chaîne. Ce n'est pas strictement nécessaire, car nous ne sommes pas particulièrement intéressés par tout ce qui dépasse le token `<eos>`, mais nous l'utiliserons dans le but d'acquérir de l'expérience avec ce type de couche. `input_shape` serait `(None, vocab_size)`, où `None` indique une séquence de longueur variable, et la forme de sortie est également `(None, vocab_size)`, comme vous pouvez le voir dans le `summary` :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 7,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Model: \"sequential\"\n",
|
||||
"_________________________________________________________________\n",
|
||||
"Layer (type) Output Shape Param # \n",
|
||||
"=================================================================\n",
|
||||
"masking (Masking) (None, None, 84) 0 \n",
|
||||
"_________________________________________________________________\n",
|
||||
"lstm (LSTM) (None, None, 128) 109056 \n",
|
||||
"_________________________________________________________________\n",
|
||||
"dense (Dense) (None, None, 84) 10836 \n",
|
||||
"=================================================================\n",
|
||||
"Total params: 119,892\n",
|
||||
"Trainable params: 119,892\n",
|
||||
"Non-trainable params: 0\n",
|
||||
"_________________________________________________________________\n",
|
||||
"15000/15000 [==============================] - 229s 15ms/step - loss: 1.5385\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"<tensorflow.python.keras.callbacks.History at 0x7fa40c1245e0>"
|
||||
]
|
||||
},
|
||||
"execution_count": 7,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"model = keras.models.Sequential([\n",
|
||||
" keras.layers.Masking(input_shape=(None,vocab_size)),\n",
|
||||
" keras.layers.LSTM(128,return_sequences=True),\n",
|
||||
" keras.layers.Dense(vocab_size,activation='softmax')\n",
|
||||
"])\n",
|
||||
"\n",
|
||||
"model.summary()\n",
|
||||
"model.compile(loss='categorical_crossentropy')\n",
|
||||
"\n",
|
||||
"model.fit(ds_train.batch(8).map(title_batch_fn))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Génération de sortie\n",
|
||||
"\n",
|
||||
"Maintenant que nous avons entraîné le modèle, nous souhaitons l'utiliser pour générer une sortie. Tout d'abord, nous avons besoin d'une méthode pour décoder le texte représenté par une séquence de numéros de tokens. Pour cela, nous pourrions utiliser la fonction `tokenizer.sequences_to_texts` ; cependant, elle ne fonctionne pas bien avec une tokenisation au niveau des caractères. Par conséquent, nous allons prendre un dictionnaire de tokens provenant du tokenizer (appelé `word_index`), construire une correspondance inversée, et écrire notre propre fonction de décodage :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 10,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"reverse_map = {val:key for key, val in tokenizer.word_index.items()}\n",
|
||||
"\n",
|
||||
"def decode(x):\n",
|
||||
" return ''.join([reverse_map[t] for t in x])"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Commençons par une chaîne `start`, que nous encodons en une séquence `inp`, puis à chaque étape, nous appelons notre réseau pour déduire le caractère suivant.\n",
|
||||
"\n",
|
||||
"La sortie du réseau `out` est un vecteur de `vocab_size` éléments représentant les probabilités de chaque jeton, et nous pouvons trouver le numéro du jeton le plus probable en utilisant `argmax`. Nous ajoutons ensuite ce caractère à la liste des jetons générés et poursuivons la génération. Ce processus de génération d'un caractère est répété `size` fois pour produire le nombre requis de caractères, et nous terminons plus tôt si le `eos_token` est rencontré.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 12,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"'Today #39;s lead to strike for the strike for the strike for the strike (AFP)'"
|
||||
]
|
||||
},
|
||||
"execution_count": 12,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"def generate(model,size=100,start='Today '):\n",
|
||||
" inp = tokenizer.texts_to_sequences([start])[0]\n",
|
||||
" chars = inp\n",
|
||||
" for i in range(size):\n",
|
||||
" out = model(tf.expand_dims(tf.one_hot(inp,vocab_size),0))[0][-1]\n",
|
||||
" nc = tf.argmax(out)\n",
|
||||
" if nc==eos_token:\n",
|
||||
" break\n",
|
||||
" chars.append(nc.numpy())\n",
|
||||
" inp = inp+[nc]\n",
|
||||
" return decode(chars)\n",
|
||||
" \n",
|
||||
"generate(model)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Échantillonnage des résultats pendant l'entraînement\n",
|
||||
"\n",
|
||||
"Étant donné que nous ne disposons d'aucune métrique utile comme *l'exactitude*, la seule manière de vérifier que notre modèle s'améliore est de **prélever des exemples** de chaînes générées pendant l'entraînement. Pour ce faire, nous utiliserons des **callbacks**, c'est-à-dire des fonctions que nous pouvons passer à la fonction `fit`, et qui seront appelées périodiquement pendant l'entraînement.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 13,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Epoch 1/3\n",
|
||||
"15000/15000 [==============================] - 226s 15ms/step - loss: 1.2703\n",
|
||||
"Today #39;s a lead in the company for the strike\n",
|
||||
"Epoch 2/3\n",
|
||||
"15000/15000 [==============================] - 227s 15ms/step - loss: 1.2057\n",
|
||||
"Today #39;s the Market Service on Security Start (AP)\n",
|
||||
"Epoch 3/3\n",
|
||||
"15000/15000 [==============================] - 226s 15ms/step - loss: 1.1752\n",
|
||||
"Today #39;s a line on the strike to start for the start\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"<tensorflow.python.keras.callbacks.History at 0x7fa40c74e3d0>"
|
||||
]
|
||||
},
|
||||
"execution_count": 13,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"sampling_callback = keras.callbacks.LambdaCallback(\n",
|
||||
" on_epoch_end = lambda batch, logs: print(generate(model))\n",
|
||||
")\n",
|
||||
"\n",
|
||||
"model.fit(ds_train.batch(8).map(title_batch_fn),callbacks=[sampling_callback],epochs=3)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Cet exemple génère déjà un texte assez bon, mais il peut être amélioré de plusieurs façons :\n",
|
||||
"\n",
|
||||
"* **Plus de texte**. Nous avons uniquement utilisé des titres pour notre tâche, mais vous pourriez vouloir expérimenter avec du texte complet. Gardez à l'esprit que les RNN ne sont pas très performants pour gérer de longues séquences, il est donc judicieux soit de les diviser en phrases plus courtes, soit de toujours entraîner sur une longueur de séquence fixe d'une valeur prédéfinie `num_chars` (par exemple, 256). Vous pourriez essayer de modifier l'exemple ci-dessus pour adopter une telle architecture, en vous inspirant du [tutoriel officiel de Keras](https://keras.io/examples/generative/lstm_character_level_text_generation/).\n",
|
||||
"\n",
|
||||
"* **LSTM multicouche**. Il est pertinent d'essayer 2 ou 3 couches de cellules LSTM. Comme mentionné dans l'unité précédente, chaque couche de LSTM extrait certains motifs du texte, et dans le cas d'un générateur au niveau des caractères, on peut s'attendre à ce que le niveau inférieur du LSTM soit responsable de l'extraction des syllabes, et les niveaux supérieurs - des mots et des combinaisons de mots. Cela peut être simplement implémenté en passant un paramètre de nombre de couches au constructeur LSTM.\n",
|
||||
"\n",
|
||||
"* Vous pourriez également vouloir expérimenter avec **les unités GRU** pour voir lesquelles donnent de meilleurs résultats, ainsi qu'avec **différentes tailles de couches cachées**. Une couche cachée trop grande peut entraîner un surapprentissage (par exemple, le réseau apprendra le texte exact), tandis qu'une taille trop petite pourrait ne pas produire de bons résultats.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Génération de texte souple et température\n",
|
||||
"\n",
|
||||
"Dans la définition précédente de `generate`, nous choisissions toujours le caractère avec la probabilité la plus élevée comme prochain caractère dans le texte généré. Cela avait pour conséquence que le texte \"cyclait\" souvent entre les mêmes séquences de caractères encore et encore, comme dans cet exemple : \n",
|
||||
"```\n",
|
||||
"today of the second the company and a second the company ...\n",
|
||||
"```\n",
|
||||
"\n",
|
||||
"Cependant, si nous examinons la distribution de probabilité pour le prochain caractère, il se peut que la différence entre quelques probabilités les plus élevées ne soit pas énorme, par exemple, un caractère peut avoir une probabilité de 0,2, un autre de 0,19, etc. Par exemple, en cherchant le prochain caractère dans la séquence '*play*', le caractère suivant pourrait tout aussi bien être un espace ou un **e** (comme dans le mot *player*).\n",
|
||||
"\n",
|
||||
"Cela nous amène à la conclusion qu'il n'est pas toujours \"juste\" de sélectionner le caractère avec la probabilité la plus élevée, car choisir le deuxième plus probable pourrait également conduire à un texte cohérent. Il est plus judicieux de **prélever un échantillon** parmi les caractères en fonction de la distribution de probabilité donnée par la sortie du réseau.\n",
|
||||
"\n",
|
||||
"Ce prélèvement peut être effectué à l'aide de la fonction `np.multinomial`, qui implémente ce que l'on appelle la **distribution multinomiale**. Une fonction qui implémente cette génération de texte **souple** est définie ci-dessous :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 33,
|
||||
"metadata": {
|
||||
"scrolled": true
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"\n",
|
||||
"--- Temperature = 0.3\n",
|
||||
"Today #39;s strike #39; to start at the store return\n",
|
||||
"On Sunday PO to Be Data Profit Up (Reuters)\n",
|
||||
"Moscow, SP wins straight to the Microsoft #39;s control of the space start\n",
|
||||
"President olding of the blast start for the strike to pay <b>...</b>\n",
|
||||
"Little red riding hood ficed to the spam countered in European <b>...</b>\n",
|
||||
"\n",
|
||||
"--- Temperature = 0.8\n",
|
||||
"Today countie strikes ryder missile faces food market blut\n",
|
||||
"On Sunday collores lose-toppy of sale of Bullment in <b>...</b>\n",
|
||||
"Moscow, IBM Diffeiting in Afghan Software Hotels (Reuters)\n",
|
||||
"President Ol Luster for Profit Peaced Raised (AP)\n",
|
||||
"Little red riding hood dace on depart talks #39; bank up\n",
|
||||
"\n",
|
||||
"--- Temperature = 1.0\n",
|
||||
"Today wits House buiting debate fixes #39; supervice stake again\n",
|
||||
"On Sunday arling digital poaching In for level\n",
|
||||
"Moscow, DS Up 7, Top Proble Protest Caprey Mamarian Strike\n",
|
||||
"President teps help of roubler stepted lessabul-Dhalitics (AFP)\n",
|
||||
"Little red riding hood signs on cash in Carter-youb\n",
|
||||
"\n",
|
||||
"--- Temperature = 1.3\n",
|
||||
"Today wits flawer ro, pSIA figat's co DroftwavesIs Talo up\n",
|
||||
"On Sunday hround elitwing wint EU Powerburlinetien\n",
|
||||
"Moscow, Bazz #39;s sentries olymen winnelds' next for Olympite Huc?\n",
|
||||
"President lost securitys from power Elections in Smiltrials\n",
|
||||
"Little red riding hood vides profit, exponituity, profitmainalist-at said listers\n",
|
||||
"\n",
|
||||
"--- Temperature = 1.8\n",
|
||||
"Today #39;It: He deat: N.KA Asside\n",
|
||||
"On Sunday i arry Par aldeup patient Wo stele1\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"ename": "KeyError",
|
||||
"evalue": "0",
|
||||
"output_type": "error",
|
||||
"traceback": [
|
||||
"\u001b[0;31m---------------------------------------------------------------------------\u001b[0m",
|
||||
"\u001b[0;31mKeyError\u001b[0m Traceback (most recent call last)",
|
||||
"\u001b[0;32m<ipython-input-33-db32367a0feb>\u001b[0m in \u001b[0;36m<module>\u001b[0;34m\u001b[0m\n\u001b[1;32m 18\u001b[0m \u001b[0mprint\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0;34mf\"\\n--- Temperature = {i}\"\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[1;32m 19\u001b[0m \u001b[0;32mfor\u001b[0m \u001b[0mj\u001b[0m \u001b[0;32min\u001b[0m \u001b[0mrange\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0;36m5\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m:\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0;32m---> 20\u001b[0;31m \u001b[0mprint\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0mgenerate_soft\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0mmodel\u001b[0m\u001b[0;34m,\u001b[0m\u001b[0msize\u001b[0m\u001b[0;34m=\u001b[0m\u001b[0;36m300\u001b[0m\u001b[0;34m,\u001b[0m\u001b[0mstart\u001b[0m\u001b[0;34m=\u001b[0m\u001b[0mwords\u001b[0m\u001b[0;34m[\u001b[0m\u001b[0mj\u001b[0m\u001b[0;34m]\u001b[0m\u001b[0;34m,\u001b[0m\u001b[0mtemperature\u001b[0m\u001b[0;34m=\u001b[0m\u001b[0mi\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0m",
|
||||
"\u001b[0;32m<ipython-input-33-db32367a0feb>\u001b[0m in \u001b[0;36mgenerate_soft\u001b[0;34m(model, size, start, temperature)\u001b[0m\n\u001b[1;32m 11\u001b[0m \u001b[0mchars\u001b[0m\u001b[0;34m.\u001b[0m\u001b[0mappend\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0mnc\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[1;32m 12\u001b[0m \u001b[0minp\u001b[0m \u001b[0;34m=\u001b[0m \u001b[0minp\u001b[0m\u001b[0;34m+\u001b[0m\u001b[0;34m[\u001b[0m\u001b[0mnc\u001b[0m\u001b[0;34m]\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0;32m---> 13\u001b[0;31m \u001b[0;32mreturn\u001b[0m \u001b[0mdecode\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0mchars\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0m\u001b[1;32m 14\u001b[0m \u001b[0;34m\u001b[0m\u001b[0m\n\u001b[1;32m 15\u001b[0m \u001b[0mwords\u001b[0m \u001b[0;34m=\u001b[0m \u001b[0;34m[\u001b[0m\u001b[0;34m'Today '\u001b[0m\u001b[0;34m,\u001b[0m\u001b[0;34m'On Sunday '\u001b[0m\u001b[0;34m,\u001b[0m\u001b[0;34m'Moscow, '\u001b[0m\u001b[0;34m,\u001b[0m\u001b[0;34m'President '\u001b[0m\u001b[0;34m,\u001b[0m\u001b[0;34m'Little red riding hood '\u001b[0m\u001b[0;34m]\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n",
|
||||
"\u001b[0;32m<ipython-input-10-3f5fa6130b1d>\u001b[0m in \u001b[0;36mdecode\u001b[0;34m(x)\u001b[0m\n\u001b[1;32m 2\u001b[0m \u001b[0;34m\u001b[0m\u001b[0m\n\u001b[1;32m 3\u001b[0m \u001b[0;32mdef\u001b[0m \u001b[0mdecode\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0mx\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m:\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0;32m----> 4\u001b[0;31m \u001b[0;32mreturn\u001b[0m \u001b[0;34m''\u001b[0m\u001b[0;34m.\u001b[0m\u001b[0mjoin\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0;34m[\u001b[0m\u001b[0mreverse_map\u001b[0m\u001b[0;34m[\u001b[0m\u001b[0mt\u001b[0m\u001b[0;34m]\u001b[0m \u001b[0;32mfor\u001b[0m \u001b[0mt\u001b[0m \u001b[0;32min\u001b[0m \u001b[0mx\u001b[0m\u001b[0;34m]\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0m",
|
||||
"\u001b[0;32m<ipython-input-10-3f5fa6130b1d>\u001b[0m in \u001b[0;36m<listcomp>\u001b[0;34m(.0)\u001b[0m\n\u001b[1;32m 2\u001b[0m \u001b[0;34m\u001b[0m\u001b[0m\n\u001b[1;32m 3\u001b[0m \u001b[0;32mdef\u001b[0m \u001b[0mdecode\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0mx\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m:\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0;32m----> 4\u001b[0;31m \u001b[0;32mreturn\u001b[0m \u001b[0;34m''\u001b[0m\u001b[0;34m.\u001b[0m\u001b[0mjoin\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0;34m[\u001b[0m\u001b[0mreverse_map\u001b[0m\u001b[0;34m[\u001b[0m\u001b[0mt\u001b[0m\u001b[0;34m]\u001b[0m \u001b[0;32mfor\u001b[0m \u001b[0mt\u001b[0m \u001b[0;32min\u001b[0m \u001b[0mx\u001b[0m\u001b[0;34m]\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0m",
|
||||
"\u001b[0;31mKeyError\u001b[0m: 0"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"def generate_soft(model,size=100,start='Today ',temperature=1.0):\n",
|
||||
" inp = tokenizer.texts_to_sequences([start])[0]\n",
|
||||
" chars = inp\n",
|
||||
" for i in range(size):\n",
|
||||
" out = model(tf.expand_dims(tf.one_hot(inp,vocab_size),0))[0][-1]\n",
|
||||
" probs = tf.exp(tf.math.log(out)/temperature).numpy().astype(np.float64)\n",
|
||||
" probs = probs/np.sum(probs)\n",
|
||||
" nc = np.argmax(np.random.multinomial(1,probs,1))\n",
|
||||
" if nc==eos_token:\n",
|
||||
" break\n",
|
||||
" chars.append(nc)\n",
|
||||
" inp = inp+[nc]\n",
|
||||
" return decode(chars)\n",
|
||||
"\n",
|
||||
"words = ['Today ','On Sunday ','Moscow, ','President ','Little red riding hood ']\n",
|
||||
" \n",
|
||||
"for i in [0.3,0.8,1.0,1.3,1.8]:\n",
|
||||
" print(f\"\\n--- Temperature = {i}\")\n",
|
||||
" for j in range(5):\n",
|
||||
" print(generate_soft(model,size=300,start=words[j],temperature=i))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Nous avons introduit un paramètre supplémentaire appelé **température**, qui est utilisé pour indiquer à quel point nous devons nous en tenir à la probabilité la plus élevée. Si la température est de 1,0, nous effectuons un échantillonnage multinomial équitable, et lorsque la température tend vers l'infini - toutes les probabilités deviennent égales, et nous sélectionnons aléatoirement le prochain caractère. Dans l'exemple ci-dessous, nous pouvons observer que le texte devient dénué de sens lorsque nous augmentons trop la température, et il ressemble à un texte \"cyclé\" généré de manière rigide lorsqu'il se rapproche de 0.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"\n---\n\n**Avertissement** : \nCe document a été traduit à l'aide du service de traduction automatique [Co-op Translator](https://github.com/Azure/co-op-translator). Bien que nous nous efforcions d'assurer l'exactitude, veuillez noter que les traductions automatisées peuvent contenir des erreurs ou des inexactitudes. Le document original dans sa langue d'origine doit être considéré comme la source faisant autorité. Pour des informations critiques, il est recommandé de recourir à une traduction professionnelle effectuée par un humain. Nous déclinons toute responsabilité en cas de malentendus ou d'interprétations erronées résultant de l'utilisation de cette traduction.\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
"interpreter": {
|
||||
"hash": "16af2a8bbb083ea23e5e41c7f5787656b2ce26968575d8763f2c4b17f9cd711f"
|
||||
},
|
||||
"kernelspec": {
|
||||
"display_name": "Python 3.8.12 ('py38')",
|
||||
"language": "python",
|
||||
"name": "python3"
|
||||
},
|
||||
"language_info": {
|
||||
"codemirror_mode": {
|
||||
"name": "ipython",
|
||||
"version": 3
|
||||
},
|
||||
"file_extension": ".py",
|
||||
"mimetype": "text/x-python",
|
||||
"name": "python",
|
||||
"nbconvert_exporter": "python",
|
||||
"pygments_lexer": "ipython3",
|
||||
"version": "3.8.12"
|
||||
},
|
||||
"coopTranslator": {
|
||||
"original_hash": "9fbb7d5fda708537649f71f5f646fcde",
|
||||
"translation_date": "2025-08-31T15:11:26+00:00",
|
||||
"source_file": "lessons/5-NLP/17-GenerativeNetworks/GenerativeTF.ipynb",
|
||||
"language_code": "fr"
|
||||
}
|
||||
},
|
||||
"nbformat": 4,
|
||||
"nbformat_minor": 4
|
||||
}
|
||||
|
|
@ -0,0 +1,353 @@
|
|||
{
|
||||
"cells": [
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"# Mécanismes d'attention et transformateurs\n",
|
||||
"\n",
|
||||
"Un inconvénient majeur des réseaux récurrents est que tous les mots d'une séquence ont le même impact sur le résultat. Cela entraîne des performances sous-optimales avec les modèles standard encodeur-décodeur LSTM pour les tâches de séquence à séquence, telles que la reconnaissance d'entités nommées et la traduction automatique. En réalité, certains mots spécifiques de la séquence d'entrée ont souvent plus d'impact sur les sorties séquentielles que d'autres.\n",
|
||||
"\n",
|
||||
"Prenons un modèle de séquence à séquence, comme la traduction automatique. Il est implémenté par deux réseaux récurrents, où un réseau (**encodeur**) compresse la séquence d'entrée dans un état caché, et un autre (**décodeur**) déploie cet état caché pour produire le résultat traduit. Le problème avec cette approche est que l'état final du réseau a du mal à se souvenir du début de la phrase, ce qui entraîne une qualité médiocre du modèle pour les phrases longues.\n",
|
||||
"\n",
|
||||
"Les **mécanismes d'attention** offrent un moyen de pondérer l'impact contextuel de chaque vecteur d'entrée sur chaque prédiction de sortie du RNN. Cela est réalisé en créant des raccourcis entre les états intermédiaires du RNN d'entrée et du RNN de sortie. Ainsi, lors de la génération du symbole de sortie $y_t$, nous prenons en compte tous les états cachés d'entrée $h_i$, avec différents coefficients de pondération $\\alpha_{t,i}$.\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"*Le modèle encodeur-décodeur avec mécanisme d'attention additive dans [Bahdanau et al., 2015](https://arxiv.org/pdf/1409.0473.pdf), cité de [ce billet de blog](https://lilianweng.github.io/lil-log/2018/06/24/attention-attention.html)*\n",
|
||||
"\n",
|
||||
"La matrice d'attention $\\{\\alpha_{i,j}\\}$ représente le degré auquel certains mots d'entrée influencent la génération d'un mot donné dans la séquence de sortie. Voici un exemple de cette matrice :\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"*Figure tirée de [Bahdanau et al., 2015](https://arxiv.org/pdf/1409.0473.pdf) (Fig.3)*\n",
|
||||
"\n",
|
||||
"Les mécanismes d'attention sont responsables de la plupart des avancées actuelles ou proches de l'état de l'art en traitement du langage naturel. Cependant, l'ajout d'attention augmente considérablement le nombre de paramètres du modèle, ce qui a entraîné des problèmes de mise à l'échelle avec les RNN. Une contrainte clé pour la mise à l'échelle des RNN est que la nature récurrente des modèles rend difficile le traitement par lots et la parallélisation de l'entraînement. Dans un RNN, chaque élément d'une séquence doit être traité dans un ordre séquentiel, ce qui signifie qu'il ne peut pas être facilement parallélisé.\n",
|
||||
"\n",
|
||||
"L'adoption des mécanismes d'attention combinée à cette contrainte a conduit à la création des modèles transformateurs, désormais à l'état de l'art, que nous connaissons et utilisons aujourd'hui, de BERT à OpenGPT3.\n",
|
||||
"\n",
|
||||
"## Modèles transformateurs\n",
|
||||
"\n",
|
||||
"Au lieu de transmettre le contexte de chaque prédiction précédente à l'étape d'évaluation suivante, les **modèles transformateurs** utilisent des **encodages positionnels** et l'attention pour capturer le contexte d'une entrée donnée dans une fenêtre de texte fournie. L'image ci-dessous montre comment les encodages positionnels avec attention peuvent capturer le contexte dans une fenêtre donnée.\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"Étant donné que chaque position d'entrée est mappée indépendamment à chaque position de sortie, les transformateurs peuvent mieux paralléliser que les RNN, ce qui permet des modèles de langage beaucoup plus grands et plus expressifs. Chaque tête d'attention peut être utilisée pour apprendre différentes relations entre les mots, ce qui améliore les tâches de traitement du langage naturel en aval.\n",
|
||||
"\n",
|
||||
"**BERT** (Bidirectional Encoder Representations from Transformers) est un réseau transformateur multi-couches très large avec 12 couches pour *BERT-base* et 24 pour *BERT-large*. Le modèle est d'abord pré-entraîné sur un corpus de texte volumineux (WikiPedia + livres) en utilisant un entraînement non supervisé (prédiction des mots masqués dans une phrase). Pendant le pré-entraînement, le modèle acquiert un niveau significatif de compréhension du langage qui peut ensuite être exploité avec d'autres ensembles de données via un ajustement fin. Ce processus est appelé **apprentissage par transfert**.\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"Il existe de nombreuses variantes des architectures de transformateurs, notamment BERT, DistilBERT, BigBird, OpenGPT3 et bien d'autres, qui peuvent être ajustées. Le [package HuggingFace](https://github.com/huggingface/) fournit un dépôt pour entraîner plusieurs de ces architectures avec PyTorch.\n",
|
||||
"\n",
|
||||
"## Utilisation de BERT pour la classification de texte\n",
|
||||
"\n",
|
||||
"Voyons comment utiliser un modèle BERT pré-entraîné pour résoudre notre tâche traditionnelle : la classification de séquences. Nous allons classifier notre ensemble de données AG News original.\n",
|
||||
"\n",
|
||||
"Tout d'abord, chargeons la bibliothèque HuggingFace et notre ensemble de données :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 10,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Loading dataset...\n",
|
||||
"Building vocab...\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"import torch\n",
|
||||
"import torchtext\n",
|
||||
"from torchnlp import *\n",
|
||||
"import transformers\n",
|
||||
"train_dataset, test_dataset, classes, vocab = load_dataset()\n",
|
||||
"vocab_len = len(vocab)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Parce que nous allons utiliser un modèle BERT pré-entraîné, nous devrons utiliser un tokenizer spécifique. Tout d'abord, nous allons charger un tokenizer associé au modèle BERT pré-entraîné.\n",
|
||||
"\n",
|
||||
"La bibliothèque HuggingFace contient un dépôt de modèles pré-entraînés, que vous pouvez utiliser simplement en spécifiant leurs noms comme arguments dans les fonctions `from_pretrained`. Tous les fichiers binaires nécessaires pour le modèle seront automatiquement téléchargés.\n",
|
||||
"\n",
|
||||
"Cependant, dans certains cas, vous devrez charger vos propres modèles. Dans ce cas, vous pouvez spécifier le répertoire contenant tous les fichiers pertinents, y compris les paramètres pour le tokenizer, le fichier `config.json` avec les paramètres du modèle, les poids binaires, etc.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 11,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"# To load the model from Internet repository using model name. \n",
|
||||
"# Use this if you are running from your own copy of the notebooks\n",
|
||||
"bert_model = 'bert-base-uncased' \n",
|
||||
"\n",
|
||||
"# To load the model from the directory on disk. Use this for Microsoft Learn module, because we have\n",
|
||||
"# prepared all required files for you.\n",
|
||||
"bert_model = './bert'\n",
|
||||
"\n",
|
||||
"tokenizer = transformers.BertTokenizer.from_pretrained(bert_model)\n",
|
||||
"\n",
|
||||
"MAX_SEQ_LEN = 128\n",
|
||||
"PAD_INDEX = tokenizer.convert_tokens_to_ids(tokenizer.pad_token)\n",
|
||||
"UNK_INDEX = tokenizer.convert_tokens_to_ids(tokenizer.unk_token)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"L'objet `tokenizer` contient la fonction `encode` qui peut être utilisée directement pour encoder du texte :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 15,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"[101, 1052, 22123, 2953, 2818, 2003, 1037, 2307, 7705, 2005, 17953, 2361, 102]"
|
||||
]
|
||||
},
|
||||
"execution_count": 15,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"tokenizer.encode('PyTorch is a great framework for NLP')"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Ensuite, créons des itérateurs que nous utiliserons pendant l'entraînement pour accéder aux données. Étant donné que BERT utilise sa propre fonction d'encodage, nous devrons définir une fonction de remplissage similaire à `padify` que nous avons définie auparavant :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"def pad_bert(b):\n",
|
||||
" # b is the list of tuples of length batch_size\n",
|
||||
" # - first element of a tuple = label, \n",
|
||||
" # - second = feature (text sequence)\n",
|
||||
" # build vectorized sequence\n",
|
||||
" v = [tokenizer.encode(x[1]) for x in b]\n",
|
||||
" # compute max length of a sequence in this minibatch\n",
|
||||
" l = max(map(len,v))\n",
|
||||
" return ( # tuple of two tensors - labels and features\n",
|
||||
" torch.LongTensor([t[0] for t in b]),\n",
|
||||
" torch.stack([torch.nn.functional.pad(torch.tensor(t),(0,l-len(t)),mode='constant',value=0) for t in v])\n",
|
||||
" )\n",
|
||||
"\n",
|
||||
"train_loader = torch.utils.data.DataLoader(train_dataset, batch_size=8, collate_fn=pad_bert, shuffle=True)\n",
|
||||
"test_loader = torch.utils.data.DataLoader(test_dataset, batch_size=8, collate_fn=pad_bert)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Dans notre cas, nous utiliserons un modèle BERT pré-entraîné appelé `bert-base-uncased`. Chargeons le modèle en utilisant le package `BertForSequenceClassification`. Cela garantit que notre modèle dispose déjà de l'architecture requise pour la classification, y compris le classificateur final. Vous verrez un message d'avertissement indiquant que les poids du classificateur final ne sont pas initialisés, et que le modèle nécessiterait un pré-entraînement - c'est tout à fait normal, car c'est exactement ce que nous allons faire !\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 9,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stderr",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Some weights of the model checkpoint at ./bert were not used when initializing BertForSequenceClassification: ['cls.predictions.bias', 'cls.predictions.transform.dense.weight', 'cls.predictions.transform.dense.bias', 'cls.predictions.decoder.weight', 'cls.seq_relationship.weight', 'cls.seq_relationship.bias', 'cls.predictions.transform.LayerNorm.weight', 'cls.predictions.transform.LayerNorm.bias']\n",
|
||||
"- This IS expected if you are initializing BertForSequenceClassification from the checkpoint of a model trained on another task or with another architecture (e.g. initializing a BertForSequenceClassification model from a BertForPreTraining model).\n",
|
||||
"- This IS NOT expected if you are initializing BertForSequenceClassification from the checkpoint of a model that you expect to be exactly identical (initializing a BertForSequenceClassification model from a BertForSequenceClassification model).\n",
|
||||
"Some weights of BertForSequenceClassification were not initialized from the model checkpoint at ./bert and are newly initialized: ['classifier.weight', 'classifier.bias']\n",
|
||||
"You should probably TRAIN this model on a down-stream task to be able to use it for predictions and inference.\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"model = transformers.BertForSequenceClassification.from_pretrained(bert_model,num_labels=4).to(device)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Nous sommes maintenant prêts à commencer l'entraînement ! Comme BERT est déjà pré-entraîné, nous souhaitons utiliser un taux d'apprentissage relativement faible afin de ne pas altérer les poids initiaux.\n",
|
||||
"\n",
|
||||
"Tout le travail important est effectué par le modèle `BertForSequenceClassification`. Lorsque nous appelons le modèle sur les données d'entraînement, il renvoie à la fois la perte (loss) et la sortie du réseau pour le minibatch d'entrée. Nous utilisons la perte pour l'optimisation des paramètres (`loss.backward()` effectue la rétropropagation), et `out` pour calculer la précision de l'entraînement en comparant les étiquettes obtenues `labs` (calculées avec `argmax`) avec les étiquettes attendues `labels`.\n",
|
||||
"\n",
|
||||
"Pour contrôler le processus, nous accumulons la perte et la précision sur plusieurs itérations, et nous les affichons tous les `report_freq` cycles d'entraînement.\n",
|
||||
"\n",
|
||||
"Cet entraînement prendra probablement beaucoup de temps, donc nous limitons le nombre d'itérations.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 6,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Loss = 1.1254194641113282, Accuracy = 0.585\n",
|
||||
"Loss = 0.6194715118408203, Accuracy = 0.83\n",
|
||||
"Loss = 0.46665248870849607, Accuracy = 0.8475\n",
|
||||
"Loss = 0.4309701919555664, Accuracy = 0.8575\n",
|
||||
"Loss = 0.35427074432373046, Accuracy = 0.8825\n",
|
||||
"Loss = 0.3306886291503906, Accuracy = 0.8975\n",
|
||||
"Loss = 0.30340143203735354, Accuracy = 0.8975\n",
|
||||
"Loss = 0.26139299392700194, Accuracy = 0.915\n",
|
||||
"Loss = 0.26708646774291994, Accuracy = 0.9225\n",
|
||||
"Loss = 0.3667240524291992, Accuracy = 0.8675\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"optimizer = torch.optim.Adam(model.parameters(), lr=2e-5)\n",
|
||||
"\n",
|
||||
"report_freq = 50\n",
|
||||
"iterations = 500 # make this larger to train for longer time!\n",
|
||||
"\n",
|
||||
"model.train()\n",
|
||||
"\n",
|
||||
"i,c = 0,0\n",
|
||||
"acc_loss = 0\n",
|
||||
"acc_acc = 0\n",
|
||||
"\n",
|
||||
"for labels,texts in train_loader:\n",
|
||||
" labels = labels.to(device)-1 # get labels in the range 0-3 \n",
|
||||
" texts = texts.to(device)\n",
|
||||
" loss, out = model(texts, labels=labels)[:2]\n",
|
||||
" labs = out.argmax(dim=1)\n",
|
||||
" acc = torch.mean((labs==labels).type(torch.float32))\n",
|
||||
" optimizer.zero_grad()\n",
|
||||
" loss.backward()\n",
|
||||
" optimizer.step()\n",
|
||||
" acc_loss += loss\n",
|
||||
" acc_acc += acc\n",
|
||||
" i+=1\n",
|
||||
" c+=1\n",
|
||||
" if i%report_freq==0:\n",
|
||||
" print(f\"Loss = {acc_loss.item()/c}, Accuracy = {acc_acc.item()/c}\")\n",
|
||||
" c = 0\n",
|
||||
" acc_loss = 0\n",
|
||||
" acc_acc = 0\n",
|
||||
" iterations-=1\n",
|
||||
" if not iterations:\n",
|
||||
" break"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Vous pouvez constater (surtout si vous augmentez le nombre d'itérations et attendez suffisamment longtemps) que la classification avec BERT nous donne une précision assez bonne ! Cela s'explique par le fait que BERT comprend déjà très bien la structure de la langue, et que nous n'avons qu'à ajuster le classificateur final. Cependant, comme BERT est un modèle volumineux, tout le processus d'entraînement prend beaucoup de temps et nécessite une puissance de calcul importante ! (GPU, et de préférence plus d'un).\n",
|
||||
"\n",
|
||||
"> **Note :** Dans notre exemple, nous utilisons l'un des plus petits modèles BERT pré-entraînés. Il existe des modèles plus grands qui sont susceptibles de donner de meilleurs résultats.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Évaluation des performances du modèle\n",
|
||||
"\n",
|
||||
"Nous pouvons maintenant évaluer les performances de notre modèle sur le jeu de données de test. La boucle d'évaluation est assez similaire à la boucle d'entraînement, mais il ne faut pas oublier de passer le modèle en mode évaluation en appelant `model.eval()`.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 10,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Final accuracy: 0.9047029702970297\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"model.eval()\n",
|
||||
"iterations = 100\n",
|
||||
"acc = 0\n",
|
||||
"i = 0\n",
|
||||
"for labels,texts in test_loader:\n",
|
||||
" labels = labels.to(device)-1 \n",
|
||||
" texts = texts.to(device)\n",
|
||||
" _, out = model(texts, labels=labels)[:2]\n",
|
||||
" labs = out.argmax(dim=1)\n",
|
||||
" acc += torch.mean((labs==labels).type(torch.float32))\n",
|
||||
" i+=1\n",
|
||||
" if i>iterations: break\n",
|
||||
" \n",
|
||||
"print(f\"Final accuracy: {acc.item()/i}\")"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## À retenir\n",
|
||||
"\n",
|
||||
"Dans cette unité, nous avons vu à quel point il est facile de prendre un modèle de langage pré-entraîné de la bibliothèque **transformers** et de l'adapter à notre tâche de classification de texte. De la même manière, les modèles BERT peuvent être utilisés pour l'extraction d'entités, les questions-réponses et d'autres tâches de NLP.\n",
|
||||
"\n",
|
||||
"Les modèles de type Transformer représentent l'état de l'art actuel en NLP, et dans la plupart des cas, ils devraient être la première solution avec laquelle vous commencez à expérimenter lorsque vous mettez en œuvre des solutions NLP personnalisées. Cependant, comprendre les principes de base des réseaux neuronaux récurrents discutés dans ce module est extrêmement important si vous souhaitez construire des modèles neuronaux avancés.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"\n---\n\n**Avertissement** : \nCe document a été traduit à l'aide du service de traduction automatique [Co-op Translator](https://github.com/Azure/co-op-translator). Bien que nous nous efforcions d'assurer l'exactitude, veuillez noter que les traductions automatisées peuvent contenir des erreurs ou des inexactitudes. Le document original dans sa langue d'origine doit être considéré comme la source faisant autorité. Pour des informations critiques, il est recommandé de recourir à une traduction professionnelle réalisée par un humain. Nous déclinons toute responsabilité en cas de malentendus ou d'interprétations erronées résultant de l'utilisation de cette traduction.\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
"kernelspec": {
|
||||
"display_name": "py37_pytorch",
|
||||
"language": "python",
|
||||
"name": "conda-env-py37_pytorch-py"
|
||||
},
|
||||
"language_info": {
|
||||
"codemirror_mode": {
|
||||
"name": "ipython",
|
||||
"version": 3
|
||||
},
|
||||
"file_extension": ".py",
|
||||
"mimetype": "text/x-python",
|
||||
"name": "python",
|
||||
"nbconvert_exporter": "python",
|
||||
"pygments_lexer": "ipython3",
|
||||
"version": "3.7.7"
|
||||
},
|
||||
"coopTranslator": {
|
||||
"original_hash": "753865967678a92dbce7d7efbd36d980",
|
||||
"translation_date": "2025-08-31T15:16:28+00:00",
|
||||
"source_file": "lessons/5-NLP/18-Transformers/TransformersPyTorch.ipynb",
|
||||
"language_code": "fr"
|
||||
}
|
||||
},
|
||||
"nbformat": 4,
|
||||
"nbformat_minor": 4
|
||||
}
|
||||
|
|
@ -0,0 +1,819 @@
|
|||
{
|
||||
"cells": [
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"# Mécanismes d'attention et transformateurs\n",
|
||||
"\n",
|
||||
"Un inconvénient majeur des réseaux récurrents est que tous les mots d'une séquence ont le même impact sur le résultat. Cela entraîne des performances sous-optimales avec les modèles standard encodeur-décodeur LSTM pour les tâches de séquence à séquence, telles que la reconnaissance d'entités nommées et la traduction automatique. En réalité, certains mots spécifiques de la séquence d'entrée ont souvent plus d'impact sur les sorties séquentielles que d'autres.\n",
|
||||
"\n",
|
||||
"Prenons un modèle de séquence à séquence, comme la traduction automatique. Il est implémenté par deux réseaux récurrents, où un réseau (**encodeur**) compresse la séquence d'entrée dans un état caché, et un autre (**décodeur**) déploie cet état caché pour produire le résultat traduit. Le problème avec cette approche est que l'état final du réseau a du mal à se souvenir du début de la phrase, ce qui entraîne une mauvaise qualité du modèle pour les phrases longues.\n",
|
||||
"\n",
|
||||
"Les **mécanismes d'attention** offrent un moyen de pondérer l'impact contextuel de chaque vecteur d'entrée sur chaque prédiction de sortie du RNN. Cela est réalisé en créant des raccourcis entre les états intermédiaires du RNN d'entrée et du RNN de sortie. Ainsi, lors de la génération du symbole de sortie $y_t$, nous prenons en compte tous les états cachés d'entrée $h_i$, avec différents coefficients de pondération $\\alpha_{t,i}$.\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"*Le modèle encodeur-décodeur avec mécanisme d'attention additive dans [Bahdanau et al., 2015](https://arxiv.org/pdf/1409.0473.pdf), cité de [cet article de blog](https://lilianweng.github.io/lil-log/2018/06/24/attention-attention.html)*\n",
|
||||
"\n",
|
||||
"La matrice d'attention $\\{\\alpha_{i,j}\\}$ représente le degré auquel certains mots d'entrée influencent la génération d'un mot donné dans la séquence de sortie. Voici un exemple de cette matrice :\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"*Figure tirée de [Bahdanau et al., 2015](https://arxiv.org/pdf/1409.0473.pdf) (Fig.3)*\n",
|
||||
"\n",
|
||||
"Les mécanismes d'attention sont responsables de la plupart des avancées actuelles ou proches de l'état de l'art en traitement du langage naturel. Cependant, l'ajout d'attention augmente considérablement le nombre de paramètres du modèle, ce qui a entraîné des problèmes de mise à l'échelle avec les RNN. Une contrainte clé pour la mise à l'échelle des RNN est que la nature récurrente des modèles rend difficile le traitement par lots et la parallélisation de l'entraînement. Dans un RNN, chaque élément d'une séquence doit être traité dans un ordre séquentiel, ce qui signifie qu'il ne peut pas être facilement parallélisé.\n",
|
||||
"\n",
|
||||
"L'adoption des mécanismes d'attention combinée à cette contrainte a conduit à la création des modèles transformateurs, désormais à l'état de l'art, que nous connaissons et utilisons aujourd'hui, de BERT à OpenGPT3.\n",
|
||||
"\n",
|
||||
"## Modèles transformateurs\n",
|
||||
"\n",
|
||||
"Au lieu de transmettre le contexte de chaque prédiction précédente à l'étape d'évaluation suivante, les **modèles transformateurs** utilisent des **encodages positionnels** et **l'attention** pour capturer le contexte d'une entrée donnée dans une fenêtre de texte fournie. L'image ci-dessous montre comment les encodages positionnels avec attention peuvent capturer le contexte dans une fenêtre donnée.\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"Étant donné que chaque position d'entrée est mappée indépendamment à chaque position de sortie, les transformateurs peuvent mieux paralléliser que les RNN, ce qui permet des modèles de langage beaucoup plus grands et plus expressifs. Chaque tête d'attention peut être utilisée pour apprendre différentes relations entre les mots, ce qui améliore les tâches de traitement du langage naturel en aval.\n",
|
||||
"\n",
|
||||
"## Construire un modèle transformateur simple\n",
|
||||
"\n",
|
||||
"Keras ne contient pas de couche transformateur intégrée, mais nous pouvons en construire une nous-mêmes. Comme précédemment, nous nous concentrerons sur la classification de texte du jeu de données AG News, mais il convient de mentionner que les modèles transformateurs donnent les meilleurs résultats pour des tâches NLP plus complexes.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 1,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"import tensorflow as tf\n",
|
||||
"from tensorflow import keras\n",
|
||||
"import tensorflow_datasets as tfds\n",
|
||||
"import numpy as np\n",
|
||||
"\n",
|
||||
"ds_train, ds_test = tfds.load('ag_news_subset').values()\n",
|
||||
"\n",
|
||||
"def extract_text(x):\n",
|
||||
" return x['title']+' '+x['description']\n",
|
||||
"\n",
|
||||
"def tupelize(x):\n",
|
||||
" return (extract_text(x),x['label'])"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Les nouvelles couches dans Keras doivent hériter de la classe `Layer` et implémenter la méthode `call`. Commençons par la couche **Positional Embedding**. Nous utiliserons [du code provenant de la documentation officielle de Keras](https://keras.io/examples/nlp/text_classification_with_transformer/). Nous supposerons que nous remplissons toutes les séquences d'entrée à une longueur `maxlen`.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 2,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"class TokenAndPositionEmbedding(keras.layers.Layer):\n",
|
||||
" def __init__(self, maxlen, vocab_size, embed_dim):\n",
|
||||
" super(TokenAndPositionEmbedding, self).__init__()\n",
|
||||
" self.token_emb = keras.layers.Embedding(input_dim=vocab_size, output_dim=embed_dim)\n",
|
||||
" self.pos_emb = keras.layers.Embedding(input_dim=maxlen, output_dim=embed_dim)\n",
|
||||
" self.maxlen = maxlen\n",
|
||||
"\n",
|
||||
" def call(self, x):\n",
|
||||
" maxlen = self.maxlen\n",
|
||||
" positions = tf.range(start=0, limit=maxlen, delta=1)\n",
|
||||
" positions = self.pos_emb(positions)\n",
|
||||
" x = self.token_emb(x)\n",
|
||||
" return x+positions"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Cette couche se compose de deux couches `Embedding` : une pour l'intégration des tokens (comme nous l'avons déjà abordé) et une pour les positions des tokens. Les positions des tokens sont générées comme une séquence de nombres naturels allant de 0 à `maxlen` à l'aide de `tf.range`, puis passées à travers la couche d'intégration. Les deux vecteurs d'intégration résultants sont ensuite additionnés, produisant une représentation intégrée positionnelle de l'entrée de forme `maxlen`$\\times$`embed_dim`.\n",
|
||||
"\n",
|
||||
"Passons maintenant à l'implémentation du bloc transformateur. Il prendra en entrée le résultat de la couche d'intégration définie précédemment :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 3,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"class TransformerBlock(keras.layers.Layer):\n",
|
||||
" def __init__(self, embed_dim, num_heads, ff_dim, rate=0.1):\n",
|
||||
" super(TransformerBlock, self).__init__()\n",
|
||||
" self.att = keras.layers.MultiHeadAttention(num_heads=num_heads, key_dim=embed_dim, name='attn')\n",
|
||||
" self.ffn = keras.Sequential(\n",
|
||||
" [keras.layers.Dense(ff_dim, activation=\"relu\"), keras.layers.Dense(embed_dim),]\n",
|
||||
" )\n",
|
||||
" self.layernorm1 = keras.layers.LayerNormalization(epsilon=1e-6)\n",
|
||||
" self.layernorm2 = keras.layers.LayerNormalization(epsilon=1e-6)\n",
|
||||
" self.dropout1 = keras.layers.Dropout(rate)\n",
|
||||
" self.dropout2 = keras.layers.Dropout(rate)\n",
|
||||
"\n",
|
||||
" def call(self, inputs, training):\n",
|
||||
" attn_output = self.att(inputs, inputs)\n",
|
||||
" attn_output = self.dropout1(attn_output, training=training)\n",
|
||||
" out1 = self.layernorm1(inputs + attn_output)\n",
|
||||
" ffn_output = self.ffn(out1)\n",
|
||||
" ffn_output = self.dropout2(ffn_output, training=training)\n",
|
||||
" return self.layernorm2(out1 + ffn_output)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Maintenant, nous sommes prêts à définir le modèle complet du transformeur :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Model: \"sequential_1\"\n",
|
||||
"_________________________________________________________________\n",
|
||||
"Layer (type) Output Shape Param # \n",
|
||||
"=================================================================\n",
|
||||
"text_vectorization (TextVect (None, 256) 0 \n",
|
||||
"_________________________________________________________________\n",
|
||||
"token_and_position_embedding (None, 256, 32) 648192 \n",
|
||||
"_________________________________________________________________\n",
|
||||
"transformer_block (Transform (None, 256, 32) 10656 \n",
|
||||
"_________________________________________________________________\n",
|
||||
"global_average_pooling1d (Gl (None, 32) 0 \n",
|
||||
"_________________________________________________________________\n",
|
||||
"dropout_2 (Dropout) (None, 32) 0 \n",
|
||||
"_________________________________________________________________\n",
|
||||
"dense_2 (Dense) (None, 20) 660 \n",
|
||||
"_________________________________________________________________\n",
|
||||
"dropout_3 (Dropout) (None, 20) 0 \n",
|
||||
"_________________________________________________________________\n",
|
||||
"dense_3 (Dense) (None, 4) 84 \n",
|
||||
"=================================================================\n",
|
||||
"Total params: 659,592\n",
|
||||
"Trainable params: 659,592\n",
|
||||
"Non-trainable params: 0\n",
|
||||
"_________________________________________________________________\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"embed_dim = 32 # Embedding size for each token\n",
|
||||
"num_heads = 2 # Number of attention heads\n",
|
||||
"ff_dim = 32 # Hidden layer size in feed forward network inside transformer\n",
|
||||
"maxlen = 256\n",
|
||||
"vocab_size = 20000\n",
|
||||
"\n",
|
||||
"model = keras.models.Sequential([\n",
|
||||
" keras.layers.experimental.preprocessing.TextVectorization(max_tokens=vocab_size,output_sequence_length=maxlen, input_shape=(1,)),\n",
|
||||
" TokenAndPositionEmbedding(maxlen, vocab_size, embed_dim),\n",
|
||||
" TransformerBlock(embed_dim, num_heads, ff_dim),\n",
|
||||
" keras.layers.GlobalAveragePooling1D(),\n",
|
||||
" keras.layers.Dropout(0.1),\n",
|
||||
" keras.layers.Dense(20, activation=\"relu\"),\n",
|
||||
" keras.layers.Dropout(0.1),\n",
|
||||
" keras.layers.Dense(4, activation=\"softmax\")\n",
|
||||
"])\n",
|
||||
"\n",
|
||||
"model.summary()"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 5,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Training tokenizer\n",
|
||||
"938/938 [==============================] - 45s 39ms/step - loss: 0.4978 - acc: 0.8068 - val_loss: 0.2808 - val_acc: 0.9124\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"<tensorflow.python.keras.callbacks.History at 0x7f9c2427a0d0>"
|
||||
]
|
||||
},
|
||||
"execution_count": 5,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"print('Training tokenizer')\n",
|
||||
"model.layers[0].adapt(ds_train.map(extract_text))\n",
|
||||
"model.compile(loss='sparse_categorical_crossentropy',metrics=['acc'], optimizer='adam')\n",
|
||||
"model.fit(ds_train.map(tupelize).batch(128),validation_data=ds_test.map(tupelize).batch(128))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Modèles Transformers BERT\n",
|
||||
"\n",
|
||||
"**BERT** (Bidirectional Encoder Representations from Transformers) est un réseau de transformateurs multi-couches très large, avec 12 couches pour *BERT-base* et 24 pour *BERT-large*. Le modèle est d'abord pré-entraîné sur un vaste corpus de données textuelles (WikiPedia + livres) en utilisant un apprentissage non supervisé (prédiction des mots masqués dans une phrase). Pendant cette phase de pré-entraînement, le modèle acquiert un niveau significatif de compréhension du langage, qui peut ensuite être exploité avec d'autres ensembles de données grâce à un ajustement fin. Ce processus est appelé **apprentissage par transfert**.\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"Il existe de nombreuses variantes des architectures Transformer, notamment BERT, DistilBERT, BigBird, OpenGPT3 et bien d'autres, qui peuvent être ajustées.\n",
|
||||
"\n",
|
||||
"Voyons comment nous pouvons utiliser un modèle BERT pré-entraîné pour résoudre notre problème classique de classification de séquences. Nous allons emprunter l'idée et une partie du code de la [documentation officielle](https://www.tensorflow.org/text/tutorials/classify_text_with_bert).\n",
|
||||
"\n",
|
||||
"Pour charger des modèles pré-entraînés, nous utiliserons **Tensorflow hub**. Tout d'abord, chargeons le vectoriseur spécifique à BERT :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 1,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"ename": "ModuleNotFoundError",
|
||||
"evalue": "No module named 'tensorflow_text'",
|
||||
"output_type": "error",
|
||||
"traceback": [
|
||||
"\u001b[1;31m---------------------------------------------------------------------------\u001b[0m",
|
||||
"\u001b[1;31mModuleNotFoundError\u001b[0m Traceback (most recent call last)",
|
||||
"\u001b[1;32m~\\AppData\\Local\\Temp/ipykernel_41180/4216669875.py\u001b[0m in \u001b[0;36m<module>\u001b[1;34m\u001b[0m\n\u001b[1;32m----> 1\u001b[1;33m \u001b[1;32mimport\u001b[0m \u001b[0mtensorflow_text\u001b[0m\u001b[1;33m\u001b[0m\u001b[1;33m\u001b[0m\u001b[0m\n\u001b[0m\u001b[0;32m 2\u001b[0m \u001b[1;32mimport\u001b[0m \u001b[0mtensorflow_hub\u001b[0m \u001b[1;32mas\u001b[0m \u001b[0mhub\u001b[0m\u001b[1;33m\u001b[0m\u001b[1;33m\u001b[0m\u001b[0m\n\u001b[0;32m 3\u001b[0m \u001b[0mvectorizer\u001b[0m \u001b[1;33m=\u001b[0m \u001b[0mhub\u001b[0m\u001b[1;33m.\u001b[0m\u001b[0mKerasLayer\u001b[0m\u001b[1;33m(\u001b[0m\u001b[1;34m'https://tfhub.dev/tensorflow/bert_en_uncased_preprocess/3'\u001b[0m\u001b[1;33m)\u001b[0m\u001b[1;33m\u001b[0m\u001b[1;33m\u001b[0m\u001b[0m\n",
|
||||
"\u001b[1;31mModuleNotFoundError\u001b[0m: No module named 'tensorflow_text'"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"import tensorflow_text \n",
|
||||
"import tensorflow_hub as hub\n",
|
||||
"vectorizer = hub.KerasLayer('https://tfhub.dev/tensorflow/bert_en_uncased_preprocess/3')"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 7,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"{'input_type_ids': <tf.Tensor: shape=(1, 128), dtype=int32, numpy=\n",
|
||||
" array([[0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,\n",
|
||||
" 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,\n",
|
||||
" 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,\n",
|
||||
" 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,\n",
|
||||
" 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,\n",
|
||||
" 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]],\n",
|
||||
" dtype=int32)>,\n",
|
||||
" 'input_word_ids': <tf.Tensor: shape=(1, 128), dtype=int32, numpy=\n",
|
||||
" array([[ 101, 1045, 2293, 19081, 102, 0, 0, 0, 0,\n",
|
||||
" 0, 0, 0, 0, 0, 0, 0, 0, 0,\n",
|
||||
" 0, 0, 0, 0, 0, 0, 0, 0, 0,\n",
|
||||
" 0, 0, 0, 0, 0, 0, 0, 0, 0,\n",
|
||||
" 0, 0, 0, 0, 0, 0, 0, 0, 0,\n",
|
||||
" 0, 0, 0, 0, 0, 0, 0, 0, 0,\n",
|
||||
" 0, 0, 0, 0, 0, 0, 0, 0, 0,\n",
|
||||
" 0, 0, 0, 0, 0, 0, 0, 0, 0,\n",
|
||||
" 0, 0, 0, 0, 0, 0, 0, 0, 0,\n",
|
||||
" 0, 0, 0, 0, 0, 0, 0, 0, 0,\n",
|
||||
" 0, 0, 0, 0, 0, 0, 0, 0, 0,\n",
|
||||
" 0, 0, 0, 0, 0, 0, 0, 0, 0,\n",
|
||||
" 0, 0, 0, 0, 0, 0, 0, 0, 0,\n",
|
||||
" 0, 0, 0, 0, 0, 0, 0, 0, 0,\n",
|
||||
" 0, 0]], dtype=int32)>,\n",
|
||||
" 'input_mask': <tf.Tensor: shape=(1, 128), dtype=int32, numpy=\n",
|
||||
" array([[1, 1, 1, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,\n",
|
||||
" 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,\n",
|
||||
" 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,\n",
|
||||
" 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,\n",
|
||||
" 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,\n",
|
||||
" 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]],\n",
|
||||
" dtype=int32)>}"
|
||||
]
|
||||
},
|
||||
"execution_count": 7,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"vectorizer(['I love transformers'])"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Il est important d'utiliser le même vectoriseur que celui avec lequel le réseau original a été entraîné. De plus, le vectoriseur BERT renvoie trois composants :\n",
|
||||
"* `input_word_ids`, qui est une séquence de numéros de tokens pour la phrase d'entrée\n",
|
||||
"* `input_mask`, qui indique quelle partie de la séquence contient l'entrée réelle et laquelle est du remplissage. Cela est similaire au masque produit par la couche `Masking`\n",
|
||||
"* `input_type_ids` est utilisé pour les tâches de modélisation de langage et permet de spécifier deux phrases d'entrée dans une seule séquence.\n",
|
||||
"\n",
|
||||
"Ensuite, nous pouvons instancier l'extracteur de caractéristiques BERT :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 8,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"bert = hub.KerasLayer('https://tfhub.dev/tensorflow/small_bert/bert_en_uncased_L-4_H-128_A-2/1')"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 9,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"pooled_output -> (1, 128)\n",
|
||||
"encoder_outputs -> 4\n",
|
||||
"sequence_output -> (1, 128, 128)\n",
|
||||
"default -> (1, 128)\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"z = bert(vectorizer(['I love transformers']))\n",
|
||||
"for i,x in z.items():\n",
|
||||
" print(f\"{i} -> { len(x) if isinstance(x, list) else x.shape }\")"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Ainsi, la couche BERT retourne plusieurs résultats utiles :\n",
|
||||
"* `pooled_output` est le résultat de la moyenne de tous les tokens dans la séquence. Vous pouvez le considérer comme une représentation sémantique intelligente de tout le réseau. Cela équivaut à la sortie de la couche `GlobalAveragePooling1D` dans notre modèle précédent.\n",
|
||||
"* `sequence_output` est la sortie de la dernière couche du transformeur (correspond à la sortie de `TransformerBlock` dans notre modèle ci-dessus).\n",
|
||||
"* `encoder_outputs` sont les sorties de toutes les couches du transformeur. Étant donné que nous avons chargé un modèle BERT à 4 couches (comme vous pouvez probablement le deviner d'après le nom, qui contient `4_H`), il possède 4 tenseurs. Le dernier est identique à `sequence_output`.\n",
|
||||
"\n",
|
||||
"Nous allons maintenant définir le modèle de classification de bout en bout. Nous utiliserons une *définition fonctionnelle de modèle*, où nous définissons l'entrée du modèle, puis fournissons une série d'expressions pour calculer sa sortie. Nous rendrons également les poids du modèle BERT non entraînables et entraînerons uniquement le classificateur final :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 10,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Model: \"model\"\n",
|
||||
"__________________________________________________________________________________________________\n",
|
||||
"Layer (type) Output Shape Param # Connected to \n",
|
||||
"==================================================================================================\n",
|
||||
"input_1 (InputLayer) [(None,)] 0 \n",
|
||||
"__________________________________________________________________________________________________\n",
|
||||
"keras_layer (KerasLayer) {'input_type_ids': ( 0 input_1[0][0] \n",
|
||||
"__________________________________________________________________________________________________\n",
|
||||
"keras_layer_1 (KerasLayer) {'pooled_output': (N 4782465 keras_layer[0][0] \n",
|
||||
" keras_layer[0][1] \n",
|
||||
" keras_layer[0][2] \n",
|
||||
"__________________________________________________________________________________________________\n",
|
||||
"dropout_4 (Dropout) (None, 128) 0 keras_layer_1[0][5] \n",
|
||||
"__________________________________________________________________________________________________\n",
|
||||
"dense_4 (Dense) (None, 4) 516 dropout_4[0][0] \n",
|
||||
"==================================================================================================\n",
|
||||
"Total params: 4,782,981\n",
|
||||
"Trainable params: 516\n",
|
||||
"Non-trainable params: 4,782,465\n",
|
||||
"__________________________________________________________________________________________________\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"inp = keras.Input(shape=(),dtype=tf.string)\n",
|
||||
"x = vectorizer(inp)\n",
|
||||
"x = bert(x)\n",
|
||||
"x = keras.layers.Dropout(0.1)(x['pooled_output'])\n",
|
||||
"out = keras.layers.Dense(4,activation='softmax')(x)\n",
|
||||
"model = keras.models.Model(inp,out)\n",
|
||||
"bert.trainable = False\n",
|
||||
"model.summary()"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 11,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"938/938 [==============================] - 528s 559ms/step - loss: 0.8056 - acc: 0.6983 - val_loss: 0.5953 - val_acc: 0.7888\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"<tensorflow.python.keras.callbacks.History at 0x7f9bb1e36d00>"
|
||||
]
|
||||
},
|
||||
"execution_count": 11,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"model.compile(loss='sparse_categorical_crossentropy',metrics=['acc'], optimizer='adam')\n",
|
||||
"model.fit(ds_train.map(tupelize).batch(128),validation_data=ds_test.map(tupelize).batch(128))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Bien que le nombre de paramètres entraînables soit faible, le processus est assez lent, car l'extracteur de caractéristiques BERT est très gourmand en calcul. Il semble que nous n'ayons pas réussi à atteindre une précision raisonnable, soit par manque d'entraînement, soit par insuffisance des paramètres du modèle.\n",
|
||||
"\n",
|
||||
"Essayons de déverrouiller les poids de BERT et de l'entraîner également. Cela nécessite un taux d'apprentissage très faible, ainsi qu'une stratégie d'entraînement plus prudente avec un **warmup**, en utilisant l'optimiseur **AdamW**. Nous utiliserons le package `tf-models-official` pour créer l'optimiseur :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 12,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Model: \"model\"\n",
|
||||
"__________________________________________________________________________________________________\n",
|
||||
"Layer (type) Output Shape Param # Connected to \n",
|
||||
"==================================================================================================\n",
|
||||
"input_1 (InputLayer) [(None,)] 0 \n",
|
||||
"__________________________________________________________________________________________________\n",
|
||||
"keras_layer (KerasLayer) {'input_type_ids': ( 0 input_1[0][0] \n",
|
||||
"__________________________________________________________________________________________________\n",
|
||||
"keras_layer_1 (KerasLayer) {'pooled_output': (N 4782465 keras_layer[0][0] \n",
|
||||
" keras_layer[0][1] \n",
|
||||
" keras_layer[0][2] \n",
|
||||
"__________________________________________________________________________________________________\n",
|
||||
"dropout_4 (Dropout) (None, 128) 0 keras_layer_1[0][5] \n",
|
||||
"__________________________________________________________________________________________________\n",
|
||||
"dense_4 (Dense) (None, 4) 516 dropout_4[0][0] \n",
|
||||
"==================================================================================================\n",
|
||||
"Total params: 4,782,981\n",
|
||||
"Trainable params: 4,782,980\n",
|
||||
"Non-trainable params: 1\n",
|
||||
"__________________________________________________________________________________________________\n",
|
||||
"938/938 [==============================] - 629s 664ms/step - loss: 0.6344 - acc: 0.7658 - val_loss: 0.4876 - val_acc: 0.8247\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"<tensorflow.python.keras.callbacks.History at 0x7f9bb0bd0070>"
|
||||
]
|
||||
},
|
||||
"execution_count": 12,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"from official.nlp import optimization \n",
|
||||
"bert.trainable=True\n",
|
||||
"model.summary()\n",
|
||||
"epochs = 3\n",
|
||||
"opt = optimization.create_optimizer(\n",
|
||||
" init_lr=3e-5,\n",
|
||||
" num_train_steps=epochs*len(ds_train),\n",
|
||||
" num_warmup_steps=0.1*epochs*len(ds_train),\n",
|
||||
" optimizer_type='adamw')\n",
|
||||
"\n",
|
||||
"model.compile(loss='sparse_categorical_crossentropy',metrics=['acc'], optimizer=opt)\n",
|
||||
"model.fit(ds_train.map(tupelize).batch(128),validation_data=ds_test.map(tupelize).batch(128))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Comme vous pouvez le constater, l'entraînement progresse assez lentement - mais vous pourriez vouloir expérimenter et entraîner le modèle pendant quelques époques (5-10) pour voir si vous pouvez obtenir un meilleur résultat par rapport aux approches que nous avons utilisées auparavant.\n",
|
||||
"\n",
|
||||
"## Bibliothèque Huggingface Transformers\n",
|
||||
"\n",
|
||||
"Une autre méthode très courante (et un peu plus simple) pour utiliser les modèles Transformer est le [package HuggingFace](https://github.com/huggingface/), qui fournit des blocs de construction simples pour différentes tâches de PNL. Il est disponible à la fois pour Tensorflow et PyTorch, un autre framework de réseaux neuronaux très populaire.\n",
|
||||
"\n",
|
||||
"> **Note** : Si vous n'êtes pas intéressé par le fonctionnement de la bibliothèque Transformers - vous pouvez passer directement à la fin de ce notebook, car vous ne verrez rien de fondamentalement différent de ce que nous avons fait précédemment. Nous allons répéter les mêmes étapes d'entraînement du modèle BERT en utilisant une bibliothèque différente et un modèle sensiblement plus grand. Par conséquent, le processus implique un entraînement assez long, donc vous pourriez simplement vouloir parcourir le code.\n",
|
||||
"\n",
|
||||
"Voyons comment notre problème peut être résolu en utilisant [Huggingface Transformers](http://huggingface.co).\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"La première chose à faire est de choisir le modèle que nous allons utiliser. En plus de certains modèles intégrés, Huggingface propose un [répertoire de modèles en ligne](https://huggingface.co/models), où vous pouvez trouver de nombreux modèles pré-entraînés par la communauté. Tous ces modèles peuvent être chargés et utilisés simplement en fournissant un nom de modèle. Tous les fichiers binaires nécessaires au modèle seront automatiquement téléchargés.\n",
|
||||
"\n",
|
||||
"Parfois, vous devrez charger vos propres modèles. Dans ce cas, vous pouvez spécifier le répertoire contenant tous les fichiers pertinents, y compris les paramètres pour le tokenizer, le fichier `config.json` avec les paramètres du modèle, les poids binaires, etc.\n",
|
||||
"\n",
|
||||
"À partir du nom du modèle, nous pouvons instancier à la fois le modèle et le tokenizer. Commençons par un tokenizer :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 2,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"import transformers\n",
|
||||
"\n",
|
||||
"# To load the model from Internet repository using model name. \n",
|
||||
"# Use this if you are running from your own copy of the notebooks\n",
|
||||
"bert_model = 'bert-base-uncased' \n",
|
||||
"\n",
|
||||
"# To load the model from the directory on disk. Use this for Microsoft Learn module, because we have\n",
|
||||
"# prepared all required files for you.\n",
|
||||
"#bert_model = './bert'\n",
|
||||
"\n",
|
||||
"tokenizer = transformers.BertTokenizer.from_pretrained(bert_model)\n",
|
||||
"\n",
|
||||
"MAX_SEQ_LEN = 128\n",
|
||||
"PAD_INDEX = tokenizer.convert_tokens_to_ids(tokenizer.pad_token)\n",
|
||||
"UNK_INDEX = tokenizer.convert_tokens_to_ids(tokenizer.unk_token)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"L'objet `tokenizer` contient la fonction `encode` qui peut être utilisée directement pour encoder du texte :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 3,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"[101, 23435, 12314, 2003, 1037, 2307, 7705, 2005, 17953, 2361, 102]"
|
||||
]
|
||||
},
|
||||
"execution_count": 3,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"tokenizer.encode('Tensorflow is a great framework for NLP')"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Nous pouvons également utiliser le tokenizer pour encoder une séquence d'une manière adaptée à son passage au modèle, c'est-à-dire en incluant les champs `token_ids`, `input_mask`, etc. Nous pouvons également spécifier que nous voulons des tenseurs Tensorflow en fournissant l'argument `return_tensors='tf'` :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"{'input_ids': <tf.Tensor: shape=(1, 5), dtype=int32, numpy=array([[ 101, 7592, 1010, 2045, 102]], dtype=int32)>, 'token_type_ids': <tf.Tensor: shape=(1, 5), dtype=int32, numpy=array([[0, 0, 0, 0, 0]], dtype=int32)>, 'attention_mask': <tf.Tensor: shape=(1, 5), dtype=int32, numpy=array([[1, 1, 1, 1, 1]], dtype=int32)>}"
|
||||
]
|
||||
},
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"tokenizer(['Hello, there'],return_tensors='tf')"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Dans notre cas, nous utiliserons un modèle BERT pré-entraîné appelé `bert-base-uncased`. *Uncased* signifie que le modèle est insensible à la casse.\n",
|
||||
"\n",
|
||||
"Lors de l'entraînement du modèle, nous devons fournir une séquence tokenisée en entrée, et pour cela, nous concevrons un pipeline de traitement des données. Étant donné que `tokenizer.encode` est une fonction Python, nous utiliserons la même approche que dans l'unité précédente en l'appelant avec `py_function` :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 31,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"def process(x):\n",
|
||||
" return tokenizer.encode(x.numpy().decode('utf-8'),return_tensors='tf',padding='max_length',max_length=MAX_SEQ_LEN,truncation=True)[0]\n",
|
||||
"\n",
|
||||
"def process_fn(x):\n",
|
||||
" s = x['title']+' '+x['description']\n",
|
||||
" e = tf.py_function(process,inp=[s],Tout=(tf.int32))\n",
|
||||
" e.set_shape(MAX_SEQ_LEN)\n",
|
||||
" return e,x['label']"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Maintenant, nous pouvons charger le modèle réel en utilisant le package `BertForSequenceClassification`. Cela garantit que notre modèle dispose déjà de l'architecture requise pour la classification, y compris le classificateur final. Vous verrez un message d'avertissement indiquant que les poids du classificateur final ne sont pas initialisés, et que le modèle nécessiterait un pré-entraînement - c'est tout à fait normal, car c'est exactement ce que nous allons faire !\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 32,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"model = transformers.TFBertForSequenceClassification.from_pretrained(bert_model,num_labels=4,output_attentions=False)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 33,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Model: \"tf_bert_for_sequence_classification_1\"\n",
|
||||
"_________________________________________________________________\n",
|
||||
"Layer (type) Output Shape Param # \n",
|
||||
"=================================================================\n",
|
||||
"bert (TFBertMainLayer) multiple 109482240 \n",
|
||||
"_________________________________________________________________\n",
|
||||
"dropout_75 (Dropout) multiple 0 \n",
|
||||
"_________________________________________________________________\n",
|
||||
"classifier (Dense) multiple 3076 \n",
|
||||
"=================================================================\n",
|
||||
"Total params: 109,485,316\n",
|
||||
"Trainable params: 109,485,316\n",
|
||||
"Non-trainable params: 0\n",
|
||||
"_________________________________________________________________\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"model.summary()"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Comme vous pouvez le voir dans `summary()`, le modèle contient presque 110 millions de paramètres ! Présumément, si nous voulons une tâche de classification simple sur un ensemble de données relativement petit, nous ne voulons pas entraîner la couche de base de BERT :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 34,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Model: \"tf_bert_for_sequence_classification_1\"\n",
|
||||
"_________________________________________________________________\n",
|
||||
"Layer (type) Output Shape Param # \n",
|
||||
"=================================================================\n",
|
||||
"bert (TFBertMainLayer) multiple 109482240 \n",
|
||||
"_________________________________________________________________\n",
|
||||
"dropout_75 (Dropout) multiple 0 \n",
|
||||
"_________________________________________________________________\n",
|
||||
"classifier (Dense) multiple 3076 \n",
|
||||
"=================================================================\n",
|
||||
"Total params: 109,485,316\n",
|
||||
"Trainable params: 3,076\n",
|
||||
"Non-trainable params: 109,482,240\n",
|
||||
"_________________________________________________________________\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"model.layers[0].trainable = False\n",
|
||||
"model.summary()"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Nous sommes maintenant prêts à commencer l'entraînement !\n",
|
||||
"\n",
|
||||
"> **Note** : L'entraînement d'un modèle BERT à grande échelle peut être très long ! C'est pourquoi nous allons seulement l'entraîner sur les 32 premiers lots. Cela sert simplement à montrer comment l'entraînement du modèle est configuré. Si vous souhaitez essayer un entraînement à grande échelle, il vous suffit de supprimer les paramètres `steps_per_epoch` et `validation_steps`, et de vous préparer à patienter !\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 30,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"32/32 [==============================] - 142s 4s/step - loss: 1.3896 - acc: 0.2500 - val_loss: 1.3863 - val_acc: 0.2480\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"<tensorflow.python.keras.callbacks.History at 0x7f1d40a4b6a0>"
|
||||
]
|
||||
},
|
||||
"execution_count": 30,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"model.compile('adam','sparse_categorical_crossentropy',['acc'])\n",
|
||||
"tf.get_logger().setLevel('ERROR')\n",
|
||||
"model.fit(ds_train.map(process_fn).batch(32),validation_data=ds_test.map(process_fn).batch(32),steps_per_epoch=32,validation_steps=2)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Si vous augmentez le nombre d'itérations, patientez suffisamment longtemps et entraînez pendant plusieurs époques, vous pouvez vous attendre à ce que la classification avec BERT offre la meilleure précision ! Cela s'explique par le fait que BERT comprend déjà très bien la structure de la langue, et qu'il suffit simplement d'ajuster le classificateur final. Cependant, comme BERT est un modèle volumineux, tout le processus d'entraînement prend beaucoup de temps et nécessite une puissance de calcul conséquente ! (GPU, et de préférence plusieurs).\n",
|
||||
"\n",
|
||||
"> **Note :** Dans notre exemple, nous utilisons l'un des plus petits modèles BERT pré-entraînés. Il existe des modèles plus grands qui sont susceptibles de produire de meilleurs résultats.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## À retenir\n",
|
||||
"\n",
|
||||
"Dans cette unité, nous avons exploré des architectures de modèles très récentes basées sur les **transformers**. Nous les avons appliquées à notre tâche de classification de texte, mais de la même manière, les modèles BERT peuvent être utilisés pour l'extraction d'entités, le questionnement automatique et d'autres tâches de NLP.\n",
|
||||
"\n",
|
||||
"Les modèles basés sur les transformers représentent l'état de l'art actuel en NLP, et dans la plupart des cas, ils devraient être la première solution à expérimenter lorsque vous implémentez des solutions NLP personnalisées. Cependant, comprendre les principes fondamentaux des réseaux neuronaux récurrents abordés dans ce module est extrêmement important si vous souhaitez concevoir des modèles neuronaux avancés.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": []
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"\n---\n\n**Avertissement** : \nCe document a été traduit à l'aide du service de traduction automatique [Co-op Translator](https://github.com/Azure/co-op-translator). Bien que nous nous efforcions d'assurer l'exactitude, veuillez noter que les traductions automatisées peuvent contenir des erreurs ou des inexactitudes. Le document original dans sa langue d'origine doit être considéré comme la source faisant autorité. Pour des informations critiques, il est recommandé de recourir à une traduction professionnelle réalisée par un humain. Nous déclinons toute responsabilité en cas de malentendus ou d'interprétations erronées résultant de l'utilisation de cette traduction.\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
"interpreter": {
|
||||
"hash": "0cb620c6d4b9f7a635928804c26cf22403d89d98d79684e4529119355ee6d5a5"
|
||||
},
|
||||
"kernelspec": {
|
||||
"display_name": "py38_tensorflow",
|
||||
"language": "python",
|
||||
"name": "python3"
|
||||
},
|
||||
"language_info": {
|
||||
"codemirror_mode": {
|
||||
"name": "ipython",
|
||||
"version": 3
|
||||
},
|
||||
"file_extension": ".py",
|
||||
"mimetype": "text/x-python",
|
||||
"name": "python",
|
||||
"nbconvert_exporter": "python",
|
||||
"pygments_lexer": "ipython3",
|
||||
"version": "3.8.12"
|
||||
},
|
||||
"coopTranslator": {
|
||||
"original_hash": "ab59c532409774988ab875f2260e8e53",
|
||||
"translation_date": "2025-08-31T15:18:11+00:00",
|
||||
"source_file": "lessons/5-NLP/18-Transformers/TransformersTF.ipynb",
|
||||
"language_code": "fr"
|
||||
}
|
||||
},
|
||||
"nbformat": 4,
|
||||
"nbformat_minor": 4
|
||||
}
|
||||
|
|
@ -0,0 +1,492 @@
|
|||
{
|
||||
"cells": [
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"# Reconnaissance d'Entités Nommées (NER)\n",
|
||||
"\n",
|
||||
"Ce notebook fait partie du [Curriculum AI pour Débutants](http://aka.ms/ai-beginners).\n",
|
||||
"\n",
|
||||
"Dans cet exemple, nous allons apprendre à entraîner un modèle de NER sur le jeu de données [Corpus Annoté pour la Reconnaissance d'Entités Nommées](https://www.kaggle.com/datasets/abhinavwalia95/entity-annotated-corpus) disponible sur Kaggle. Avant de commencer, veuillez télécharger le fichier [ner_dataset.csv](https://www.kaggle.com/datasets/abhinavwalia95/entity-annotated-corpus?resource=download&select=ner_dataset.csv) dans le répertoire actuel.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 62,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"import pandas as pd\n",
|
||||
"from tensorflow import keras\n",
|
||||
"import numpy as np"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Préparation du jeu de données\n",
|
||||
"\n",
|
||||
"Nous commencerons par lire le jeu de données dans un dataframe. Si vous souhaitez en savoir plus sur l'utilisation de Pandas, consultez une [leçon sur le traitement des données](https://github.com/microsoft/Data-Science-For-Beginners/tree/main/2-Working-With-Data/07-python) dans notre [Data Science pour les débutants](http://aka.ms/datascience-beginners)\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 3,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/html": [
|
||||
"<div>\n",
|
||||
"<style scoped>\n",
|
||||
" .dataframe tbody tr th:only-of-type {\n",
|
||||
" vertical-align: middle;\n",
|
||||
" }\n",
|
||||
"\n",
|
||||
" .dataframe tbody tr th {\n",
|
||||
" vertical-align: top;\n",
|
||||
" }\n",
|
||||
"\n",
|
||||
" .dataframe thead th {\n",
|
||||
" text-align: right;\n",
|
||||
" }\n",
|
||||
"</style>\n",
|
||||
"<table border=\"1\" class=\"dataframe\">\n",
|
||||
" <thead>\n",
|
||||
" <tr style=\"text-align: right;\">\n",
|
||||
" <th></th>\n",
|
||||
" <th>Sentence #</th>\n",
|
||||
" <th>Word</th>\n",
|
||||
" <th>POS</th>\n",
|
||||
" <th>Tag</th>\n",
|
||||
" </tr>\n",
|
||||
" </thead>\n",
|
||||
" <tbody>\n",
|
||||
" <tr>\n",
|
||||
" <th>0</th>\n",
|
||||
" <td>Sentence: 1</td>\n",
|
||||
" <td>Thousands</td>\n",
|
||||
" <td>NNS</td>\n",
|
||||
" <td>O</td>\n",
|
||||
" </tr>\n",
|
||||
" <tr>\n",
|
||||
" <th>1</th>\n",
|
||||
" <td>NaN</td>\n",
|
||||
" <td>of</td>\n",
|
||||
" <td>IN</td>\n",
|
||||
" <td>O</td>\n",
|
||||
" </tr>\n",
|
||||
" <tr>\n",
|
||||
" <th>2</th>\n",
|
||||
" <td>NaN</td>\n",
|
||||
" <td>demonstrators</td>\n",
|
||||
" <td>NNS</td>\n",
|
||||
" <td>O</td>\n",
|
||||
" </tr>\n",
|
||||
" <tr>\n",
|
||||
" <th>3</th>\n",
|
||||
" <td>NaN</td>\n",
|
||||
" <td>have</td>\n",
|
||||
" <td>VBP</td>\n",
|
||||
" <td>O</td>\n",
|
||||
" </tr>\n",
|
||||
" <tr>\n",
|
||||
" <th>4</th>\n",
|
||||
" <td>NaN</td>\n",
|
||||
" <td>marched</td>\n",
|
||||
" <td>VBN</td>\n",
|
||||
" <td>O</td>\n",
|
||||
" </tr>\n",
|
||||
" </tbody>\n",
|
||||
"</table>\n",
|
||||
"</div>"
|
||||
],
|
||||
"text/plain": [
|
||||
" Sentence # Word POS Tag\n",
|
||||
"0 Sentence: 1 Thousands NNS O\n",
|
||||
"1 NaN of IN O\n",
|
||||
"2 NaN demonstrators NNS O\n",
|
||||
"3 NaN have VBP O\n",
|
||||
"4 NaN marched VBN O"
|
||||
]
|
||||
},
|
||||
"execution_count": 3,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"df = pd.read_csv('ner_dataset.csv',encoding='unicode-escape')\n",
|
||||
"df.head()"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Obtenons des étiquettes uniques et créons des dictionnaires de correspondance que nous pouvons utiliser pour convertir les étiquettes en numéros de classe :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"array(['O', 'B-geo', 'B-gpe', 'B-per', 'I-geo', 'B-org', 'I-org', 'B-tim',\n",
|
||||
" 'B-art', 'I-art', 'I-per', 'I-gpe', 'I-tim', 'B-nat', 'B-eve',\n",
|
||||
" 'I-eve', 'I-nat'], dtype=object)"
|
||||
]
|
||||
},
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"tags = df.Tag.unique()\n",
|
||||
"tags"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 8,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"'O'"
|
||||
]
|
||||
},
|
||||
"execution_count": 8,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"id2tag = dict(enumerate(tags))\n",
|
||||
"tag2id = { v : k for k,v in id2tag.items() }\n",
|
||||
"\n",
|
||||
"id2tag[0]"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Maintenant, nous devons faire la même chose avec le vocabulaire. Pour simplifier, nous allons créer un vocabulaire sans tenir compte de la fréquence des mots ; dans la vie réelle, vous pourriez utiliser le vectoriseur de Keras et limiter le nombre de mots.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 14,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"vocab = set(df['Word'].apply(lambda x: x.lower()))\n",
|
||||
"id2word = { i+1 : v for i,v in enumerate(vocab) }\n",
|
||||
"id2word[0] = '<UNK>'\n",
|
||||
"vocab.add('<UNK>')\n",
|
||||
"word2id = { v : k for k,v in id2word.items() }"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Nous devons créer un ensemble de données de phrases pour l'entraînement. Bouclons à travers l'ensemble de données original et séparons toutes les phrases individuelles en `X` (listes de mots) et `Y` (liste de tokens) :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 41,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"X,Y = [],[]\n",
|
||||
"s,t = [],[]\n",
|
||||
"for i,row in df[['Sentence #','Word','Tag']].iterrows():\n",
|
||||
" if pd.isna(row['Sentence #']):\n",
|
||||
" s.append(row['Word'])\n",
|
||||
" t.append(row['Tag'])\n",
|
||||
" else:\n",
|
||||
" if len(s)>0:\n",
|
||||
" X.append(s)\n",
|
||||
" Y.append(t)\n",
|
||||
" s,t = [row['Word']],[row['Tag']]\n",
|
||||
"X.append(s)\n",
|
||||
"Y.append(t)\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": []
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 93,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"([10386,\n",
|
||||
" 23515,\n",
|
||||
" 4134,\n",
|
||||
" 29620,\n",
|
||||
" 7954,\n",
|
||||
" 13583,\n",
|
||||
" 21193,\n",
|
||||
" 12222,\n",
|
||||
" 27322,\n",
|
||||
" 18258,\n",
|
||||
" 5815,\n",
|
||||
" 15880,\n",
|
||||
" 5355,\n",
|
||||
" 25242,\n",
|
||||
" 31327,\n",
|
||||
" 18258,\n",
|
||||
" 27067,\n",
|
||||
" 23515,\n",
|
||||
" 26444,\n",
|
||||
" 14412,\n",
|
||||
" 358,\n",
|
||||
" 26551,\n",
|
||||
" 5011,\n",
|
||||
" 30558],\n",
|
||||
" [0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 2, 0, 0, 0, 0, 0])"
|
||||
]
|
||||
},
|
||||
"execution_count": 93,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"def vectorize(seq):\n",
|
||||
" return [word2id[x.lower()] for x in seq]\n",
|
||||
"\n",
|
||||
"def tagify(seq):\n",
|
||||
" return [tag2id[x] for x in seq]\n",
|
||||
"\n",
|
||||
"Xv = list(map(vectorize,X))\n",
|
||||
"Yv = list(map(tagify,Y))\n",
|
||||
"\n",
|
||||
"Xv[0], Yv[0]"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Pour simplifier, nous allons compléter toutes les phrases avec 0 tokens jusqu'à la longueur maximale. Dans la vie réelle, nous pourrions vouloir utiliser une stratégie plus intelligente et compléter les séquences uniquement au sein d'un mini-lot.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 51,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"X_data = keras.preprocessing.sequence.pad_sequences(Xv,padding='post')\n",
|
||||
"Y_data = keras.preprocessing.sequence.pad_sequences(Yv,padding='post')"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Définir un réseau de classification de tokens\n",
|
||||
"\n",
|
||||
"Nous utiliserons un réseau bidirectionnel LSTM à deux couches pour la classification de tokens. Afin d'appliquer un classificateur dense à chacune des sorties de la dernière couche LSTM, nous utiliserons la construction `TimeDistributed`, qui réplique la même couche dense à chacune des sorties du LSTM à chaque étape :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 94,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Model: \"sequential_3\"\n",
|
||||
"_________________________________________________________________\n",
|
||||
" Layer (type) Output Shape Param # \n",
|
||||
"=================================================================\n",
|
||||
" embedding_4 (Embedding) (None, 104, 300) 9545400 \n",
|
||||
" \n",
|
||||
" bidirectional_6 (Bidirectio (None, 104, 200) 320800 \n",
|
||||
" nal) \n",
|
||||
" \n",
|
||||
" bidirectional_7 (Bidirectio (None, 104, 200) 240800 \n",
|
||||
" nal) \n",
|
||||
" \n",
|
||||
" time_distributed_3 (TimeDis (None, 104, 17) 3417 \n",
|
||||
" tributed) \n",
|
||||
" \n",
|
||||
"=================================================================\n",
|
||||
"Total params: 10,110,417\n",
|
||||
"Trainable params: 10,110,417\n",
|
||||
"Non-trainable params: 0\n",
|
||||
"_________________________________________________________________\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"maxlen = X_data.shape[1]\n",
|
||||
"vocab_size = len(vocab)\n",
|
||||
"num_tags = len(tags)\n",
|
||||
"model = keras.models.Sequential([\n",
|
||||
" keras.layers.Embedding(vocab_size, 300, input_length=maxlen),\n",
|
||||
" keras.layers.Bidirectional(keras.layers.LSTM(units=100, activation='tanh', return_sequences=True)),\n",
|
||||
" keras.layers.Bidirectional(keras.layers.LSTM(units=100, activation='tanh', return_sequences=True)),\n",
|
||||
" keras.layers.TimeDistributed(keras.layers.Dense(num_tags, activation='softmax'))\n",
|
||||
"])\n",
|
||||
"model.compile(loss='sparse_categorical_crossentropy',optimizer='adam',metrics=['acc'])\n",
|
||||
"model.summary()"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Notez ici que nous spécifions explicitement `maxlen` pour notre jeu de données - si nous voulons que le réseau puisse gérer des séquences de longueur variable, nous devons être un peu plus astucieux lors de la définition du réseau.\n",
|
||||
"\n",
|
||||
"Passons maintenant à l'entraînement du modèle. Pour des raisons de rapidité, nous n'entraînerons que pendant une seule époque, mais vous pouvez essayer de prolonger la durée d'entraînement. De plus, vous pourriez vouloir séparer une partie du jeu de données comme jeu de données d'entraînement, afin d'observer la précision de validation.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 57,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"1499/1499 [==============================] - 740s 488ms/step - loss: 0.0667 - acc: 0.9841\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"<keras.callbacks.History at 0x16f0bb2a310>"
|
||||
]
|
||||
},
|
||||
"execution_count": 57,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"model.fit(X_data,Y_data)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Tester le Résultat\n",
|
||||
"\n",
|
||||
"Voyons maintenant comment notre modèle de reconnaissance d'entités fonctionne sur une phrase d'exemple :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 91,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"sent = 'John Smith went to Paris to attend a conference in cancer development institute'\n",
|
||||
"words = sent.lower().split()\n",
|
||||
"v = keras.preprocessing.sequence.pad_sequences([[word2id[x] for x in words]],padding='post',maxlen=maxlen)\n",
|
||||
"res = model(v)[0]"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 92,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"john -> B-per\n",
|
||||
"smith -> I-per\n",
|
||||
"went -> O\n",
|
||||
"to -> O\n",
|
||||
"paris -> B-geo\n",
|
||||
"to -> O\n",
|
||||
"attend -> O\n",
|
||||
"a -> O\n",
|
||||
"conference -> O\n",
|
||||
"in -> O\n",
|
||||
"cancer -> B-org\n",
|
||||
"development -> I-org\n",
|
||||
"institute -> I-org\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"r = np.argmax(res.numpy(),axis=1)\n",
|
||||
"for i,w in zip(r,words):\n",
|
||||
" print(f\"{w} -> {id2tag[i]}\")"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## À retenir\n",
|
||||
"\n",
|
||||
"Même un modèle LSTM simple donne des résultats raisonnables pour la reconnaissance d'entités nommées (NER). Cependant, pour obtenir des résultats nettement meilleurs, vous pourriez envisager d'utiliser de grands modèles de langage pré-entraînés comme BERT. La formation de BERT pour la NER à l'aide de la bibliothèque Huggingface Transformers est décrite [ici](https://huggingface.co/course/chapter7/2?fw=pt).\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"\n---\n\n**Avertissement** : \nCe document a été traduit à l'aide du service de traduction automatique [Co-op Translator](https://github.com/Azure/co-op-translator). Bien que nous nous efforcions d'assurer l'exactitude, veuillez noter que les traductions automatisées peuvent contenir des erreurs ou des inexactitudes. Le document original dans sa langue d'origine doit être considéré comme la source faisant autorité. Pour des informations critiques, il est recommandé de recourir à une traduction professionnelle réalisée par un humain. Nous déclinons toute responsabilité en cas de malentendus ou d'interprétations erronées résultant de l'utilisation de cette traduction.\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
"interpreter": {
|
||||
"hash": "16af2a8bbb083ea23e5e41c7f5787656b2ce26968575d8763f2c4b17f9cd711f"
|
||||
},
|
||||
"kernelspec": {
|
||||
"display_name": "Python 3.8.12 ('py38')",
|
||||
"language": "python",
|
||||
"name": "python3"
|
||||
},
|
||||
"language_info": {
|
||||
"codemirror_mode": {
|
||||
"name": "ipython",
|
||||
"version": 3
|
||||
},
|
||||
"file_extension": ".py",
|
||||
"mimetype": "text/x-python",
|
||||
"name": "python",
|
||||
"nbconvert_exporter": "python",
|
||||
"pygments_lexer": "ipython3",
|
||||
"version": "3.8.12"
|
||||
},
|
||||
"orig_nbformat": 4,
|
||||
"coopTranslator": {
|
||||
"original_hash": "254d25052dcca4ef84f59a05f2935bdc",
|
||||
"translation_date": "2025-08-31T15:20:25+00:00",
|
||||
"source_file": "lessons/5-NLP/19-NER/NER-TF.ipynb",
|
||||
"language_code": "fr"
|
||||
}
|
||||
},
|
||||
"nbformat": 4,
|
||||
"nbformat_minor": 2
|
||||
}
|
||||
|
|
@ -0,0 +1,325 @@
|
|||
{
|
||||
"cells": [
|
||||
{
|
||||
"attachments": {},
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Expérimenter avec OpenAI GPT\n",
|
||||
"\n",
|
||||
"Ce notebook fait partie du [Programme pour débutants en IA](http://aka.ms/ai-beginners).\n",
|
||||
"\n",
|
||||
"Dans ce notebook, nous allons explorer comment nous pouvons utiliser le modèle OpenAI-GPT avec la bibliothèque `transformers` de Hugging Face.\n",
|
||||
"\n",
|
||||
"Sans plus attendre, lançons le pipeline de génération de texte et commençons à générer !\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 1,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stderr",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"c:\\Users\\bethanycheum\\Desktop\\AI-For-Beginners\\.venv\\lib\\site-packages\\tqdm\\auto.py:21: TqdmWarning: IProgress not found. Please update jupyter and ipywidgets. See https://ipywidgets.readthedocs.io/en/stable/user_install.html\n",
|
||||
" from .autonotebook import tqdm as notebook_tqdm\n",
|
||||
"Downloading model.safetensors: 100%|██████████| 479M/479M [04:28<00:00, 1.78MB/s] \n",
|
||||
"c:\\Users\\bethanycheum\\Desktop\\AI-For-Beginners\\.venv\\lib\\site-packages\\huggingface_hub\\file_download.py:133: UserWarning: `huggingface_hub` cache-system uses symlinks by default to efficiently store duplicated files but your machine does not support them in C:\\Users\\bethanycheum\\.cache\\huggingface\\hub. Caching files will still work but in a degraded version that might require more space on your disk. This warning can be disabled by setting the `HF_HUB_DISABLE_SYMLINKS_WARNING` environment variable. For more details, see https://huggingface.co/docs/huggingface_hub/how-to-cache#limitations.\n",
|
||||
"To support symlinks on Windows, you either need to activate Developer Mode or to run Python as an administrator. In order to see activate developer mode, see this article: https://docs.microsoft.com/en-us/windows/apps/get-started/enable-your-device-for-development\n",
|
||||
" warnings.warn(message)\n",
|
||||
"Some weights of OpenAIGPTLMHeadModel were not initialized from the model checkpoint at openai-gpt and are newly initialized: ['position_ids']\n",
|
||||
"You should probably TRAIN this model on a down-stream task to be able to use it for predictions and inference.\n",
|
||||
"Downloading (…)neration_config.json: 100%|██████████| 74.0/74.0 [00:00<00:00, 48.8kB/s]\n",
|
||||
"Downloading (…)olve/main/vocab.json: 100%|██████████| 816k/816k [00:00<00:00, 1.76MB/s]\n",
|
||||
"Downloading (…)olve/main/merges.txt: 100%|██████████| 458k/458k [00:00<00:00, 1.11MB/s]\n",
|
||||
"Downloading (…)/main/tokenizer.json: 100%|██████████| 1.27M/1.27M [00:00<00:00, 2.12MB/s]\n",
|
||||
"Xformers is not installed correctly. If you want to use memory_efficient_attention to accelerate training use the following command to install Xformers\n",
|
||||
"pip install xformers.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"[{'generated_text': \"Hello! I am a neural network, and I want to say that i apologize for not coming to you yourself, for not helping you, and that i was too busy getting dressed and studying for a midterm. you know, the kind where the teachers are like that and they come in pairs with their boyfriends, but not with theirs. it's true, that i have had a girlfriend, and i'm only going on wednesdays and thursdays because i was too busy with college, but maybe\"},\n",
|
||||
" {'generated_text': 'Hello! I am a neural network, and I want to say that we have been blessed with a wonderful gift ; no one of us has died at all. and our spirits are strong, very strong. in one very lucky moment of luck for you, all has been given direction and destiny, and for us there are no more mysteries. the earth has been chosen for you, and that earth is now ours, and you must be forever in our hearts. \" \\n the words, as one,'},\n",
|
||||
" {'generated_text': 'Hello! I am a neural network, and I want to say that if you would just turn and face the general, you would have a nice day. \" \\n \" sure thing, \" said one of the soldiers, and started to run. the rest of the soldiers followed, shouting. the general turned to general zulu, raising his arm. the general said something in his native language, and the general immediately started to run. zulu started to move toward the wall, with the'},\n",
|
||||
" {'generated_text': 'Hello! I am a neural network, and I want to say that i am not a doctor but an anthropologist to you, a specialist, a specialist in the field of astrobiological biology, and that i am very much involved in this investigation. i am not sure, i am not certain, but i can confirm your conclusions and therefore i will go to the top. i have a colleague who has just returned from this expedition and his findings confirm that you are a specialist. that is, he'},\n",
|
||||
" {'generated_text': \"Hello! I am a neural network, and I want to say that everyone here is in agreement that no matter how many times i say to myself,'he was never a man of action on the battlefield,'or'he 'll never take a chance at killing any civilians,'or'he 'll never let his men go undefended against enemy forces of this caliber,'or'that's just what i need in a day like today. \\n you see, there are only three groups that\"}]"
|
||||
]
|
||||
},
|
||||
"execution_count": 1,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"from transformers import pipeline\n",
|
||||
"\n",
|
||||
"model_name = 'openai-gpt' \n",
|
||||
"\n",
|
||||
"generator = pipeline('text-generation', model=model_name)\n",
|
||||
"\n",
|
||||
"generator(\"Hello! I am a neural network, and I want to say that\", max_length=100, num_return_sequences=5)\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"attachments": {},
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Conception de prompts\n",
|
||||
"\n",
|
||||
"Dans certains cas, vous pouvez utiliser directement la génération openai-gpt en concevant des prompts appropriés. Regardez les exemples ci-dessous :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 2,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"[{'generated_text': 'Synonyms of a word cat: the same cat i used to stare at, and you in'},\n",
|
||||
" {'generated_text': 'Synonyms of a word cat: cat of the woods, cat of the hills, cat of'},\n",
|
||||
" {'generated_text': 'Synonyms of a word cat: you! \\n \" it\\'s a girl. \" i said'},\n",
|
||||
" {'generated_text': \"Synonyms of a word cat: big cat. but how come, we didn't hear it\"},\n",
|
||||
" {'generated_text': 'Synonyms of a word cat: \" mea - o - c \" which makes them sound'}]"
|
||||
]
|
||||
},
|
||||
"execution_count": 2,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"generator(\"Synonyms of a word cat:\", max_length=20, num_return_sequences=5)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 3,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"[{'generated_text': 'I love when you say this -> Positive\\nI have myself -> Negative\\nThis is awful for you to say this -> positive this is so horrible - > positive that your brother is gay - >'},\n",
|
||||
" {'generated_text': 'I love when you say this -> Positive\\nI have myself -> Negative\\nThis is awful for you to say this -> negative i will bring this on you -, < positive am i, i'},\n",
|
||||
" {'generated_text': 'I love when you say this -> Positive\\nI have myself -> Negative\\nThis is awful for you to say this -> negative i have self - esteem i must take it - : \\n - -'},\n",
|
||||
" {'generated_text': 'I love when you say this -> Positive\\nI have myself -> Negative\\nThis is awful for you to say this -> negative this is - : \\n if it were true that the devil would have'},\n",
|
||||
" {'generated_text': \"I love when you say this -> Positive\\nI have myself -> Negative\\nThis is awful for you to say this -> positive i have you - > positive it's a bad thing, > positive\"}]"
|
||||
]
|
||||
},
|
||||
"execution_count": 3,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"generator(\"I love when you say this -> Positive\\nI have myself -> Negative\\nThis is awful for you to say this ->\", max_length=40, num_return_sequences=5)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"[{'generated_text': 'Translate English to French: cat => chat, dog => chien, student => new and unusual. there were no more words to be'},\n",
|
||||
" {'generated_text': 'Translate English to French: cat => chat, dog => chien, student => student \\n his eyes were huge in his lean face as'},\n",
|
||||
" {'generated_text': \"Translate English to French: cat => chat, dog => chien, student => the teacher's words, their words, their words.\"}]"
|
||||
]
|
||||
},
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"generator(\"Translate English to French: cat => chat, dog => chien, student => \", top_k=50, max_length=30, num_return_sequences=3)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 5,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"[{'generated_text': 'People who liked the movie The Matrix also liked it, and there was the movie of the first man after us. \\n i wanted to laugh at how stupid these stupid actors were. no, they were'},\n",
|
||||
" {'generated_text': \"People who liked the movie The Matrix also liked the movie, and the film was the result. and that's when the man in the story was brought into reality, after a few decades. \\n a\"},\n",
|
||||
" {'generated_text': 'People who liked the movie The Matrix also liked the movie the matrix, because there was a very old movie movie called the matrix, where there was a great super hero, and the super hero came out'},\n",
|
||||
" {'generated_text': \"People who liked the movie The Matrix also liked the movie that didn't have a chance to pay cash, if they could afford it. most often they got a good deal and a lot of money,\"},\n",
|
||||
" {'generated_text': \"People who liked the movie The Matrix also liked the movie, and i didn't seem to have the same problem. \\n i 'd met the other half of my family. i spent most of my time\"}]"
|
||||
]
|
||||
},
|
||||
"execution_count": 5,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"generator(\"People who liked the movie The Matrix also liked \", max_length=40, num_return_sequences=5)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Stratégies d'échantillonnage de texte\n",
|
||||
"\n",
|
||||
"Jusqu'à présent, nous avons utilisé une stratégie d'échantillonnage **gloutonne** simple, où nous sélectionnons le mot suivant en fonction de la probabilité la plus élevée. Voici comment cela fonctionne :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 6,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"[{'generated_text': 'It was early evening when I can back from work. I usually work late, but this time it was an exception. When I entered a room, I saw my friend, a young man, sprawled across the bed in his bed. \\n \" hi, i\\'m mike eptirard. \" \\n there was silence on the other side of the door. i listened for any trace of life but there was nothing. my heart began to pound, i was starting to sweat, i took out my wallet'},\n",
|
||||
" {'generated_text': 'It was early evening when I can back from work. I usually work late, but this time it was an exception. When I entered a room, I saw my mother on the bed, hugging her legs to her chest and sobbing. i saw my dad and mother from the corner of my eye. \\n elfin face was covered in tears as i entered the room. my dad and mother also wept ; just as they did every other time i came to work. but this time, they had different faces'},\n",
|
||||
" {'generated_text': 'It was early evening when I can back from work. I usually work late, but this time it was an exception. When I entered a room, I saw the room had changed because it was dark. it still smelled like a hospital. a new light shined through from a vent in the ceiling. i found myself in a bathroom and a small room with a sink and a wall of glass. the bathroom billion years ago. not so different from all of the rest of the apartment. \\n now...'},\n",
|
||||
" {'generated_text': 'It was early evening when I can back from work. I usually work late, but this time it was an exception. When I entered a room, I saw a large woman with dark hair and pale skin. she was asleep, but i noticed a faint movement of her face. i could sense she was awake. i got up and walked over to her. \\n \" hello miss. i am inspector michael o\\'dell ; we are investigating the case against you. i wanted to ask if you were the'},\n",
|
||||
" {'generated_text': 'It was early evening when I can back from work. I usually work late, but this time it was an exception. When I entered a room, I saw i had an empty table and three empty chairs. that was all i needed. i had left a note on a table in the center of the room and had a pen in hand. \" \\n \" i think what you were doing was something he was doing to her. \" \\n \" yeah, \" i nodded with a grin. \" i'}]"
|
||||
]
|
||||
},
|
||||
"execution_count": 6,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"prompt = \"It was early evening when I can back from work. I usually work late, but this time it was an exception. When I entered a room, I saw\"\n",
|
||||
"generator(prompt,max_length=100,num_return_sequences=5)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"**La recherche par faisceau** permet au générateur d'explorer plusieurs directions (*faisceaux*) de génération de texte et de sélectionner celles avec le score global le plus élevé. Vous pouvez effectuer une recherche par faisceau en fournissant le paramètre `num_beams`. Vous pouvez également spécifier `no_repeat_ngram_size` pour pénaliser le modèle en cas de répétition de n-grammes d'une taille donnée.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 7,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"[{'generated_text': 'It was early evening when I can back from work. I usually work late, but this time it was an exception. When I entered a room, I saw a man sitting in a chair with his head in his hands. he didn\\'t look up as i approached. \\n \" excuse me, sir, \" i said. \" can i help you? \" \\n the man looked up at me. his eyes were red - rimmed and his face was pale, as if he hadn\\'t slept in days'},\n",
|
||||
" {'generated_text': 'It was early evening when I can back from work. I usually work late, but this time it was an exception. When I entered a room, I saw a man sitting at a desk in the middle of the room. he had his back to me, so i couldn\\'t see what he was doing. \" \\n \" what did he look like? \" i asked as i sat down on the bed next to her. \\n she took a deep breath and looked at me with tears in her eyes'},\n",
|
||||
" {'generated_text': 'It was early evening when I can back from work. I usually work late, but this time it was an exception. When I entered a room, I saw a woman sitting on the bed, reading a book. she looked up at me and smiled. \\n \" hi, \" she said. \" can i help you? \" \\n i sat down next to her and looked around the room. the walls were white, and there was a large window in the middle of the wall that looked out on'},\n",
|
||||
" {'generated_text': 'It was early evening when I can back from work. I usually work late, but this time it was an exception. When I entered a room, I saw a man sitting at a table in the middle of the room. he looked up as i walked in, and when he saw me, he got up and walked over to me. \\n \" can i help you? \" he asked as he put his hand on the small of my back and led me to a chair at the other end of'},\n",
|
||||
" {'generated_text': 'It was early evening when I can back from work. I usually work late, but this time it was an exception. When I entered a room, I saw a woman sitting on the edge of her bed, reading a book. she looked up at me and smiled. \\n \" hello, \" she said. \" can i help you? \" \\n i didn\\'t know what to say, so i just sat down in the chair next to the bed and looked at her. her hair was dark brown'}]"
|
||||
]
|
||||
},
|
||||
"execution_count": 7,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"prompt = \"It was early evening when I can back from work. I usually work late, but this time it was an exception. When I entered a room, I saw\"\n",
|
||||
"generator(prompt,max_length=100,num_return_sequences=5,num_beams=10,no_repeat_ngram_size=2)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"**L'échantillonnage** sélectionne le mot suivant de manière non déterministe, en utilisant la distribution de probabilité retournée par le modèle. Vous activez l'échantillonnage en utilisant le paramètre `do_sample=True`. Vous pouvez également spécifier la `temperature`, pour rendre le modèle plus ou moins déterministe.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 8,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"[{'generated_text': 'It was early evening when I can back from work. I usually work late, but this time it was an exception. When I entered a room, I saw her. she was on the bed, but she looked very different. \\n \" honey, what\\'s the matter? \" i asked. \\n she sat up. \" i can\\'t believe it\\'s real. i\\'ve been dreaming about you for the last two days. \" \\n \" i can\\'t believe it either. i guess that\\'s how'}]"
|
||||
]
|
||||
},
|
||||
"execution_count": 8,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"prompt = \"It was early evening when I can back from work. I usually work late, but this time it was an exception. When I entered a room, I saw\"\n",
|
||||
"generator(prompt,max_length=100,do_sample=True,temperature=0.8)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Nous pouvons également fournir des paramètres supplémentaires pour l'échantillonnage : \n",
|
||||
"* `top_k` spécifie le nombre d'options de mots à considérer lors de l'utilisation de l'échantillonnage. Cela réduit les chances d'obtenir des mots étranges (de faible probabilité) dans notre texte. \n",
|
||||
"* `top_p` est similaire, mais nous choisissons le plus petit sous-ensemble des mots les plus probables, dont la probabilité totale est supérieure à p. \n",
|
||||
"\n",
|
||||
"N'hésitez pas à expérimenter en ajoutant ces paramètres. \n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"attachments": {},
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Affiner vos modèles\n",
|
||||
"\n",
|
||||
"Vous pouvez également [affiner votre modèle](https://learn.microsoft.com/en-us/azure/cognitive-services/openai/how-to/fine-tuning?pivots=programming-language-studio?WT.mc_id=academic-77998-bethanycheum) avec votre propre jeu de données. Cela vous permettra d'ajuster le style du texte tout en conservant la majeure partie du modèle linguistique.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"\n---\n\n**Avertissement** : \nCe document a été traduit à l'aide du service de traduction automatique [Co-op Translator](https://github.com/Azure/co-op-translator). Bien que nous nous efforcions d'assurer l'exactitude, veuillez noter que les traductions automatisées peuvent contenir des erreurs ou des inexactitudes. Le document original dans sa langue d'origine doit être considéré comme la source faisant autorité. Pour des informations critiques, il est recommandé de faire appel à une traduction professionnelle humaine. Nous déclinons toute responsabilité en cas de malentendus ou d'interprétations erronées résultant de l'utilisation de cette traduction.\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
"interpreter": {
|
||||
"hash": "16af2a8bbb083ea23e5e41c7f5787656b2ce26968575d8763f2c4b17f9cd711f"
|
||||
},
|
||||
"kernelspec": {
|
||||
"display_name": "Python 3.8.12 ('py38')",
|
||||
"language": "python",
|
||||
"name": "python3"
|
||||
},
|
||||
"language_info": {
|
||||
"codemirror_mode": {
|
||||
"name": "ipython",
|
||||
"version": 3
|
||||
},
|
||||
"file_extension": ".py",
|
||||
"mimetype": "text/x-python",
|
||||
"name": "python",
|
||||
"nbconvert_exporter": "python",
|
||||
"pygments_lexer": "ipython3",
|
||||
"version": "3.10.11"
|
||||
},
|
||||
"orig_nbformat": 4,
|
||||
"coopTranslator": {
|
||||
"original_hash": "d4ff89615d38924a55594f16d6d20678",
|
||||
"translation_date": "2025-08-31T15:19:40+00:00",
|
||||
"source_file": "lessons/5-NLP/20-LangModels/GPT-PyTorch.ipynb",
|
||||
"language_code": "fr"
|
||||
}
|
||||
},
|
||||
"nbformat": 4,
|
||||
"nbformat_minor": 2
|
||||
}
|
||||
|
|
@ -0,0 +1,49 @@
|
|||
{
|
||||
"cells": [
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Devoir : Équations diophantiennes\n",
|
||||
"\n",
|
||||
"> Ce devoir fait partie du [programme AI for Beginners](http://github.com/microsoft/ai-for-beginners) et s'inspire de [cet article](https://habr.com/post/128704/).\n",
|
||||
"\n",
|
||||
"Votre objectif est de résoudre une **équation diophantienne** - une équation avec des racines entières et des coefficients entiers. Par exemple, considérez l'équation suivante :\n",
|
||||
"\n",
|
||||
"$$a+2b+3c+4d=30$$\n",
|
||||
"\n",
|
||||
"Vous devez trouver des racines entières $a$,$b$,$c$,$d\\in\\mathbb{N}$ qui satisfont cette équation.\n",
|
||||
"\n",
|
||||
"Conseils :\n",
|
||||
"1. Vous pouvez considérer que les racines se situent dans l'intervalle [0;30].\n",
|
||||
"1. Comme un gène, envisagez d'utiliser la liste des valeurs des racines.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": []
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"\n---\n\n**Avertissement** : \nCe document a été traduit à l'aide du service de traduction automatique [Co-op Translator](https://github.com/Azure/co-op-translator). Bien que nous nous efforcions d'assurer l'exactitude, veuillez noter que les traductions automatisées peuvent contenir des erreurs ou des inexactitudes. Le document original dans sa langue d'origine doit être considéré comme la source faisant autorité. Pour des informations critiques, il est recommandé de recourir à une traduction professionnelle réalisée par un humain. Nous déclinons toute responsabilité en cas de malentendus ou d'interprétations erronées résultant de l'utilisation de cette traduction.\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
"language_info": {
|
||||
"name": "python"
|
||||
},
|
||||
"orig_nbformat": 4,
|
||||
"coopTranslator": {
|
||||
"original_hash": "a967e1fa1e11ab2b6467b19349a4a9aa",
|
||||
"translation_date": "2025-08-31T14:22:06+00:00",
|
||||
"source_file": "lessons/6-Other/21-GeneticAlgorithms/Diophantine.ipynb",
|
||||
"language_code": "fr"
|
||||
}
|
||||
},
|
||||
"nbformat": 4,
|
||||
"nbformat_minor": 2
|
||||
}
|
||||
File diff suppressed because one or more lines are too long
|
|
@ -0,0 +1,501 @@
|
|||
{
|
||||
"cells": [
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"# Entraîner un RL à équilibrer un Cartpole\n",
|
||||
"\n",
|
||||
"Ce notebook fait partie du [programme AI for Beginners](http://aka.ms/ai-beginners). Il s'inspire du [tutoriel officiel de PyTorch](https://pytorch.org/tutorials/intermediate/reinforcement_q_learning.html) et de [cette implémentation Cartpole avec PyTorch](https://github.com/yc930401/Actor-Critic-pytorch).\n",
|
||||
"\n",
|
||||
"Dans cet exemple, nous utiliserons le RL pour entraîner un modèle à équilibrer une barre sur un chariot qui peut se déplacer à gauche et à droite sur une échelle horizontale. Nous utiliserons l'environnement [OpenAI Gym](https://www.gymlibrary.ml/) pour simuler la barre.\n",
|
||||
"\n",
|
||||
"> **Note** : Vous pouvez exécuter le code de cette leçon localement (par exemple, depuis Visual Studio Code), auquel cas la simulation s'ouvrira dans une nouvelle fenêtre. Lorsque vous exécutez le code en ligne, il peut être nécessaire d'apporter quelques ajustements au code, comme décrit [ici](https://towardsdatascience.com/rendering-openai-gym-envs-on-binder-and-google-colab-536f99391cc7).\n",
|
||||
"\n",
|
||||
"Nous commencerons par nous assurer que Gym est installé :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"import sys\n",
|
||||
"!{sys.executable} -m pip install gym"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Créons maintenant l'environnement CartPole et voyons comment l'utiliser. Un environnement possède les propriétés suivantes :\n",
|
||||
"\n",
|
||||
"* **Action space** est l'ensemble des actions possibles que nous pouvons effectuer à chaque étape de la simulation \n",
|
||||
"* **Observation space** est l'ensemble des observations que nous pouvons réaliser \n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"import gym\n",
|
||||
"\n",
|
||||
"env = gym.make(\"CartPole-v1\")\n",
|
||||
"\n",
|
||||
"print(f\"Action space: {env.action_space}\")\n",
|
||||
"print(f\"Observation space: {env.observation_space}\")"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Voyons comment fonctionne la simulation. La boucle suivante exécute la simulation jusqu'à ce que `env.step` ne renvoie plus le drapeau de terminaison `done`. Nous choisirons des actions de manière aléatoire en utilisant `env.action_space.sample()`, ce qui signifie que l'expérience échouera probablement très rapidement (l'environnement CartPole se termine lorsque la vitesse du CartPole, sa position ou son angle dépassent certaines limites).\n",
|
||||
"\n",
|
||||
"> La simulation s'ouvrira dans une nouvelle fenêtre. Vous pouvez exécuter le code plusieurs fois et observer son comportement.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"env.reset()\n",
|
||||
"\n",
|
||||
"done = False\n",
|
||||
"total_reward = 0\n",
|
||||
"while not done:\n",
|
||||
" env.render()\n",
|
||||
" obs, rew, done, info = env.step(env.action_space.sample())\n",
|
||||
" total_reward += rew\n",
|
||||
" print(f\"{obs} -> {rew}\")\n",
|
||||
"print(f\"Total reward: {total_reward}\")"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Vous pouvez remarquer que les observations contiennent 4 nombres. Ce sont :\n",
|
||||
"- Position du chariot\n",
|
||||
"- Vitesse du chariot\n",
|
||||
"- Angle de la tige\n",
|
||||
"- Taux de rotation de la tige\n",
|
||||
"\n",
|
||||
"`rew` est la récompense que nous recevons à chaque étape. Vous pouvez constater que dans l'environnement CartPole, vous recevez 1 point de récompense pour chaque étape de simulation, et l'objectif est de maximiser la récompense totale, c'est-à-dire le temps pendant lequel le CartPole peut rester en équilibre sans tomber.\n",
|
||||
"\n",
|
||||
"Pendant l'apprentissage par renforcement, notre objectif est d'entraîner une **politique** $\\pi$, qui pour chaque état $s$ nous indiquera quelle action $a$ entreprendre, donc essentiellement $a = \\pi(s)$.\n",
|
||||
"\n",
|
||||
"Si vous souhaitez une solution probabiliste, vous pouvez considérer la politique comme renvoyant un ensemble de probabilités pour chaque action, c'est-à-dire que $\\pi(a|s)$ représenterait la probabilité que nous devrions entreprendre l'action $a$ dans l'état $s$.\n",
|
||||
"\n",
|
||||
"## Méthode du Gradient de Politique\n",
|
||||
"\n",
|
||||
"Dans l'algorithme d'apprentissage par renforcement le plus simple, appelé **Gradient de Politique**, nous allons entraîner un réseau de neurones à prédire la prochaine action.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"import numpy as np\n",
|
||||
"import matplotlib.pyplot as plt\n",
|
||||
"import torch\n",
|
||||
"\n",
|
||||
"num_inputs = 4\n",
|
||||
"num_actions = 2\n",
|
||||
"\n",
|
||||
"model = torch.nn.Sequential(\n",
|
||||
" torch.nn.Linear(num_inputs, 128, bias=False, dtype=torch.float32),\n",
|
||||
" torch.nn.ReLU(),\n",
|
||||
" torch.nn.Linear(128, num_actions, bias = False, dtype=torch.float32),\n",
|
||||
" torch.nn.Softmax(dim=1)\n",
|
||||
")"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Nous allons entraîner le réseau en réalisant de nombreuses expériences et en mettant à jour notre réseau après chaque exécution. Définissons une fonction qui exécutera l'expérience et renverra les résultats (le **trace** ainsi nommé) - tous les états, actions (et leurs probabilités recommandées), et récompenses :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"def run_episode(max_steps_per_episode = 10000,render=False): \n",
|
||||
" states, actions, probs, rewards = [],[],[],[]\n",
|
||||
" state = env.reset()\n",
|
||||
" for _ in range(max_steps_per_episode):\n",
|
||||
" if render:\n",
|
||||
" env.render()\n",
|
||||
" action_probs = model(torch.from_numpy(np.expand_dims(state,0)))[0]\n",
|
||||
" action = np.random.choice(num_actions, p=np.squeeze(action_probs.detach().numpy()))\n",
|
||||
" nstate, reward, done, info = env.step(action)\n",
|
||||
" if done:\n",
|
||||
" break\n",
|
||||
" states.append(state)\n",
|
||||
" actions.append(action)\n",
|
||||
" probs.append(action_probs.detach().numpy())\n",
|
||||
" rewards.append(reward)\n",
|
||||
" state = nstate\n",
|
||||
" return np.vstack(states), np.vstack(actions), np.vstack(probs), np.vstack(rewards)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Vous pouvez exécuter un épisode avec un réseau non entraîné et observer que la récompense totale (AKA durée de l'épisode) est très faible :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"s, a, p, r = run_episode()\n",
|
||||
"print(f\"Total reward: {np.sum(r)}\")"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"L'un des aspects délicats de l'algorithme de gradient de politique est d'utiliser **les récompenses actualisées**. L'idée est que nous calculons le vecteur des récompenses totales à chaque étape du jeu, et pendant ce processus, nous actualisons les premières récompenses en utilisant un coefficient $gamma$. Nous normalisons également le vecteur résultant, car nous l'utiliserons comme poids pour influencer notre entraînement :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"eps = 0.0001\n",
|
||||
"\n",
|
||||
"def discounted_rewards(rewards,gamma=0.99,normalize=True):\n",
|
||||
" ret = []\n",
|
||||
" s = 0\n",
|
||||
" for r in rewards[::-1]:\n",
|
||||
" s = r + gamma * s\n",
|
||||
" ret.insert(0, s)\n",
|
||||
" if normalize:\n",
|
||||
" ret = (ret-np.mean(ret))/(np.std(ret)+eps)\n",
|
||||
" return ret"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Passons maintenant à l'entraînement proprement dit ! Nous allons exécuter 300 épisodes, et à chaque épisode, nous effectuerons les étapes suivantes :\n",
|
||||
"\n",
|
||||
"1. Exécuter l'expérience et collecter la trace.\n",
|
||||
"2. Calculer la différence (`gradients`) entre les actions effectuées et les probabilités prédites. Plus cette différence est faible, plus nous sommes certains d'avoir pris la bonne décision.\n",
|
||||
"3. Calculer les récompenses actualisées et multiplier les gradients par ces récompenses actualisées - cela garantit que les étapes avec des récompenses plus élevées auront un impact plus important sur le résultat final que celles avec des récompenses plus faibles.\n",
|
||||
"4. Les actions cibles attendues pour notre réseau neuronal seront en partie issues des probabilités prédites pendant l'exécution, et en partie des gradients calculés. Nous utiliserons le paramètre `alpha` pour déterminer dans quelle mesure les gradients et les récompenses sont pris en compte - c'est ce qu'on appelle le *taux d'apprentissage* de l'algorithme de renforcement.\n",
|
||||
"5. Enfin, nous entraînons notre réseau sur les états et les actions attendues, puis nous répétons le processus.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"optimizer = torch.optim.Adam(model.parameters(), lr=0.01)\n",
|
||||
"\n",
|
||||
"def train_on_batch(x, y):\n",
|
||||
" x = torch.from_numpy(x)\n",
|
||||
" y = torch.from_numpy(y)\n",
|
||||
" optimizer.zero_grad()\n",
|
||||
" predictions = model(x)\n",
|
||||
" loss = -torch.mean(torch.log(predictions) * y)\n",
|
||||
" loss.backward()\n",
|
||||
" optimizer.step()\n",
|
||||
" return loss"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"alpha = 1e-4\n",
|
||||
"\n",
|
||||
"history = []\n",
|
||||
"for epoch in range(300):\n",
|
||||
" states, actions, probs, rewards = run_episode()\n",
|
||||
" one_hot_actions = np.eye(2)[actions.T][0]\n",
|
||||
" gradients = one_hot_actions-probs\n",
|
||||
" dr = discounted_rewards(rewards)\n",
|
||||
" gradients *= dr\n",
|
||||
" target = alpha*np.vstack([gradients])+probs\n",
|
||||
" train_on_batch(states,target)\n",
|
||||
" history.append(np.sum(rewards))\n",
|
||||
" if epoch%100==0:\n",
|
||||
" print(f\"{epoch} -> {np.sum(rewards)}\")\n",
|
||||
"\n",
|
||||
"plt.plot(history)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Maintenant, lançons l'épisode avec le rendu pour voir le résultat :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"_ = run_episode(render=True)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Espérons que vous pouvez constater que la tige peut maintenant s'équilibrer assez bien !\n",
|
||||
"\n",
|
||||
"## Modèle Acteur-Critique\n",
|
||||
"\n",
|
||||
"Le modèle Acteur-Critique est une évolution des gradients de politique, dans lequel nous construisons un réseau neuronal pour apprendre à la fois la politique et les récompenses estimées. Le réseau aura deux sorties (ou vous pouvez le voir comme deux réseaux distincts) :\n",
|
||||
"* **Acteur** recommandera l'action à entreprendre en nous donnant la distribution de probabilité des états, comme dans le modèle de gradient de politique.\n",
|
||||
"* **Critique** estimera quelle serait la récompense issue de ces actions. Il renvoie les récompenses totales estimées dans le futur pour l'état donné.\n",
|
||||
"\n",
|
||||
"Définissons un tel modèle :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"from itertools import count\n",
|
||||
"import torch.nn.functional as F"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"device = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\n",
|
||||
"env = gym.make(\"CartPole-v1\")\n",
|
||||
"\n",
|
||||
"state_size = env.observation_space.shape[0]\n",
|
||||
"action_size = env.action_space.n\n",
|
||||
"lr = 0.0001\n",
|
||||
"\n",
|
||||
"class Actor(torch.nn.Module):\n",
|
||||
" def __init__(self, state_size, action_size):\n",
|
||||
" super(Actor, self).__init__()\n",
|
||||
" self.state_size = state_size\n",
|
||||
" self.action_size = action_size\n",
|
||||
" self.linear1 = torch.nn.Linear(self.state_size, 128)\n",
|
||||
" self.linear2 = torch.nn.Linear(128, 256)\n",
|
||||
" self.linear3 = torch.nn.Linear(256, self.action_size)\n",
|
||||
"\n",
|
||||
" def forward(self, state):\n",
|
||||
" output = F.relu(self.linear1(state))\n",
|
||||
" output = F.relu(self.linear2(output))\n",
|
||||
" output = self.linear3(output)\n",
|
||||
" distribution = torch.distributions.Categorical(F.softmax(output, dim=-1))\n",
|
||||
" return distribution\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"class Critic(torch.nn.Module):\n",
|
||||
" def __init__(self, state_size, action_size):\n",
|
||||
" super(Critic, self).__init__()\n",
|
||||
" self.state_size = state_size\n",
|
||||
" self.action_size = action_size\n",
|
||||
" self.linear1 = torch.nn.Linear(self.state_size, 128)\n",
|
||||
" self.linear2 = torch.nn.Linear(128, 256)\n",
|
||||
" self.linear3 = torch.nn.Linear(256, 1)\n",
|
||||
"\n",
|
||||
" def forward(self, state):\n",
|
||||
" output = F.relu(self.linear1(state))\n",
|
||||
" output = F.relu(self.linear2(output))\n",
|
||||
" value = self.linear3(output)\n",
|
||||
" return value"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Nous devrions légèrement modifier nos fonctions `discounted_rewards` et `run_episode` :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"def discounted_rewards(next_value, rewards, masks, gamma=0.99):\n",
|
||||
" R = next_value\n",
|
||||
" returns = []\n",
|
||||
" for step in reversed(range(len(rewards))):\n",
|
||||
" R = rewards[step] + gamma * R * masks[step]\n",
|
||||
" returns.insert(0, R)\n",
|
||||
" return returns\n",
|
||||
"\n",
|
||||
"def run_episode(actor, critic, n_iters):\n",
|
||||
" optimizerA = torch.optim.Adam(actor.parameters())\n",
|
||||
" optimizerC = torch.optim.Adam(critic.parameters())\n",
|
||||
" for iter in range(n_iters):\n",
|
||||
" state = env.reset()\n",
|
||||
" log_probs = []\n",
|
||||
" values = []\n",
|
||||
" rewards = []\n",
|
||||
" masks = []\n",
|
||||
" entropy = 0\n",
|
||||
" env.reset()\n",
|
||||
"\n",
|
||||
" for i in count():\n",
|
||||
" env.render()\n",
|
||||
" state = torch.FloatTensor(state).to(device)\n",
|
||||
" dist, value = actor(state), critic(state)\n",
|
||||
"\n",
|
||||
" action = dist.sample()\n",
|
||||
" next_state, reward, done, _ = env.step(action.cpu().numpy())\n",
|
||||
"\n",
|
||||
" log_prob = dist.log_prob(action).unsqueeze(0)\n",
|
||||
" entropy += dist.entropy().mean()\n",
|
||||
"\n",
|
||||
" log_probs.append(log_prob)\n",
|
||||
" values.append(value)\n",
|
||||
" rewards.append(torch.tensor([reward], dtype=torch.float, device=device))\n",
|
||||
" masks.append(torch.tensor([1-done], dtype=torch.float, device=device))\n",
|
||||
"\n",
|
||||
" state = next_state\n",
|
||||
"\n",
|
||||
" if done:\n",
|
||||
" print('Iteration: {}, Score: {}'.format(iter, i))\n",
|
||||
" break\n",
|
||||
"\n",
|
||||
"\n",
|
||||
" next_state = torch.FloatTensor(next_state).to(device)\n",
|
||||
" next_value = critic(next_state)\n",
|
||||
" returns = discounted_rewards(next_value, rewards, masks)\n",
|
||||
"\n",
|
||||
" log_probs = torch.cat(log_probs)\n",
|
||||
" returns = torch.cat(returns).detach()\n",
|
||||
" values = torch.cat(values)\n",
|
||||
"\n",
|
||||
" advantage = returns - values\n",
|
||||
"\n",
|
||||
" actor_loss = -(log_probs * advantage.detach()).mean()\n",
|
||||
" critic_loss = advantage.pow(2).mean()\n",
|
||||
"\n",
|
||||
" optimizerA.zero_grad()\n",
|
||||
" optimizerC.zero_grad()\n",
|
||||
" actor_loss.backward()\n",
|
||||
" critic_loss.backward()\n",
|
||||
" optimizerA.step()\n",
|
||||
" optimizerC.step()\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Maintenant, nous allons exécuter la boucle principale d'entraînement. Nous utiliserons un processus d'entraînement manuel du réseau en calculant les fonctions de perte appropriées et en mettant à jour les paramètres du réseau :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"\n",
|
||||
"actor = Actor(state_size, action_size).to(device)\n",
|
||||
"critic = Critic(state_size, action_size).to(device)\n",
|
||||
"run_episode(actor, critic, n_iters=100)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": []
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"env.close()"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Points clés\n",
|
||||
"\n",
|
||||
"Nous avons vu deux algorithmes de RL dans cette démonstration : le gradient de politique simple et l'acteur-critique plus sophistiqué. Vous pouvez constater que ces algorithmes fonctionnent avec des notions abstraites d'état, d'action et de récompense - ce qui leur permet d'être appliqués à des environnements très différents.\n",
|
||||
"\n",
|
||||
"L'apprentissage par renforcement nous permet d'apprendre la meilleure stratégie pour résoudre un problème simplement en observant la récompense finale. Le fait de ne pas avoir besoin de jeux de données étiquetés nous permet de répéter les simulations plusieurs fois afin d'optimiser nos modèles. Cependant, il reste encore de nombreux défis dans le domaine du RL, que vous pourrez découvrir si vous décidez de vous concentrer davantage sur cette branche fascinante de l'IA.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"\n---\n\n**Avertissement** : \nCe document a été traduit à l'aide du service de traduction automatique [Co-op Translator](https://github.com/Azure/co-op-translator). Bien que nous nous efforcions d'assurer l'exactitude, veuillez noter que les traductions automatisées peuvent contenir des erreurs ou des inexactitudes. Le document original dans sa langue d'origine doit être considéré comme la source faisant autorité. Pour des informations critiques, il est recommandé de recourir à une traduction humaine professionnelle. Nous déclinons toute responsabilité en cas de malentendus ou d'interprétations erronées résultant de l'utilisation de cette traduction.\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
"kernelspec": {
|
||||
"display_name": "Python 3.10.4 64-bit",
|
||||
"language": "python",
|
||||
"name": "python3"
|
||||
},
|
||||
"language_info": {
|
||||
"codemirror_mode": {
|
||||
"name": "ipython",
|
||||
"version": 3
|
||||
},
|
||||
"file_extension": ".py",
|
||||
"mimetype": "text/x-python",
|
||||
"name": "python",
|
||||
"nbconvert_exporter": "python",
|
||||
"pygments_lexer": "ipython3",
|
||||
"version": "3.10.4"
|
||||
},
|
||||
"orig_nbformat": 4,
|
||||
"vscode": {
|
||||
"interpreter": {
|
||||
"hash": "916dbcbb3f70747c44a77c7bcd40155683ae19c65e1c03b4aa3499c5328201f1"
|
||||
}
|
||||
},
|
||||
"coopTranslator": {
|
||||
"original_hash": "04f8d9978cd11281d81dd037cbf6ce20",
|
||||
"translation_date": "2025-08-31T14:25:30+00:00",
|
||||
"source_file": "lessons/6-Other/22-DeepRL/CartPole-RL-PyTorch.ipynb",
|
||||
"language_code": "fr"
|
||||
}
|
||||
},
|
||||
"nbformat": 4,
|
||||
"nbformat_minor": 2
|
||||
}
|
||||
File diff suppressed because one or more lines are too long
|
|
@ -0,0 +1,109 @@
|
|||
{
|
||||
"cells": [
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"# Entraîner une voiture de montagne à s'échapper\n",
|
||||
"\n",
|
||||
"Travail pratique issu du [Curriculum AI pour débutants](https://github.com/microsoft/ai-for-beginners).\n",
|
||||
"\n",
|
||||
"Votre objectif est d'entraîner un agent RL à contrôler [Mountain Car](https://www.gymlibrary.ml/environments/classic_control/mountain_car/) dans l'environnement OpenAI.\n",
|
||||
"\n",
|
||||
"Commençons par créer l'environnement :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 1,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"import gym\n",
|
||||
"env = gym.make('MountainCar-v0')"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Voyons à quoi ressemble l'expérience aléatoire :\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 2,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"state = env.reset()\n",
|
||||
"while True:\n",
|
||||
" env.render()\n",
|
||||
" action = env.action_space.sample()\n",
|
||||
" state, reward, done, info = env.step(action)\n",
|
||||
" if done:\n",
|
||||
" break"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": []
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"## Lost of code here"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 3,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"env.close()"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"\n---\n\n**Avertissement** : \nCe document a été traduit à l'aide du service de traduction automatique [Co-op Translator](https://github.com/Azure/co-op-translator). Bien que nous nous efforcions d'assurer l'exactitude, veuillez noter que les traductions automatisées peuvent contenir des erreurs ou des inexactitudes. Le document original dans sa langue d'origine doit être considéré comme la source faisant autorité. Pour des informations critiques, il est recommandé de recourir à une traduction humaine professionnelle. Nous déclinons toute responsabilité en cas de malentendus ou d'interprétations erronées résultant de l'utilisation de cette traduction.\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
"interpreter": {
|
||||
"hash": "16af2a8bbb083ea23e5e41c7f5787656b2ce26968575d8763f2c4b17f9cd711f"
|
||||
},
|
||||
"kernelspec": {
|
||||
"display_name": "Python 3.8.12 ('py38')",
|
||||
"language": "python",
|
||||
"name": "python3"
|
||||
},
|
||||
"language_info": {
|
||||
"codemirror_mode": {
|
||||
"name": "ipython",
|
||||
"version": 3
|
||||
},
|
||||
"file_extension": ".py",
|
||||
"mimetype": "text/x-python",
|
||||
"name": "python",
|
||||
"nbconvert_exporter": "python",
|
||||
"pygments_lexer": "ipython3",
|
||||
"version": "3.8.12"
|
||||
},
|
||||
"coopTranslator": {
|
||||
"original_hash": "f062b3b18449593ef8e0fcc029868781",
|
||||
"translation_date": "2025-08-31T14:27:52+00:00",
|
||||
"source_file": "lessons/6-Other/22-DeepRL/lab/MountainCar.ipynb",
|
||||
"language_code": "fr"
|
||||
}
|
||||
},
|
||||
"nbformat": 4,
|
||||
"nbformat_minor": 4
|
||||
}
|
||||
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
|
|
@ -1,8 +1,8 @@
|
|||
<!--
|
||||
CO_OP_TRANSLATOR_METADATA:
|
||||
{
|
||||
"original_hash": "f3a6b0ddf7e6e3f33b2a543baf086dc9",
|
||||
"translation_date": "2025-08-24T09:45:36+00:00",
|
||||
"original_hash": "07191303b7ea2aff1d47e2b0fe4bb862",
|
||||
"translation_date": "2025-08-31T14:20:21+00:00",
|
||||
"source_file": "README.md",
|
||||
"language_code": "hi"
|
||||
}
|
||||
|
|
@ -21,13 +21,24 @@ CO_OP_TRANSLATOR_METADATA:
|
|||
|
||||
[](https://discord.gg/zxKYvhSnVp?WT.mc_id=academic-000002-leestott)
|
||||
|
||||
# शुरुआती लोगों के लिए कृत्रिम बुद्धिमत्ता - एक पाठ्यक्रम
|
||||
# शुरुआती लोगों के लिए कृत्रिम बुद्धिमत्ता - एक पाठ्यक्रम
|
||||
|
||||
| द्वारा ](./lessons/sketchnotes/ai-overview.png)|
|
||||
||
|
||||
|:---:|
|
||||
| शुरुआती लोगों के लिए AI - _[@girlie_mac](https://twitter.com/girlie_mac) द्वारा स्केच नोट_ |
|
||||
| शुरुआती लोगों के लिए AI - _[@girlie_mac](https://twitter.com/girlie_mac) द्वारा स्केच_ |
|
||||
|
||||
**कृत्रिम बुद्धिमत्ता** (AI) की दुनिया को हमारे 12-सप्ताह, 24-पाठ वाले पाठ्यक्रम के साथ खोजें! इसमें व्यावहारिक पाठ, क्विज़ और प्रयोगशालाएं शामिल हैं। यह पाठ्यक्रम शुरुआती लोगों के लिए अनुकूल है और इसमें TensorFlow और PyTorch जैसे टूल्स के साथ-साथ AI में नैतिकता भी शामिल है।
|
||||
**कृत्रिम बुद्धिमत्ता** (AI) की दुनिया को हमारे 12-सप्ताह, 24-पाठ वाले पाठ्यक्रम के साथ खोजें! इसमें व्यावहारिक पाठ, क्विज़ और प्रयोगशालाएं शामिल हैं। यह पाठ्यक्रम शुरुआती लोगों के लिए अनुकूल है और इसमें TensorFlow और PyTorch जैसे उपकरणों के साथ-साथ AI में नैतिकता भी शामिल है।
|
||||
|
||||
### 🌐 बहुभाषी समर्थन
|
||||
|
||||
#### GitHub Action के माध्यम से समर्थित (स्वचालित और हमेशा अद्यतन)
|
||||
|
||||
[French](../fr/README.md) | [Spanish](../es/README.md) | [German](../de/README.md) | [Russian](../ru/README.md) | [Arabic](../ar/README.md) | [Persian (Farsi)](../fa/README.md) | [Urdu](../ur/README.md) | [Chinese (Simplified)](../zh/README.md) | [Chinese (Traditional, Macau)](../mo/README.md) | [Chinese (Traditional, Hong Kong)](../hk/README.md) | [Chinese (Traditional, Taiwan)](../tw/README.md) | [Japanese](../ja/README.md) | [Korean](../ko/README.md) | [Hindi](./README.md) | [Bengali](../bn/README.md) | [Marathi](../mr/README.md) | [Nepali](../ne/README.md) | [Punjabi (Gurmukhi)](../pa/README.md) | [Portuguese (Portugal)](../pt/README.md) | [Portuguese (Brazil)](../br/README.md) | [Italian](../it/README.md) | [Polish](../pl/README.md) | [Turkish](../tr/README.md) | [Greek](../el/README.md) | [Thai](../th/README.md) | [Swedish](../sv/README.md) | [Danish](../da/README.md) | [Norwegian](../no/README.md) | [Finnish](../fi/README.md) | [Dutch](../nl/README.md) | [Hebrew](../he/README.md) | [Vietnamese](../vi/README.md) | [Indonesian](../id/README.md) | [Malay](../ms/README.md) | [Tagalog (Filipino)](../tl/README.md) | [Swahili](../sw/README.md) | [Hungarian](../hu/README.md) | [Czech](../cs/README.md) | [Slovak](../sk/README.md) | [Romanian](../ro/README.md) | [Bulgarian](../bg/README.md) | [Serbian (Cyrillic)](../sr/README.md) | [Croatian](../hr/README.md) | [Slovenian](../sl/README.md) | [Ukrainian](../uk/README.md) | [Burmese (Myanmar)](../my/README.md)
|
||||
|
||||
**यदि आप अतिरिक्त अनुवाद चाहते हैं, तो समर्थित भाषाओं की सूची [यहां](https://github.com/Azure/co-op-translator/blob/main/getting_started/supported-languages.md) उपलब्ध है।**
|
||||
|
||||
## समुदाय से जुड़ें
|
||||
[](https://discord.gg/kzRShWzttr)
|
||||
|
||||
## आप क्या सीखेंगे
|
||||
|
||||
|
|
@ -35,23 +46,23 @@ CO_OP_TRANSLATOR_METADATA:
|
|||
|
||||
इस पाठ्यक्रम में, आप सीखेंगे:
|
||||
|
||||
* कृत्रिम बुद्धिमत्ता के विभिन्न दृष्टिकोण, जिसमें **ज्ञान प्रतिनिधित्व** और तर्क के साथ "पुराने अच्छे" प्रतीकात्मक दृष्टिकोण ([GOFAI](https://en.wikipedia.org/wiki/Symbolic_artificial_intelligence)) शामिल हैं।
|
||||
* **न्यूरल नेटवर्क्स** और **डीप लर्निंग**, जो आधुनिक AI के केंद्र में हैं। हम इन महत्वपूर्ण विषयों के पीछे के विचारों को दो सबसे लोकप्रिय फ्रेमवर्क - [TensorFlow](http://Tensorflow.org) और [PyTorch](http://pytorch.org) में कोड का उपयोग करके समझाएंगे।
|
||||
* छवियों और पाठ के साथ काम करने के लिए **न्यूरल आर्किटेक्चर**। हम हाल के मॉडलों को कवर करेंगे लेकिन अत्याधुनिक तकनीकों में थोड़े पीछे हो सकते हैं।
|
||||
* कृत्रिम बुद्धिमत्ता के विभिन्न दृष्टिकोण, जिसमें "पुराने" प्रतीकात्मक दृष्टिकोण शामिल हैं, जैसे **ज्ञान प्रतिनिधित्व** और तर्क ([GOFAI](https://en.wikipedia.org/wiki/Symbolic_artificial_intelligence))।
|
||||
* **न्यूरल नेटवर्क** और **डीप लर्निंग**, जो आधुनिक AI के केंद्र में हैं। हम इन महत्वपूर्ण विषयों के पीछे के विचारों को [TensorFlow](http://Tensorflow.org) और [PyTorch](http://pytorch.org) जैसे लोकप्रिय फ्रेमवर्क्स के कोड के माध्यम से समझाएंगे।
|
||||
* छवियों और पाठ के साथ काम करने के लिए **न्यूरल आर्किटेक्चर**। हम हाल के मॉडलों को कवर करेंगे लेकिन अत्याधुनिक तकनीकों में थोड़ी कमी हो सकती है।
|
||||
* कम लोकप्रिय AI दृष्टिकोण, जैसे **जेनेटिक एल्गोरिदम** और **मल्टी-एजेंट सिस्टम्स**।
|
||||
|
||||
हम इस पाठ्यक्रम में क्या कवर नहीं करेंगे:
|
||||
इस पाठ्यक्रम में हम क्या कवर नहीं करेंगे:
|
||||
|
||||
> [इस पाठ्यक्रम के लिए सभी अतिरिक्त संसाधन हमारी Microsoft Learn संग्रह में खोजें](https://learn.microsoft.com/en-us/collections/7w28iy2xrqzdj0?WT.mc_id=academic-77998-bethanycheum)
|
||||
> [इस पाठ्यक्रम के लिए सभी अतिरिक्त संसाधन हमारे Microsoft Learn संग्रह में खोजें](https://learn.microsoft.com/en-us/collections/7w28iy2xrqzdj0?WT.mc_id=academic-77998-bethanycheum)
|
||||
|
||||
* **व्यवसाय में AI** का उपयोग करने के व्यावसायिक मामले। Microsoft Learn पर [व्यवसाय उपयोगकर्ताओं के लिए AI का परिचय](https://docs.microsoft.com/learn/paths/introduction-ai-for-business-users/?WT.mc_id=academic-77998-bethanycheum) या [AI बिजनेस स्कूल](https://www.microsoft.com/ai/ai-business-school/?WT.mc_id=academic-77998-bethanycheum), जो [INSEAD](https://www.insead.edu/) के सहयोग से विकसित किया गया है, को लेने पर विचार करें।
|
||||
* **व्यवसाय में AI** का उपयोग करने के व्यावसायिक मामले। इसके लिए, [व्यवसाय उपयोगकर्ताओं के लिए AI का परिचय](https://docs.microsoft.com/learn/paths/introduction-ai-for-business-users/?WT.mc_id=academic-77998-bethanycheum) या [AI बिजनेस स्कूल](https://www.microsoft.com/ai/ai-business-school/?WT.mc_id=academic-77998-bethanycheum) पर विचार करें।
|
||||
* **क्लासिक मशीन लर्निंग**, जिसे हमारे [शुरुआती लोगों के लिए मशीन लर्निंग पाठ्यक्रम](http://github.com/Microsoft/ML-for-Beginners) में अच्छी तरह से वर्णित किया गया है।
|
||||
* **[कॉग्निटिव सर्विसेज](https://azure.microsoft.com/services/cognitive-services/?WT.mc_id=academic-77998-bethanycheum)** का उपयोग करके बनाए गए व्यावहारिक AI अनुप्रयोग। इसके लिए, हम Microsoft Learn के [विज़न](https://docs.microsoft.com/learn/paths/create-computer-vision-solutions-azure-cognitive-services/?WT.mc_id=academic-77998-bethanycheum), [नेचुरल लैंग्वेज प्रोसेसिंग](https://docs.microsoft.com/learn/paths/explore-natural-language-processing/?WT.mc_id=academic-77998-bethanycheum), **[Azure OpenAI Service के साथ जनरेटिव AI](https://learn.microsoft.com/en-us/training/paths/develop-ai-solutions-azure-openai/?WT.mc_id=academic-77998-bethanycheum)** और अन्य मॉड्यूल से शुरुआत करने की सिफारिश करते हैं।
|
||||
* विशिष्ट ML **क्लाउड फ्रेमवर्क**, जैसे [Azure Machine Learning](https://azure.microsoft.com/services/machine-learning/?WT.mc_id=academic-77998-bethanycheum), [Microsoft Fabric](https://learn.microsoft.com/en-us/training/paths/get-started-fabric/?WT.mc_id=academic-77998-bethanycheum), या [Azure Databricks](https://docs.microsoft.com/learn/paths/data-engineer-azure-databricks?WT.mc_id=academic-77998-bethanycheum)। [Azure Machine Learning के साथ मशीन लर्निंग समाधान बनाएं और संचालित करें](https://docs.microsoft.com/learn/paths/build-ai-solutions-with-azure-ml-service/?WT.mc_id=academic-77998-bethanycheum) और [Azure Databricks के साथ मशीन लर्निंग समाधान बनाएं और संचालित करें](https://docs.microsoft.com/learn/paths/build-operate-machine-learning-solutions-azure-databricks/?WT.mc_id=academic-77998-bethanycheum) लर्निंग पाथ्स का उपयोग करने पर विचार करें।
|
||||
* **संवादी AI** और **चैट बॉट्स**। इसके लिए एक अलग [संवादी AI समाधान बनाएं](https://docs.microsoft.com/learn/paths/create-conversational-ai-solutions/?WT.mc_id=academic-77998-bethanycheum) लर्निंग पाथ है, और आप अधिक विवरण के लिए [इस ब्लॉग पोस्ट](https://soshnikov.com/azure/hello-bot-conversational-ai-on-microsoft-platform/) का भी संदर्भ ले सकते हैं।
|
||||
* डीप लर्निंग के पीछे का **गहन गणित**। इसके लिए, हम Ian Goodfellow, Yoshua Bengio और Aaron Courville द्वारा लिखित [Deep Learning](https://www.amazon.com/Deep-Learning-Adaptive-Computation-Machine/dp/0262035618) की सिफारिश करेंगे, जो ऑनलाइन भी उपलब्ध है [https://www.deeplearningbook.org/](https://www.deeplearningbook.org/)।
|
||||
* **[कॉग्निटिव सर्विसेज](https://azure.microsoft.com/services/cognitive-services/?WT.mc_id=academic-77998-bethanycheum)** का उपयोग करके व्यावहारिक AI अनुप्रयोग। इसके लिए, [Generative AI with Azure OpenAI Service](https://learn.microsoft.com/en-us/training/paths/develop-ai-solutions-azure-openai/?WT.mc_id=academic-77998-bethanycheum) जैसे मॉड्यूल से शुरुआत करें।
|
||||
* **क्लाउड फ्रेमवर्क्स**, जैसे [Azure Machine Learning](https://azure.microsoft.com/services/machine-learning/?WT.mc_id=academic-77998-bethanycheum)।
|
||||
* **संवादी AI** और **चैट बॉट्स**। इसके लिए [Create conversational AI solutions](https://docs.microsoft.com/learn/paths/create-conversational-ai-solutions/?WT.mc_id=academic-77998-bethanycheum) पर विचार करें।
|
||||
* **डीप लर्निंग के पीछे गहन गणित**। इसके लिए, [Deep Learning](https://www.amazon.com/Deep-Learning-Adaptive-Computation-Machine/dp/0262035618) पुस्तक की सिफारिश की जाती है।
|
||||
|
||||
क्लाउड में _AI_ विषयों के लिए एक सरल परिचय के लिए, आप [Azure पर कृत्रिम बुद्धिमत्ता के साथ शुरुआत करें](https://docs.microsoft.com/learn/paths/get-started-with-artificial-intelligence-on-azure/?WT.mc_id=academic-77998-bethanycheum) लर्निंग पाथ लेने पर विचार कर सकते हैं।
|
||||
क्लाउड में _AI_ के लिए एक सरल परिचय के लिए, [Get started with artificial intelligence on Azure](https://docs.microsoft.com/learn/paths/get-started-with-artificial-intelligence-on-azure/?WT.mc_id=academic-77998-bethanycheum) पर विचार करें।
|
||||
|
||||
# सामग्री
|
||||
|
||||
|
|
@ -61,69 +72,69 @@ CO_OP_TRANSLATOR_METADATA:
|
|||
| I | [**AI का परिचय**](./lessons/1-Intro/README.md) | | |
|
||||
| 01 | [AI का परिचय और इतिहास](./lessons/1-Intro/README.md) | - | - |
|
||||
| II | **प्रतीकात्मक AI** |
|
||||
| 02 | [ज्ञान प्रतिनिधित्व और विशेषज्ञ प्रणाली](./lessons/2-Symbolic/README.md) | [विशेषज्ञ प्रणाली](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/2-Symbolic/Animals.ipynb) / [ऑन्टोलॉजी](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/2-Symbolic/FamilyOntology.ipynb) /[कॉन्सेप्ट ग्राफ](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/2-Symbolic/MSConceptGraph.ipynb) | |
|
||||
| III | [**न्यूरल नेटवर्क्स का परिचय**](./lessons/3-NeuralNetworks/README.md) |||
|
||||
| 03 | [परसेप्ट्रॉन](./lessons/3-NeuralNetworks/03-Perceptron/README.md) | [नोटबुक](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/3-NeuralNetworks/03-Perceptron/Perceptron.ipynb) | [प्रयोगशाला](./lessons/3-NeuralNetworks/03-Perceptron/lab/README.md) |
|
||||
| 04 | [मल्टी-लेयर्ड परसेप्ट्रॉन और अपना फ्रेमवर्क बनाना](./lessons/3-NeuralNetworks/04-OwnFramework/README.md) | [नोटबुक](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/3-NeuralNetworks/04-OwnFramework/OwnFramework.ipynb) | [प्रयोगशाला](./lessons/3-NeuralNetworks/04-OwnFramework/lab/README.md) |
|
||||
| 05 | [फ्रेमवर्क्स (PyTorch/TensorFlow) और ओवरफिटिंग का परिचय](./lessons/3-NeuralNetworks/05-Frameworks/README.md) | [PyTorch](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/3-NeuralNetworks/05-Frameworks/IntroPyTorch.ipynb) / [Keras](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/3-NeuralNetworks/05-Frameworks/IntroKeras.ipynb) / [TensorFlow](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/3-NeuralNetworks/05-Frameworks/IntroKerasTF.ipynb) | [प्रयोगशाला](./lessons/3-NeuralNetworks/05-Frameworks/lab/README.md) |
|
||||
| IV | [**कंप्यूटर विज़न**](./lessons/4-ComputerVision/README.md) | [PyTorch](https://docs.microsoft.com/learn/modules/intro-computer-vision-pytorch/?WT.mc_id=academic-77998-cacaste) / [TensorFlow](https://docs.microsoft.com/learn/modules/intro-computer-vision-TensorFlow/?WT.mc_id=academic-77998-cacaste)| [Microsoft Azure पर कंप्यूटर विज़न का अन्वेषण करें](https://learn.microsoft.com/en-us/collections/7w28iy2xrqzdj0?WT.mc_id=academic-77998-bethanycheum) |
|
||||
| 06 | [कंप्यूटर विज़न का परिचय। OpenCV](./lessons/4-ComputerVision/06-IntroCV/README.md) | [नोटबुक](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/4-ComputerVision/06-IntroCV/OpenCV.ipynb) | [प्रयोगशाला](./lessons/4-ComputerVision/06-IntroCV/lab/README.md) |
|
||||
| 07 | [कन्वोल्यूशनल न्यूरल नेटवर्क्स](./lessons/4-ComputerVision/07-ConvNets/README.md) और [CNN आर्किटेक्चर](./lessons/4-ComputerVision/07-ConvNets/CNN_Architectures.md) | [PyTorch](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/4-ComputerVision/07-ConvNets/ConvNetsPyTorch.ipynb) /[TensorFlow](https://microsoft.github.io/AI-For-Beginners/lessons/4-ComputerVision/07-ConvNets/ConvNetsTF.ipynb) | [प्रयोगशाला](./lessons/4-ComputerVision/07-ConvNets/lab/README.md) |
|
||||
| 08 | [प्री-ट्रेंड नेटवर्क्स और ट्रांसफर लर्निंग](./lessons/4-ComputerVision/08-TransferLearning/README.md) और [ट्रेनिंग ट्रिक्स](./lessons/4-ComputerVision/08-TransferLearning/TrainingTricks.md) | [PyTorch](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/4-ComputerVision/08-TransferLearning/TransferLearningPyTorch.ipynb) / [TensorFlow](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/3-NeuralNetworks/05-Frameworks/IntroKerasTF.ipynb) | [लैब](./lessons/4-ComputerVision/08-TransferLearning/lab/README.md) |
|
||||
| 09 | [ऑटोएन्कोडर्स और VAEs](./lessons/4-ComputerVision/09-Autoencoders/README.md) | [PyTorch](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/4-ComputerVision/09-Autoencoders/AutoEncodersPyTorch.ipynb) / [TensorFlow](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/4-ComputerVision/09-Autoencoders/AutoencodersTF.ipynb) | |
|
||||
| 10 | [जेनरेटिव एडवर्सेरियल नेटवर्क्स और आर्टिस्टिक स्टाइल ट्रांसफर](./lessons/4-ComputerVision/10-GANs/README.md) | [PyTorch](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/4-ComputerVision/10-GANs/GANPyTorch.ipynb) / [TensorFlow](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/4-ComputerVision/10-GANs/GANTF.ipynb) | |
|
||||
| 11 | [ऑब्जेक्ट डिटेक्शन](./lessons/4-ComputerVision/11-ObjectDetection/README.md) | [TensorFlow](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/4-ComputerVision/11-ObjectDetection/ObjectDetection.ipynb) | [लैब](./lessons/4-ComputerVision/11-ObjectDetection/lab/README.md) |
|
||||
| 12 | [सेमांटिक सेगमेंटेशन. U-Net](./lessons/4-ComputerVision/12-Segmentation/README.md) | [PyTorch](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/4-ComputerVision/12-Segmentation/SemanticSegmentationPytorch.ipynb) / [TensorFlow](../../(https:/github.com/microsoft/AI-For-Beginners/blob/main/lessons/4-ComputerVision/12-Segmentation/SemanticSegmentationTF.ipynb)) | |
|
||||
| V | [**नेचुरल लैंग्वेज प्रोसेसिंग**](./lessons/5-NLP/README.md) | [PyTorch](https://docs.microsoft.com/learn/modules/intro-natural-language-processing-pytorch/?WT.mc_id=academic-77998-cacaste) /[TensorFlow](https://docs.microsoft.com/learn/modules/intro-natural-language-processing-TensorFlow/?WT.mc_id=academic-77998-cacaste) | [Microsoft Azure पर नेचुरल लैंग्वेज प्रोसेसिंग एक्सप्लोर करें](https://learn.microsoft.com/en-us/collections/7w28iy2xrqzdj0?WT.mc_id=academic-77998-bethanycheum)|
|
||||
| 13 | [टेक्स्ट रिप्रेजेंटेशन. Bow/TF-IDF](./lessons/5-NLP/13-TextRep/README.md) | [PyTorch](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/5-NLP/13-TextRep/TextRepresentationPyTorch.ipynb) / [TensorFlow](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/5-NLP/13-TextRep/TextRepresentationTF.ipynb) | |
|
||||
| 14 | [सेमांटिक वर्ड एम्बेडिंग्स. Word2Vec और GloVe](./lessons/5-NLP/14-Embeddings/README.md) | [PyTorch](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/5-NLP/14-Embeddings/EmbeddingsPyTorch.ipynb) / [TensorFlow](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/5-NLP/14-Embeddings/EmbeddingsTF.ipynb) | |
|
||||
| 15 | [लैंग्वेज मॉडलिंग. अपने एम्बेडिंग्स को ट्रेन करें](./lessons/5-NLP/15-LanguageModeling/README.md) | [PyTorch](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/5-NLP/15-LanguageModeling/CBoW-PyTorch.ipynb) / [TensorFlow](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/5-NLP/15-LanguageModeling/CBoW-TF.ipynb) | [लैब](./lessons/5-NLP/15-LanguageModeling/lab/README.md) |
|
||||
| 16 | [रिकरेंट न्यूरल नेटवर्क्स](./lessons/5-NLP/16-RNN/README.md) | [PyTorch](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/5-NLP/16-RNN/RNNPyTorch.ipynb) / [TensorFlow](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/5-NLP/16-RNN/RNNTF.ipynb) | |
|
||||
| 17 | [जेनरेटिव रिकरेंट नेटवर्क्स](./lessons/5-NLP/17-GenerativeNetworks/README.md) | [PyTorch](https://microsoft.github.io/AI-For-Beginners/lessons/5-NLP/17-GenerativeNetworks/GenerativePyTorch.md) / [TensorFlow](https://microsoft.github.io/AI-For-Beginners/lessons/5-NLP/17-GenerativeNetworks/GenerativeTF.md) | [लैब](./lessons/5-NLP/17-GenerativeNetworks/lab/README.md) |
|
||||
| 18 | [ट्रांसफॉर्मर्स. BERT.](./lessons/5-NLP/18-Transformers/READMEtransformers.md) | [PyTorch](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/5-NLP/18-Transformers/TransformersPyTorch.ipynb) /[TensorFlow](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/5-NLP/18-Transformers/TransformersTF.ipynb) | |
|
||||
| 19 | [नेम्ड एंटिटी रिकग्निशन](./lessons/5-NLP/19-NER/README.md) | [TensorFlow](https://microsoft.github.io/AI-For-Beginners/lessons/5-NLP/19-NER/NER-TF.ipynb) | [लैब](./lessons/5-NLP/19-NER/lab/README.md) |
|
||||
| 20 | [लार्ज लैंग्वेज मॉडल्स, प्रॉम्प्ट प्रोग्रामिंग और फ्यू-शॉट टास्क्स](./lessons/5-NLP/20-LangModels/READMELargeLang.md) | [PyTorch](https://microsoft.github.io/AI-For-Beginners/lessons/5-NLP/20-LangModels/GPT-PyTorch.ipynb) | |
|
||||
| 02 | [ज्ञान प्रतिनिधित्व और विशेषज्ञ प्रणाली](./lessons/2-Symbolic/README.md) | [विशेषज्ञ प्रणाली](./lessons/2-Symbolic/Animals.ipynb) / [ऑन्टोलॉजी](./lessons/2-Symbolic/FamilyOntology.ipynb) /[कॉन्सेप्ट ग्राफ](./lessons/2-Symbolic/MSConceptGraph.ipynb) | |
|
||||
| III | [**न्यूरल नेटवर्क का परिचय**](./lessons/3-NeuralNetworks/README.md) |||
|
||||
| 03 | [परसेप्ट्रॉन](./lessons/3-NeuralNetworks/03-Perceptron/README.md) | [नोटबुक](./lessons/3-NeuralNetworks/03-Perceptron/Perceptron.ipynb) | [प्रयोगशाला](./lessons/3-NeuralNetworks/03-Perceptron/lab/README.md) |
|
||||
| 04 | [मल्टी-लेयर्ड परसेप्ट्रॉन और अपना फ्रेमवर्क बनाना](./lessons/3-NeuralNetworks/04-OwnFramework/README.md) | [नोटबुक](./lessons/3-NeuralNetworks/04-OwnFramework/OwnFramework.ipynb) | [प्रयोगशाला](./lessons/3-NeuralNetworks/04-OwnFramework/lab/README.md) |
|
||||
| 05 | [फ्रेमवर्क्स (PyTorch/TensorFlow) और ओवरफिटिंग का परिचय](./lessons/3-NeuralNetworks/05-Frameworks/README.md) | [PyTorch](./lessons/3-NeuralNetworks/05-Frameworks/IntroPyTorch.ipynb) / [Keras](./lessons/3-NeuralNetworks/05-Frameworks/IntroKeras.ipynb) / [TensorFlow](./lessons/3-NeuralNetworks/05-Frameworks/IntroKerasTF.ipynb) | [प्रयोगशाला](./lessons/3-NeuralNetworks/05-Frameworks/lab/README.md) |
|
||||
| IV | [**कंप्यूटर विज़न**](./lessons/4-ComputerVision/README.md) | [PyTorch](https://docs.microsoft.com/learn/modules/intro-computer-vision-pytorch/?WT.mc_id=academic-77998-cacaste) / [TensorFlow](https://docs.microsoft.com/learn/modules/intro-computer-vision-TensorFlow/?WT.mc_id=academic-77998-cacaste)| [Microsoft Azure पर कंप्यूटर विज़न का अन्वेषण करें](https://learn.microsoft.com/en-us/collections/7w28iy2xrqzdj0?WT.mc_id=academic-77998-bethanycheum) |
|
||||
| 06 | [कंप्यूटर विज़न का परिचय। OpenCV](./lessons/4-ComputerVision/06-IntroCV/README.md) | [नोटबुक](./lessons/4-ComputerVision/06-IntroCV/OpenCV.ipynb) | [प्रयोगशाला](./lessons/4-ComputerVision/06-IntroCV/lab/README.md) |
|
||||
| 07 | [कन्वोल्यूशनल न्यूरल नेटवर्क्स](./lessons/4-ComputerVision/07-ConvNets/README.md) और [CNN आर्किटेक्चर](./lessons/4-ComputerVision/07-ConvNets/CNN_Architectures.md) | [PyTorch](./lessons/4-ComputerVision/07-ConvNets/ConvNetsPyTorch.ipynb) /[TensorFlow](./lessons/4-ComputerVision/07-ConvNets/ConvNetsTF.ipynb) | [प्रयोगशाला](./lessons/4-ComputerVision/07-ConvNets/lab/README.md) |
|
||||
| 08 | [प्री-ट्रेंड नेटवर्क्स और ट्रांसफर लर्निंग](./lessons/4-ComputerVision/08-TransferLearning/README.md) और [ट्रेनिंग ट्रिक्स](./lessons/4-ComputerVision/08-TransferLearning/TrainingTricks.md) | [PyTorch](./lessons/4-ComputerVision/08-TransferLearning/TransferLearningPyTorch.ipynb) / [TensorFlow](./lessons/3-NeuralNetworks/05-Frameworks/IntroKerasTF.ipynb) | [प्रयोगशाला](./lessons/4-ComputerVision/08-TransferLearning/lab/README.md) |
|
||||
| 09 | [ऑटोएन्कोडर्स और VAEs](./lessons/4-ComputerVision/09-Autoencoders/README.md) | [PyTorch](./lessons/4-ComputerVision/09-Autoencoders/AutoEncodersPyTorch.ipynb) / [TensorFlow](./lessons/4-ComputerVision/09-Autoencoders/AutoencodersTF.ipynb) | |
|
||||
| 10 | [जेनरेटिव एडवर्सेरियल नेटवर्क्स और आर्टिस्टिक स्टाइल ट्रांसफर](./lessons/4-ComputerVision/10-GANs/README.md) | [PyTorch](./lessons/4-ComputerVision/10-GANs/GANPyTorch.ipynb) / [TensorFlow](./lessons/4-ComputerVision/10-GANs/GANTF.ipynb) | |
|
||||
| 11 | [ऑब्जेक्ट डिटेक्शन](./lessons/4-ComputerVision/11-ObjectDetection/README.md) | [TensorFlow](./lessons/4-ComputerVision/11-ObjectDetection/ObjectDetection.ipynb) | [प्रयोगशाला](./lessons/4-ComputerVision/11-ObjectDetection/lab/README.md) |
|
||||
| 12 | [सेमांटिक सेगमेंटेशन। U-Net](./lessons/4-ComputerVision/12-Segmentation/README.md) | [PyTorch](./lessons/4-ComputerVision/12-Segmentation/SemanticSegmentationPytorch.ipynb) / [TensorFlow](./lessons/4-ComputerVision/12-Segmentation/SemanticSegmentationTF.ipynb) | |
|
||||
| V | [**नेचुरल लैंग्वेज प्रोसेसिंग**](./lessons/5-NLP/README.md) | [PyTorch](https://docs.microsoft.com/learn/modules/intro-natural-language-processing-pytorch/?WT.mc_id=academic-77998-cacaste) /[TensorFlow](https://docs.microsoft.com/learn/modules/intro-natural-language-processing-TensorFlow/?WT.mc_id=academic-77998-cacaste) | [Microsoft Azure पर नेचुरल लैंग्वेज प्रोसेसिंग का अन्वेषण करें](https://learn.microsoft.com/en-us/collections/7w28iy2xrqzdj0?WT.mc_id=academic-77998-bethanycheum)|
|
||||
| 13 | [टेक्स्ट रिप्रेजेंटेशन। Bow/TF-IDF](./lessons/5-NLP/13-TextRep/README.md) | [PyTorch](./lessons/5-NLP/13-TextRep/TextRepresentationPyTorch.ipynb) / [TensorFlow](./lessons/5-NLP/13-TextRep/TextRepresentationTF.ipynb) | |
|
||||
| 14 | [सेमांटिक वर्ड एम्बेडिंग्स। Word2Vec और GloVe](./lessons/5-NLP/14-Embeddings/README.md) | [PyTorch](./lessons/5-NLP/14-Embeddings/EmbeddingsPyTorch.ipynb) / [TensorFlow](./lessons/5-NLP/14-Embeddings/EmbeddingsTF.ipynb) | |
|
||||
| 15 | [लैंग्वेज मॉडलिंग। अपने एम्बेडिंग्स को ट्रेन करें](./lessons/5-NLP/15-LanguageModeling/README.md) | [PyTorch](./lessons/5-NLP/15-LanguageModeling/CBoW-PyTorch.ipynb) / [TensorFlow](./lessons/5-NLP/15-LanguageModeling/CBoW-TF.ipynb) | [प्रयोगशाला](./lessons/5-NLP/15-LanguageModeling/lab/README.md) |
|
||||
| 16 | [रिकरेंट न्यूरल नेटवर्क्स](./lessons/5-NLP/16-RNN/README.md) | [PyTorch](./lessons/5-NLP/16-RNN/RNNPyTorch.ipynb) / [TensorFlow](./lessons/5-NLP/16-RNN/RNNTF.ipynb) | |
|
||||
| 17 | [जेनरेटिव रिकारेंट नेटवर्क्स](./lessons/5-NLP/17-GenerativeNetworks/README.md) | [PyTorch](./lessons/5-NLP/17-GenerativeNetworks/GenerativePyTorch.md) / [TensorFlow](./lessons/5-NLP/17-GenerativeNetworks/GenerativeTF.md) | [प्रयोगशाला](./lessons/5-NLP/17-GenerativeNetworks/lab/README.md) |
|
||||
| 18 | [ट्रांसफॉर्मर्स। BERT.](./lessons/5-NLP/18-Transformers/READMEtransformers.md) | [PyTorch](./lessons/5-NLP/18-Transformers/TransformersPyTorch.ipynb) /[TensorFlow](./lessons/5-NLP/18-Transformers/TransformersTF.ipynb) | |
|
||||
| 19 | [नेम्ड एंटिटी रिकग्निशन](./lessons/5-NLP/19-NER/README.md) | [TensorFlow](./lessons/5-NLP/19-NER/NER-TF.ipynb) | [प्रयोगशाला](./lessons/5-NLP/19-NER/lab/README.md) |
|
||||
| 20 | [लार्ज लैंग्वेज मॉडल्स, प्रॉम्प्ट प्रोग्रामिंग और फ्यू-शॉट टास्क्स](./lessons/5-NLP/20-LangModels/READMELargeLang.md) | [PyTorch](./lessons/5-NLP/20-LangModels/GPT-PyTorch.ipynb) | |
|
||||
| VI | **अन्य AI तकनीकें** || |
|
||||
| 21 | [जेनेटिक एल्गोरिदम्स](./lessons/6-Other/21-GeneticAlgorithms/README.md) | [Notebook](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/6-Other/21-GeneticAlgorithms/Genetic.ipynb) | |
|
||||
| 22 | [डीप रिइंफोर्समेंट लर्निंग](./lessons/6-Other/22-DeepRL/README.md) | [PyTorch](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/6-Other/22-DeepRL/CartPole-RL-PyTorch.ipynb) /[TensorFlow](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/6-Other/22-DeepRL/CartPole-RL-TF.ipynb) | [लैब](./lessons/6-Other/22-DeepRL/lab/README.md) |
|
||||
| 21 | [जेनेटिक एल्गोरिदम](./lessons/6-Other/21-GeneticAlgorithms/README.md) | [नोटबुक](./lessons/6-Other/21-GeneticAlgorithms/Genetic.ipynb) | |
|
||||
| 22 | [डीप रिइंफोर्समेंट लर्निंग](./lessons/6-Other/22-DeepRL/README.md) | [PyTorch](./lessons/6-Other/22-DeepRL/CartPole-RL-PyTorch.ipynb) /[TensorFlow](./lessons/6-Other/22-DeepRL/CartPole-RL-TF.ipynb) | [प्रयोगशाला](./lessons/6-Other/22-DeepRL/lab/README.md) |
|
||||
| 23 | [मल्टी-एजेंट सिस्टम्स](./lessons/6-Other/23-MultiagentSystems/README.md) | | |
|
||||
| VII | **AI एथिक्स** | | |
|
||||
| 24 | [AI एथिक्स और जिम्मेदार AI](./lessons/7-Ethics/README.md) | [Microsoft Learn: जिम्मेदार AI सिद्धांत](https://docs.microsoft.com/learn/paths/responsible-ai-business-principles/?WT.mc_id=academic-77998-cacaste) | |
|
||||
| IX | **एक्स्ट्रा** | | |
|
||||
| 25 | [मल्टी-मोडल नेटवर्क्स, CLIP और VQGAN](./lessons/X-Extras/X1-MultiModal/README.md) | [Notebook](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/X-Extras/X1-MultiModal/Clip.ipynb) | |
|
||||
| VII | **AI नैतिकता** | | |
|
||||
| 24 | [AI नैतिकता और जिम्मेदार AI](./lessons/7-Ethics/README.md) | [Microsoft Learn: जिम्मेदार AI सिद्धांत](https://docs.microsoft.com/learn/paths/responsible-ai-business-principles/?WT.mc_id=academic-77998-cacaste) | |
|
||||
| IX | **अतिरिक्त** | | |
|
||||
| 25 | [मल्टी-मोडल नेटवर्क्स, CLIP और VQGAN](./lessons/X-Extras/X1-MultiModal/README.md) | [नोटबुक](./lessons/X-Extras/X1-MultiModal/Clip.ipynb) | |
|
||||
|
||||
## प्रत्येक पाठ में शामिल है
|
||||
|
||||
* प्री-रीडिंग सामग्री
|
||||
* निष्पादन योग्य Jupyter नोटबुक्स, जो अक्सर फ्रेमवर्क (**PyTorch** या **TensorFlow**) के लिए विशिष्ट होती हैं। निष्पादन योग्य नोटबुक में बहुत सैद्धांतिक सामग्री भी होती है, इसलिए विषय को समझने के लिए आपको नोटबुक का कम से कम एक संस्करण (PyTorch या TensorFlow) देखना होगा।
|
||||
* **लैब्स** कुछ विषयों के लिए उपलब्ध हैं, जो आपको सीखी गई सामग्री को किसी विशिष्ट समस्या पर लागू करने का अवसर देते हैं।
|
||||
* कुछ सेक्शन में [**MS Learn**](https://learn.microsoft.com/en-us/collections/7w28iy2xrqzdj0?WT.mc_id=academic-77998-bethanycheum) मॉड्यूल्स के लिंक शामिल हैं जो संबंधित विषयों को कवर करते हैं।
|
||||
* पूर्व-पढ़ाई सामग्री
|
||||
* निष्पादन योग्य Jupyter नोटबुक्स, जो अक्सर फ्रेमवर्क (**PyTorch** या **TensorFlow**) के लिए विशिष्ट होती हैं। निष्पादन योग्य नोटबुक में बहुत सारा सैद्धांतिक सामग्री भी होती है, इसलिए विषय को समझने के लिए आपको कम से कम एक संस्करण (PyTorch या TensorFlow) को पढ़ना होगा।
|
||||
* **प्रयोगशालाएं**, जो कुछ विषयों के लिए उपलब्ध हैं, आपको सीखी गई सामग्री को किसी विशिष्ट समस्या पर लागू करने का अवसर देती हैं।
|
||||
* कुछ अनुभागों में [**MS Learn**](https://learn.microsoft.com/en-us/collections/7w28iy2xrqzdj0?WT.mc_id=academic-77998-bethanycheum) मॉड्यूल्स के लिंक शामिल हैं, जो संबंधित विषयों को कवर करते हैं।
|
||||
|
||||
## शुरुआत कैसे करें
|
||||
## शुरुआत करें
|
||||
|
||||
- हमने आपके विकास पर्यावरण को सेटअप करने में मदद करने के लिए एक [सेटअप पाठ](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/0-course-setup/setup.md) बनाया है।
|
||||
- शिक्षकों के लिए, हमने आपके लिए एक [पाठ्यक्रम सेटअप पाठ](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/0-course-setup/for-teachers.md) भी बनाया है!
|
||||
- [VSCode या Codepace में कोड कैसे चलाएं](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/0-course-setup/how-to-run.md)
|
||||
- हमने आपके विकास पर्यावरण को सेटअप करने में मदद के लिए एक [सेटअप पाठ](./lessons/0-course-setup/setup.md) बनाया है।
|
||||
- शिक्षकों के लिए, हमने आपके लिए एक [पाठ्यक्रम सेटअप पाठ](./lessons/0-course-setup/for-teachers.md) भी बनाया है!
|
||||
- [VSCode या Codepace में कोड कैसे चलाएं](./lessons/0-course-setup/how-to-run.md)
|
||||
|
||||
इन चरणों का पालन करें:
|
||||
इन चरणों का पालन करें:
|
||||
|
||||
रेपॉजिटरी को फोर्क करें: इस पेज के ऊपर-दाईं ओर "Fork" बटन पर क्लिक करें।
|
||||
रिपॉजिटरी को फोर्क करें: इस पेज के ऊपर-दाईं ओर "Fork" बटन पर क्लिक करें।
|
||||
|
||||
रेपॉजिटरी को क्लोन करें: `git clone https://github.com/microsoft/AI-For-Beginners.git`
|
||||
रिपॉजिटरी को क्लोन करें: `git clone https://github.com/microsoft/AI-For-Beginners.git`
|
||||
|
||||
इस रेपो को स्टार (🌟) करना न भूलें ताकि इसे बाद में आसानी से ढूंढ सकें।
|
||||
इस रिपॉजिटरी को बाद में आसानी से खोजने के लिए इसे स्टार (🌟) करना न भूलें।
|
||||
|
||||
## अन्य शिक्षार्थियों से मिलें
|
||||
|
||||
हमारे [आधिकारिक AI Discord सर्वर](https://aka.ms/genai-discord?WT.mc_id=academic-105485-bethanycheum) में शामिल हों ताकि आप इस कोर्स को लेने वाले अन्य शिक्षार्थियों से मिल सकें, नेटवर्क बना सकें और सहायता प्राप्त कर सकें।
|
||||
हमारे [आधिकारिक AI Discord सर्वर](https://aka.ms/genai-discord?WT.mc_id=academic-105485-bethanycheum) से जुड़ें, इस कोर्स को लेने वाले अन्य शिक्षार्थियों से मिलने और नेटवर्क बनाने के लिए और सहायता प्राप्त करने के लिए।
|
||||
|
||||
यदि आपके पास उत्पाद प्रतिक्रिया या प्रश्न हैं, तो हमारे [Azure AI Foundry Developer Forum](https://aka.ms/foundry/forum) पर जाएं।
|
||||
यदि आपके पास उत्पाद प्रतिक्रिया या प्रश्न हैं, तो हमारे [Azure AI Foundry डेवलपर फोरम](https://aka.ms/foundry/forum) पर जाएं।
|
||||
|
||||
## क्विज़
|
||||
> **क्विज़ के बारे में एक नोट**: सभी क्विज़ Quiz-app फ़ोल्डर में etc\quiz-app में संग्रहीत हैं। इन्हें पाठों के भीतर से जोड़ा गया है। क्विज़ ऐप को स्थानीय रूप से चलाया जा सकता है या Azure पर तैनात किया जा सकता है; `quiz-app` फ़ोल्डर में दिए गए निर्देशों का पालन करें। इन्हें धीरे-धीरे स्थानीयकृत किया जा रहा है।
|
||||
## क्विज़
|
||||
> **क्विज़ के बारे में एक नोट**: सभी क्विज़ Quiz-app फ़ोल्डर में etc\quiz-app में संग्रहीत हैं, या [ऑनलाइन यहां](https://ff-quizzes.netlify.app/) उपलब्ध हैं। इन्हें पाठों के भीतर से जोड़ा गया है। क्विज़ ऐप को स्थानीय रूप से चलाया जा सकता है या Azure पर डिप्लॉय किया जा सकता है; `quiz-app` फ़ोल्डर में दिए गए निर्देशों का पालन करें। इन्हें धीरे-धीरे स्थानीय भाषाओं में अनुवादित किया जा रहा है।
|
||||
## मदद चाहिए
|
||||
|
||||
क्या आपके पास सुझाव हैं या आपने वर्तनी या कोड में त्रुटियां पाई हैं? एक समस्या उठाएं या एक पुल अनुरोध बनाएं।
|
||||
क्या आपके पास सुझाव हैं या आपने वर्तनी या कोड में त्रुटियां पाई हैं? एक मुद्दा उठाएं या एक पुल अनुरोध बनाएं।
|
||||
|
||||
## विशेष धन्यवाद
|
||||
|
||||
|
|
@ -137,11 +148,11 @@ CO_OP_TRANSLATOR_METADATA:
|
|||
|
||||
हमारी टीम अन्य पाठ्यक्रम भी बनाती है! इन्हें देखें:
|
||||
|
||||
- [शुरुआती के लिए जनरेटिव AI](https://aka.ms/genai-beginners)
|
||||
- [शुरुआती के लिए जनरेटिव AI .NET](https://github.com/microsoft/Generative-AI-for-beginners-dotnet)
|
||||
- [जावास्क्रिप्ट के साथ जनरेटिव AI](https://github.com/microsoft/generative-ai-with-javascript)
|
||||
- [जावा के साथ जनरेटिव AI](https://github.com/microsoft/Generative-AI-for-beginners-java)
|
||||
- [शुरुआती के लिए AI](https://aka.ms/ai-beginners)
|
||||
- [शुरुआती के लिए जनरेटिव एआई](https://aka.ms/genai-beginners)
|
||||
- [शुरुआती के लिए जनरेटिव एआई .NET](https://github.com/microsoft/Generative-AI-for-beginners-dotnet)
|
||||
- [जावास्क्रिप्ट के साथ जनरेटिव एआई](https://github.com/microsoft/generative-ai-with-javascript)
|
||||
- [जावा के साथ जनरेटिव एआई](https://github.com/microsoft/Generative-AI-for-beginners-java)
|
||||
- [शुरुआती के लिए एआई](https://aka.ms/ai-beginners)
|
||||
- [शुरुआती के लिए डेटा साइंस](https://aka.ms/datascience-beginners)
|
||||
- [शुरुआती के लिए मशीन लर्निंग](https://aka.ms/ml-beginners)
|
||||
- [शुरुआती के लिए साइबर सुरक्षा](https://github.com/microsoft/Security-101)
|
||||
|
|
@ -152,5 +163,7 @@ CO_OP_TRANSLATOR_METADATA:
|
|||
- [C#/.NET डेवलपर्स के लिए GitHub Copilot में महारत हासिल करना](https://github.com/microsoft/mastering-github-copilot-for-dotnet-csharp-developers)
|
||||
- [अपना खुद का Copilot एडवेंचर चुनें](https://github.com/microsoft/CopilotAdventures)
|
||||
|
||||
---
|
||||
|
||||
**अस्वीकरण**:
|
||||
यह दस्तावेज़ AI अनुवाद सेवा [Co-op Translator](https://github.com/Azure/co-op-translator) का उपयोग करके अनुवादित किया गया है। जबकि हम सटीकता के लिए प्रयासरत हैं, कृपया ध्यान दें कि स्वचालित अनुवाद में त्रुटियां या अशुद्धियां हो सकती हैं। मूल भाषा में उपलब्ध मूल दस्तावेज़ को आधिकारिक स्रोत माना जाना चाहिए। महत्वपूर्ण जानकारी के लिए, पेशेवर मानव अनुवाद की सिफारिश की जाती है। इस अनुवाद के उपयोग से उत्पन्न किसी भी गलतफहमी या गलत व्याख्या के लिए हम उत्तरदायी नहीं हैं।
|
||||
यह दस्तावेज़ AI अनुवाद सेवा [Co-op Translator](https://github.com/Azure/co-op-translator) का उपयोग करके अनुवादित किया गया है। जबकि हम सटीकता के लिए प्रयासरत हैं, कृपया ध्यान दें कि स्वचालित अनुवाद में त्रुटियां या अशुद्धियां हो सकती हैं। मूल भाषा में उपलब्ध मूल दस्तावेज़ को प्रामाणिक स्रोत माना जाना चाहिए। महत्वपूर्ण जानकारी के लिए, पेशेवर मानव अनुवाद की सिफारिश की जाती है। इस अनुवाद के उपयोग से उत्पन्न किसी भी गलतफहमी या गलत व्याख्या के लिए हम उत्तरदायी नहीं हैं।
|
||||
|
|
@ -0,0 +1,478 @@
|
|||
{
|
||||
"cells": [
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {
|
||||
"collapsed": true
|
||||
},
|
||||
"source": [
|
||||
"# पशु विशेषज्ञ प्रणाली को लागू करना\n",
|
||||
"\n",
|
||||
"[AI for Beginners Curriculum](http://github.com/microsoft/ai-for-beginners) से एक उदाहरण।\n",
|
||||
"\n",
|
||||
"इस उदाहरण में, हम एक सरल ज्ञान-आधारित प्रणाली को लागू करेंगे जो कुछ शारीरिक विशेषताओं के आधार पर किसी पशु का निर्धारण करती है। इस प्रणाली को निम्नलिखित AND-OR पेड़ द्वारा दर्शाया जा सकता है (यह पूरे पेड़ का एक हिस्सा है, हम आसानी से इसमें और नियम जोड़ सकते हैं):\n",
|
||||
"\n",
|
||||
"\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## हमारा खुद का एक्सपर्ट सिस्टम शेल बैकवर्ड इंफरेंस के साथ\n",
|
||||
"\n",
|
||||
"आइए प्रोडक्शन रूल्स के आधार पर नॉलेज रिप्रेजेंटेशन के लिए एक सरल भाषा को परिभाषित करने की कोशिश करें। हम नियमों को परिभाषित करने के लिए Python क्लासेस का उपयोग करेंगे। मुख्य रूप से 3 प्रकार की क्लासेस होंगी:\n",
|
||||
"* `Ask` एक प्रश्न को दर्शाता है जिसे उपयोगकर्ता से पूछा जाना है। इसमें संभावित उत्तरों का सेट होता है।\n",
|
||||
"* `If` एक नियम को दर्शाता है, और यह केवल नियम की सामग्री को संग्रहीत करने के लिए एक सिंटैक्टिक शॉर्टकट है।\n",
|
||||
"* `AND`/`OR` क्लासेस हैं जो ट्री की AND/OR शाखाओं को दर्शाती हैं। ये केवल अंदर दिए गए तर्कों की सूची को संग्रहीत करती हैं। कोड को सरल बनाने के लिए, सभी कार्यक्षमता पैरेंट क्लास `Content` में परिभाषित की गई है।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 1,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"class Ask():\n",
|
||||
" def __init__(self,choices=['y','n']):\n",
|
||||
" self.choices = choices\n",
|
||||
" def ask(self):\n",
|
||||
" if max([len(x) for x in self.choices])>1:\n",
|
||||
" for i,x in enumerate(self.choices):\n",
|
||||
" print(\"{0}. {1}\".format(i,x),flush=True)\n",
|
||||
" x = int(input())\n",
|
||||
" return self.choices[x]\n",
|
||||
" else:\n",
|
||||
" print(\"/\".join(self.choices),flush=True)\n",
|
||||
" return input()\n",
|
||||
"\n",
|
||||
"class Content():\n",
|
||||
" def __init__(self,x):\n",
|
||||
" self.x=x\n",
|
||||
" \n",
|
||||
"class If(Content):\n",
|
||||
" pass\n",
|
||||
"\n",
|
||||
"class AND(Content):\n",
|
||||
" pass\n",
|
||||
"\n",
|
||||
"class OR(Content):\n",
|
||||
" pass"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"हमारे सिस्टम में, कार्यशील स्मृति में **तथ्यों** की सूची **गुण-वैल्यू जोड़ों** के रूप में होगी। ज्ञान आधार को एक बड़े शब्दकोश के रूप में परिभाषित किया जा सकता है जो क्रियाओं (नए तथ्य जो कार्यशील स्मृति में डाले जाने चाहिए) को शर्तों से जोड़ता है, जो AND-OR अभिव्यक्तियों के रूप में व्यक्त किए जाते हैं। साथ ही, कुछ तथ्यों को `पूछा` जा सकता है।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 2,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"rules = {\n",
|
||||
" 'default': Ask(['y','n']),\n",
|
||||
" 'color' : Ask(['red-brown','black and white','other']),\n",
|
||||
" 'pattern' : Ask(['dark stripes','dark spots']),\n",
|
||||
" 'mammal': If(OR(['hair','gives milk'])),\n",
|
||||
" 'carnivor': If(OR([AND(['sharp teeth','claws','forward-looking eyes']),'eats meat'])),\n",
|
||||
" 'ungulate': If(['mammal',OR(['has hooves','chews cud'])]),\n",
|
||||
" 'bird': If(OR(['feathers',AND(['flies','lies eggs'])])),\n",
|
||||
" 'animal:monkey' : If(['mammal','carnivor','color:red-brown','pattern:dark spots']),\n",
|
||||
" 'animal:tiger' : If(['mammal','carnivor','color:red-brown','pattern:dark stripes']),\n",
|
||||
" 'animal:giraffe' : If(['ungulate','long neck','long legs','pattern:dark spots']),\n",
|
||||
" 'animal:zebra' : If(['ungulate','pattern:dark stripes']),\n",
|
||||
" 'animal:ostrich' : If(['bird','long nech','color:black and white','cannot fly']),\n",
|
||||
" 'animal:pinguin' : If(['bird','swims','color:black and white','cannot fly']),\n",
|
||||
" 'animal:albatross' : If(['bird','flies well'])\n",
|
||||
"}"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"पिछड़े निष्कर्षण (backward inference) को करने के लिए, हम `Knowledgebase` क्लास को परिभाषित करेंगे। इसमें शामिल होंगे:\n",
|
||||
"* कार्यशील `memory` - एक डिक्शनरी जो गुणों (attributes) को उनके मानों (values) से जोड़ती है।\n",
|
||||
"* Knowledgebase के `rules` - जैसा कि ऊपर परिभाषित किया गया है।\n",
|
||||
"\n",
|
||||
"दो मुख्य विधियाँ (methods) हैं:\n",
|
||||
"* `get` - किसी गुण का मान प्राप्त करने के लिए, यदि आवश्यक हो तो निष्कर्षण (inference) करते हुए। उदाहरण के लिए, `get('color')` किसी रंग के स्लॉट का मान प्राप्त करेगा (यह आवश्यक होने पर पूछेगा और कार्यशील मेमोरी में बाद के उपयोग के लिए मान को संग्रहीत करेगा)। यदि हम `get('color:blue')` पूछते हैं, तो यह रंग पूछेगा और फिर `y`/`n` मान लौटाएगा, जो रंग पर निर्भर करेगा।\n",
|
||||
"* `eval` - वास्तविक निष्कर्षण करता है, यानी AND/OR ट्री को पार करता है, उप-लक्ष्यों (sub-goals) का मूल्यांकन करता है, आदि।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 3,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"class KnowledgeBase():\n",
|
||||
" def __init__(self,rules):\n",
|
||||
" self.rules = rules\n",
|
||||
" self.memory = {}\n",
|
||||
" \n",
|
||||
" def get(self,name):\n",
|
||||
" if ':' in name:\n",
|
||||
" k,v = name.split(':')\n",
|
||||
" vv = self.get(k)\n",
|
||||
" return 'y' if v==vv else 'n'\n",
|
||||
" if name in self.memory.keys():\n",
|
||||
" return self.memory[name]\n",
|
||||
" for fld in self.rules.keys():\n",
|
||||
" if fld==name or fld.startswith(name+\":\"):\n",
|
||||
" # print(\" + proving {}\".format(fld))\n",
|
||||
" value = 'y' if fld==name else fld.split(':')[1]\n",
|
||||
" res = self.eval(self.rules[fld],field=name)\n",
|
||||
" if res!='y' and res!='n' and value=='y':\n",
|
||||
" self.memory[name] = res\n",
|
||||
" return res\n",
|
||||
" if res=='y':\n",
|
||||
" self.memory[name] = value\n",
|
||||
" return value\n",
|
||||
" # field is not found, using default\n",
|
||||
" res = self.eval(self.rules['default'],field=name)\n",
|
||||
" self.memory[name]=res\n",
|
||||
" return res\n",
|
||||
" \n",
|
||||
" def eval(self,expr,field=None):\n",
|
||||
" # print(\" + eval {}\".format(expr))\n",
|
||||
" if isinstance(expr,Ask):\n",
|
||||
" print(field)\n",
|
||||
" return expr.ask()\n",
|
||||
" elif isinstance(expr,If):\n",
|
||||
" return self.eval(expr.x)\n",
|
||||
" elif isinstance(expr,AND) or isinstance(expr,list):\n",
|
||||
" expr = expr.x if isinstance(expr,AND) else expr\n",
|
||||
" for x in expr:\n",
|
||||
" if self.eval(x)=='n':\n",
|
||||
" return 'n'\n",
|
||||
" return 'y'\n",
|
||||
" elif isinstance(expr,OR):\n",
|
||||
" for x in expr.x:\n",
|
||||
" if self.eval(x)=='y':\n",
|
||||
" return 'y'\n",
|
||||
" return 'n'\n",
|
||||
" elif isinstance(expr,str):\n",
|
||||
" return self.get(expr)\n",
|
||||
" else:\n",
|
||||
" print(\"Unknown expr: {}\".format(expr))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"अब चलिए हमारे पशु ज्ञानकोष को परिभाषित करते हैं और परामर्श करते हैं। ध्यान दें कि यह कॉल आपसे प्रश्न पूछेगा। आप `y`/`n` टाइप करके हां-ना वाले प्रश्नों का उत्तर दे सकते हैं, या लंबे बहुविकल्पीय उत्तरों वाले प्रश्नों के लिए संख्या (0..N) निर्दिष्ट कर सकते हैं।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 4,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"hair\n",
|
||||
"y/n\n",
|
||||
"sharp teeth\n",
|
||||
"y/n\n",
|
||||
"claws\n",
|
||||
"y/n\n",
|
||||
"forward-looking eyes\n",
|
||||
"y/n\n",
|
||||
"color\n",
|
||||
"0. red-brown\n",
|
||||
"1. black and white\n",
|
||||
"2. other\n",
|
||||
"has hooves\n",
|
||||
"y/n\n",
|
||||
"long neck\n",
|
||||
"y/n\n",
|
||||
"long legs\n",
|
||||
"y/n\n",
|
||||
"pattern\n",
|
||||
"0. dark stripes\n",
|
||||
"1. dark spots\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"'giraffe'"
|
||||
]
|
||||
},
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"kb = KnowledgeBase(rules)\n",
|
||||
"kb.get('animal')"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## PyKnow का उपयोग करके Forward Inference\n",
|
||||
"\n",
|
||||
"अगले उदाहरण में, हम ज्ञान प्रतिनिधित्व के लिए एक लाइब्रेरी [PyKnow](https://github.com/buguroo/pyknow/) का उपयोग करके forward inference को लागू करने की कोशिश करेंगे। **PyKnow** एक लाइब्रेरी है जो Python में forward inference सिस्टम बनाने के लिए डिज़ाइन की गई है, और यह पुराने क्लासिकल सिस्टम [CLIPS](http://www.clipsrules.net/index.html) के समान है।\n",
|
||||
"\n",
|
||||
"हम forward chaining को खुद भी बिना ज्यादा समस्या के लागू कर सकते थे, लेकिन साधारण (naive) कार्यान्वयन आमतौर पर बहुत प्रभावी नहीं होते। अधिक प्रभावी नियम मिलान (rule matching) के लिए एक विशेष एल्गोरिदम [Rete](https://en.wikipedia.org/wiki/Rete_algorithm) का उपयोग किया जाता है।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 5,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Collecting git+https://github.com/buguroo/pyknow/\n",
|
||||
" Cloning https://github.com/buguroo/pyknow/ to /tmp/pip-req-build-3cqeulyl\n",
|
||||
" Running command git clone --filter=blob:none --quiet https://github.com/buguroo/pyknow/ /tmp/pip-req-build-3cqeulyl\n",
|
||||
" Resolved https://github.com/buguroo/pyknow/ to commit 48818336f2e9a126f1964f2d8dc22d37ff800fe8\n",
|
||||
" Preparing metadata (setup.py) ... \u001b[?25ldone\n",
|
||||
"\u001b[?25hCollecting frozendict==1.2\n",
|
||||
" Using cached frozendict-1.2.tar.gz (2.6 kB)\n",
|
||||
" Preparing metadata (setup.py) ... \u001b[?25ldone\n",
|
||||
"\u001b[?25hCollecting schema==0.6.7\n",
|
||||
" Using cached schema-0.6.7-py2.py3-none-any.whl (14 kB)\n",
|
||||
"Building wheels for collected packages: pyknow, frozendict\n",
|
||||
" Building wheel for pyknow (setup.py) ... \u001b[?25ldone\n",
|
||||
"\u001b[?25h Created wheel for pyknow: filename=pyknow-1.7.0-py3-none-any.whl size=34228 sha256=b7de5b09292c4007667c72f69b98d5a1b5f7324ff15f9dd8e077c3d5f7aade42\n",
|
||||
" Stored in directory: /tmp/pip-ephem-wheel-cache-k7jpave7/wheels/81/1a/d3/f6c15dbe1955598a37755215f2a10449e7418500d7bd4b9508\n",
|
||||
" Building wheel for frozendict (setup.py) ... \u001b[?25ldone\n",
|
||||
"\u001b[?25h Created wheel for frozendict: filename=frozendict-1.2-py3-none-any.whl size=3148 sha256=2863d55c240d2409cddf05ccfe600591f8478681549fc97555c47c90dc6bb160\n",
|
||||
" Stored in directory: /home/rg/.cache/pip/wheels/49/ac/f8/cb8120244e710bdb479c86198b03c7b08c3c2d3d2bf448fd6e\n",
|
||||
"Successfully built pyknow frozendict\n",
|
||||
"Installing collected packages: schema, frozendict, pyknow\n",
|
||||
"Successfully installed frozendict-1.2 pyknow-1.7.0 schema-0.6.7\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"import sys\n",
|
||||
"!{sys.executable} -m pip install git+https://github.com/buguroo/pyknow/"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 13,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"from pyknow import *\n",
|
||||
"#import pyknow"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"हम अपने सिस्टम को `KnowledgeEngine` का सबक्लास बनाकर एक क्लास के रूप में परिभाषित करेंगे। प्रत्येक नियम को `@Rule` एनोटेशन के साथ एक अलग फ़ंक्शन द्वारा परिभाषित किया जाता है, जो यह निर्दिष्ट करता है कि नियम कब सक्रिय होना चाहिए। नियम के अंदर, हम `declare` फ़ंक्शन का उपयोग करके नए तथ्य जोड़ सकते हैं, और उन तथ्यों को जोड़ने से आगे की अनुमान इंजन द्वारा कुछ और नियमों को कॉल किया जाएगा।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 14,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"class Animals(KnowledgeEngine):\n",
|
||||
" @Rule(OR(\n",
|
||||
" AND(Fact('sharp teeth'),Fact('claws'),Fact('forward looking eyes')),\n",
|
||||
" Fact('eats meat')))\n",
|
||||
" def cornivor(self):\n",
|
||||
" self.declare(Fact('carnivor'))\n",
|
||||
" \n",
|
||||
" @Rule(OR(Fact('hair'),Fact('gives milk')))\n",
|
||||
" def mammal(self):\n",
|
||||
" self.declare(Fact('mammal'))\n",
|
||||
"\n",
|
||||
" @Rule(Fact('mammal'),\n",
|
||||
" OR(Fact('has hooves'),Fact('chews cud')))\n",
|
||||
" def hooves(self):\n",
|
||||
" self.declare('ungulate')\n",
|
||||
" \n",
|
||||
" @Rule(OR(Fact('feathers'),AND(Fact('flies'),Fact('lays eggs'))))\n",
|
||||
" def bird(self):\n",
|
||||
" self.declare('bird')\n",
|
||||
" \n",
|
||||
" @Rule(Fact('mammal'),Fact('carnivor'),\n",
|
||||
" Fact(color='red-brown'),\n",
|
||||
" Fact(pattern='dark spots'))\n",
|
||||
" def monkey(self):\n",
|
||||
" self.declare(Fact(animal='monkey'))\n",
|
||||
"\n",
|
||||
" @Rule(Fact('mammal'),Fact('carnivor'),\n",
|
||||
" Fact(color='red-brown'),\n",
|
||||
" Fact(pattern='dark stripes'))\n",
|
||||
" def tiger(self):\n",
|
||||
" self.declare(Fact(animal='tiger'))\n",
|
||||
"\n",
|
||||
" @Rule(Fact('ungulate'),\n",
|
||||
" Fact('long neck'),\n",
|
||||
" Fact('long legs'),\n",
|
||||
" Fact(pattern='dark spots'))\n",
|
||||
" def giraffe(self):\n",
|
||||
" self.declare(Fact(animal='giraffe'))\n",
|
||||
"\n",
|
||||
" @Rule(Fact('ungulate'),\n",
|
||||
" Fact(pattern='dark stripes'))\n",
|
||||
" def zebra(self):\n",
|
||||
" self.declare(Fact(animal='zebra'))\n",
|
||||
"\n",
|
||||
" @Rule(Fact('bird'),\n",
|
||||
" Fact('long neck'),\n",
|
||||
" Fact('cannot fly'),\n",
|
||||
" Fact(color='black and white'))\n",
|
||||
" def straus(self):\n",
|
||||
" self.declare(Fact(animal='ostrich'))\n",
|
||||
"\n",
|
||||
" @Rule(Fact('bird'),\n",
|
||||
" Fact('swims'),\n",
|
||||
" Fact('cannot fly'),\n",
|
||||
" Fact(color='black and white'))\n",
|
||||
" def pinguin(self):\n",
|
||||
" self.declare(Fact(animal='pinguin'))\n",
|
||||
"\n",
|
||||
" @Rule(Fact('bird'),\n",
|
||||
" Fact('flies well'))\n",
|
||||
" def albatros(self):\n",
|
||||
" self.declare(Fact(animal='albatross'))\n",
|
||||
" \n",
|
||||
" @Rule(Fact(animal=MATCH.a))\n",
|
||||
" def print_result(self,a):\n",
|
||||
" print('Animal is {}'.format(a))\n",
|
||||
" \n",
|
||||
" def factz(self,l):\n",
|
||||
" for x in l:\n",
|
||||
" self.declare(x)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"एक बार जब हम एक ज्ञान आधार परिभाषित कर लेते हैं, तो हम अपनी कार्यशील स्मृति को कुछ प्रारंभिक तथ्यों से भरते हैं, और फिर `run()` विधि को कॉल करते हैं ताकि निष्कर्षण किया जा सके। आप परिणामस्वरूप देख सकते हैं कि नए निष्कर्षित तथ्य कार्यशील स्मृति में जोड़े जाते हैं, जिसमें जानवर के बारे में अंतिम तथ्य भी शामिल है (यदि हमने सभी प्रारंभिक तथ्यों को सही तरीके से सेट किया है)।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 15,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Animal is tiger\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"FactList([(0, InitialFact()),\n",
|
||||
" (1, Fact(color='red-brown')),\n",
|
||||
" (2, Fact(pattern='dark stripes')),\n",
|
||||
" (3, Fact('sharp teeth')),\n",
|
||||
" (4, Fact('claws')),\n",
|
||||
" (5, Fact('forward looking eyes')),\n",
|
||||
" (6, Fact('gives milk')),\n",
|
||||
" (7, Fact('mammal')),\n",
|
||||
" (8, Fact('carnivor')),\n",
|
||||
" (9, Fact(animal='tiger'))])"
|
||||
]
|
||||
},
|
||||
"execution_count": 15,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"ex1 = Animals()\n",
|
||||
"ex1.reset()\n",
|
||||
"ex1.factz([\n",
|
||||
" Fact(color='red-brown'),\n",
|
||||
" Fact(pattern='dark stripes'),\n",
|
||||
" Fact('sharp teeth'),\n",
|
||||
" Fact('claws'),\n",
|
||||
" Fact('forward looking eyes'),\n",
|
||||
" Fact('gives milk')])\n",
|
||||
"ex1.run()\n",
|
||||
"ex1.facts"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": []
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"\n---\n\n**अस्वीकरण**: \nयह दस्तावेज़ AI अनुवाद सेवा [Co-op Translator](https://github.com/Azure/co-op-translator) का उपयोग करके अनुवादित किया गया है। जबकि हम सटीकता सुनिश्चित करने का प्रयास करते हैं, कृपया ध्यान दें कि स्वचालित अनुवाद में त्रुटियां या अशुद्धियां हो सकती हैं। मूल भाषा में उपलब्ध मूल दस्तावेज़ को प्रामाणिक स्रोत माना जाना चाहिए। महत्वपूर्ण जानकारी के लिए, पेशेवर मानव अनुवाद की सिफारिश की जाती है। इस अनुवाद के उपयोग से उत्पन्न किसी भी गलतफहमी या गलत व्याख्या के लिए हम उत्तरदायी नहीं हैं।\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
"kernelspec": {
|
||||
"display_name": "Python 3.7.4 64-bit (conda)",
|
||||
"metadata": {
|
||||
"interpreter": {
|
||||
"hash": "86193a1ab0ba47eac1c69c1756090baa3b420b3eea7d4aafab8b85f8b312f0c5"
|
||||
}
|
||||
},
|
||||
"name": "python3"
|
||||
},
|
||||
"language_info": {
|
||||
"codemirror_mode": {
|
||||
"name": "ipython",
|
||||
"version": 3
|
||||
},
|
||||
"file_extension": ".py",
|
||||
"mimetype": "text/x-python",
|
||||
"name": "python",
|
||||
"nbconvert_exporter": "python",
|
||||
"pygments_lexer": "ipython3",
|
||||
"version": "3.11.2"
|
||||
},
|
||||
"coopTranslator": {
|
||||
"original_hash": "ab2bd97b0453415b89a469284609a8ce",
|
||||
"translation_date": "2025-08-31T14:55:07+00:00",
|
||||
"source_file": "lessons/2-Symbolic/Animals.ipynb",
|
||||
"language_code": "hi"
|
||||
}
|
||||
},
|
||||
"nbformat": 4,
|
||||
"nbformat_minor": 2
|
||||
}
|
||||
|
|
@ -0,0 +1,593 @@
|
|||
{
|
||||
"cells": [
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {
|
||||
"collapsed": true
|
||||
},
|
||||
"source": [
|
||||
"# परिवार संबंध ऑंटोलॉजी\n",
|
||||
"\n",
|
||||
"यह उदाहरण [AI for Beginners Curriculum](http://github.com/microsoft/ai-for-beginners) का हिस्सा है, और इसे [इस ब्लॉग पोस्ट](https://habr.com/post/270857/) से प्रेरित होकर बनाया गया है।\n",
|
||||
"\n",
|
||||
"मुझे हमेशा परिवार में लोगों के बीच विभिन्न संबंधों को याद रखना मुश्किल लगता है। इस उदाहरण में, हम एक ऑंटोलॉजी लेंगे जो परिवार संबंधों को परिभाषित करती है, और वास्तविक वंशावली वृक्ष का उपयोग करेंगे, और फिर दिखाएंगे कि हम स्वचालित निष्कर्षण करके सभी रिश्तेदारों को कैसे खोज सकते हैं।\n",
|
||||
"\n",
|
||||
"### वंशावली वृक्ष प्राप्त करना\n",
|
||||
"\n",
|
||||
"उदाहरण के रूप में, हम [रोमानोव ज़ार परिवार](https://en.wikipedia.org/wiki/House_of_Romanov) का वंशावली वृक्ष लेंगे। परिवार संबंधों का वर्णन करने के लिए सबसे सामान्य प्रारूप [GEDCOM](https://en.wikipedia.org/wiki/GEDCOM) है। हम GEDCOM प्रारूप में रोमानोव परिवार का वंशावली वृक्ष लेंगे:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 1,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"0 HEAD\n",
|
||||
"1 CHAR UTF8\n",
|
||||
"1 GEDC\n",
|
||||
"2 VERS 5.5\n",
|
||||
"0 @0@ INDI\n",
|
||||
"1 NAME Mihail Fedorovich /Romanov/\n",
|
||||
"1 SEX M\n",
|
||||
"1 BIRT\n",
|
||||
"2 DATE 1613\n",
|
||||
"1 DEAT \n",
|
||||
"2 DATE 1645\n",
|
||||
"1 FAMS @41@\n",
|
||||
"0 @1@ INDI\n",
|
||||
"1 NAME Evdokija Lukjanovna /Streshneva/\n",
|
||||
"1 SEX F\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"!head -15 data/tsars.ged"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"GEDCOM फ़ाइल का उपयोग करने के लिए, हम `python-gedcom` लाइब्रेरी का उपयोग कर सकते हैं:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 2,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Collecting python-gedcom\n",
|
||||
" Downloading python_gedcom-1.0.0-py2.py3-none-any.whl (35 kB)\n",
|
||||
"Installing collected packages: python-gedcom\n",
|
||||
"Successfully installed python-gedcom-1.0.0\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"import sys\n",
|
||||
"!{sys.executable} -m pip install python-gedcom"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"यह पुस्तकालय फ़ाइल पार्सिंग से संबंधित कुछ तकनीकी समस्याओं को दूर करता है, लेकिन यह अभी भी हमें पेड़ में सभी व्यक्तियों और परिवारों तक काफी निम्न-स्तरीय पहुंच प्रदान करता है। यहां बताया गया है कि हम फ़ाइल को कैसे पार्स कर सकते हैं और सभी व्यक्तियों की सूची दिखा सकते हैं:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 3,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"from gedcom.parser import Parser\n",
|
||||
"from gedcom.element.individual import IndividualElement\n",
|
||||
"from gedcom.element.family import FamilyElement\n",
|
||||
"g = Parser()\n",
|
||||
"g.parse_file('data/tsars.ged')"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 4,
|
||||
"metadata": {
|
||||
"scrolled": true,
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"[('@0@', ('Mihail Fedorovich', 'Romanov')),\n",
|
||||
" ('@1@', ('Evdokija Lukjanovna', 'Streshneva')),\n",
|
||||
" ('@2@', ('Aleksej Mihajlovich', 'Romanov')),\n",
|
||||
" ('@3@', ('Marija Ilinichna', 'Miloslavskaja')),\n",
|
||||
" ('@4@', ('Natalja Kirillovna', 'Naryshkina')),\n",
|
||||
" ('@5@', ('Marfa Matveevna', 'Apraksina')),\n",
|
||||
" ('@6@', ('Fedor Alekseevich', 'Romanov')),\n",
|
||||
" ('@7@', ('Sofja Aleksevna', 'Romanova')),\n",
|
||||
" ('@8@', ('Ivan V Alekseevich', 'Romanov')),\n",
|
||||
" ('@9@', ('Praskovja Fedorovna', 'Saltykova')),\n",
|
||||
" ('@10@', ('Ekaterina Ivanovna', 'Romanova')),\n",
|
||||
" ('@11@', ('Anna Ivanovna', 'Romanova')),\n",
|
||||
" ('@12@', ('Fridrih Vilgelm', 'Kurlandskij')),\n",
|
||||
" ('@13@', ('Karl Leopold', 'Meklenburg-Shverinskij')),\n",
|
||||
" ('@14@', ('Anna Leopoldovna', 'Meklenburg-Shverinskaja')),\n",
|
||||
" ('@15@', ('Anton Ulrih', 'Braunshvejg-Volfenbjuttelskij')),\n",
|
||||
" ('@16@', ('Ivan VI Antonovich', 'Braunshvejg-Volfenbjuttelskij')),\n",
|
||||
" ('@17@', ('Petr I Alekseevich', 'Romanov')),\n",
|
||||
" ('@18@', ('Evdokija Fedorovna', 'Lopuhina')),\n",
|
||||
" ('@19@', ('Ekaterina I Alekseevna', 'Mihajlova')),\n",
|
||||
" ('@20@', ('Aleksej Petrovich', 'Romanov')),\n",
|
||||
" ('@21@', ('Sharlotta Kristina', 'Braunshvejg-Volfenbjuttelskaja')),\n",
|
||||
" ('@22@', ('Petr II Alekseevich', 'Romanov')),\n",
|
||||
" ('@23@', ('Anna Petrovna', 'Romanova')),\n",
|
||||
" ('@24@', ('Elizaveta Petrovna', 'Romanova')),\n",
|
||||
" ('@25@', ('Karl Fridrih', 'Golshtejn-Gottorpskij')),\n",
|
||||
" ('@26@', ('Petr III Fedorovich', 'Romanov')),\n",
|
||||
" ('@27@', ('Ekaterina II', 'Alekseevna')),\n",
|
||||
" ('@28@', ('Pavel I Petrovich', 'Romanov')),\n",
|
||||
" ('@29@', ('Natalja Alekseevna', 'Gessen-Darmshtadskaja')),\n",
|
||||
" ('@30@', ('Marija Fedorovna', 'Vjurtembergskaja')),\n",
|
||||
" ('@31@', ('Aleksandr I Pavlovich', 'Romanov')),\n",
|
||||
" ('@32@', ('Elizaveta Alekseevna', 'Baden-Durlahskaja')),\n",
|
||||
" ('@33@', ('Nikolaj I Pavlovich', 'Romanov')),\n",
|
||||
" ('@34@', ('Aleksandra Fedorovna', 'Prusskaja')),\n",
|
||||
" ('@35@', ('Aleksandr II Nikolaevich', 'Romanov')),\n",
|
||||
" ('@36@', ('Marija Aleksandrovna', 'Gessenskaja')),\n",
|
||||
" ('@37@', ('Aleksandr III Aleksandrovich', 'Romanov')),\n",
|
||||
" ('@38@', ('Marija Fedorovna', 'Datskaja')),\n",
|
||||
" ('@39@', ('Nikolaj II Aleksandrovich', 'Romanov')),\n",
|
||||
" ('@40@', ('Aleksandra Fedorovna', 'Gessenskaja'))]"
|
||||
]
|
||||
},
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"d = g.get_element_dictionary()\n",
|
||||
"[ (k,v.get_name()) for k,v in d.items() if isinstance(v,IndividualElement)]"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"यहां बताया गया है कि हम परिवारों के बारे में जानकारी कैसे प्राप्त कर सकते हैं। ध्यान दें कि यह हमें **पहचानकर्ताओं** की एक सूची देता है, और यदि हमें अधिक स्पष्टता चाहिए तो हमें उन्हें नामों में बदलने की आवश्यकता है:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 5,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"[('@41@', ['@0@', '@1@', '@2@']),\n",
|
||||
" ('@42@', ['@2@', '@3@', '@6@', '@7@', '@8@']),\n",
|
||||
" ('@43@', ['@8@', '@9@', '@10@', '@11@']),\n",
|
||||
" ('@44@', ['@13@', '@10@', '@14@']),\n",
|
||||
" ('@45@', ['@15@', '@14@', '@16@']),\n",
|
||||
" ('@46@', ['@2@', '@4@', '@17@']),\n",
|
||||
" ('@47@', ['@17@', '@18@', '@20@']),\n",
|
||||
" ('@48@', ['@20@', '@21@', '@22@']),\n",
|
||||
" ('@49@', ['@17@', '@19@', '@23@', '@24@']),\n",
|
||||
" ('@50@', ['@25@', '@23@', '@26@']),\n",
|
||||
" ('@51@', ['@26@', '@27@', '@28@']),\n",
|
||||
" ('@52@', ['@28@', '@30@', '@31@', '@33@']),\n",
|
||||
" ('@53@', ['@33@', '@34@', '@35@']),\n",
|
||||
" ('@54@', ['@35@', '@36@', '@37@']),\n",
|
||||
" ('@55@', ['@37@', '@38@', '@39@'])]"
|
||||
]
|
||||
},
|
||||
"execution_count": 5,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"d = g.get_element_dictionary()\n",
|
||||
"[ (k,[x.get_value() for x in v.get_child_elements()]) for k,v in d.items() if isinstance(v,FamilyElement)]"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"### परिवार ऑंटोलॉजी प्राप्त करना\n",
|
||||
"\n",
|
||||
"अब, आइए [परिवार ऑंटोलॉजी](https://raw.githubusercontent.com/blokhin/genealogical-trees/master/data/header.ttl) पर नज़र डालें, जिसे सेमांटिक वेब ट्रिपलेट्स के सेट के रूप में परिभाषित किया गया है। इस ऑंटोलॉजी में `isUncleOf`, `isCousinOf` और कई अन्य संबंधों को परिभाषित किया गया है। ये सभी संबंध मूल प्रेडिकेट्स `isMotherOf`, `isFatherOf`, `isBrotherOf` और `isSisterOf` के आधार पर परिभाषित किए गए हैं। हम ऑंटोलॉजी का उपयोग करके स्वचालित तर्क के माध्यम से अन्य सभी संबंधों को निकालेंगे।\n",
|
||||
"\n",
|
||||
"यहाँ `isAuntOf` प्रॉपर्टी की एक नमूना परिभाषा दी गई है, जिसे `isSisterOf` और `isParentOf` के संयोजन के रूप में परिभाषित किया गया है (*आंटी किसी के माता-पिता की बहन होती है*)।\n",
|
||||
"\n",
|
||||
"```\n",
|
||||
"fhkb:isAuntOf a owl:ObjectProperty ;\n",
|
||||
" rdfs:domain fhkb:Woman ;\n",
|
||||
" rdfs:range fhkb:Person ;\n",
|
||||
" owl:propertyChainAxiom ( fhkb:isSisterOf fhkb:isParentOf ) .\n",
|
||||
"```\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 6,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"@prefix fhkb: <http://www.example.com/genealogy.owl#> .\n",
|
||||
"@prefix owl: <http://www.w3.org/2002/07/owl#> .\n",
|
||||
"@prefix rdf: <http://www.w3.org/1999/02/22-rdf-syntax-ns#> .\n",
|
||||
"@prefix rdfs: <http://www.w3.org/2000/01/rdf-schema#> .\n",
|
||||
"@prefix xml: <http://www.w3.org/XML/1998/namespace> .\n",
|
||||
"@prefix xsd: <http://www.w3.org/2001/XMLSchema#> .\n",
|
||||
"\n",
|
||||
"<http://www.example.com/genealogy.owl#> a owl:Ontology .\n",
|
||||
"\n",
|
||||
"fhkb:DomainEntity a owl:Class .\n",
|
||||
"\n",
|
||||
"fhkb:Man a owl:Class ;\n",
|
||||
" owl:equivalentClass [ a owl:Class ;\n",
|
||||
" owl:intersectionOf ( fhkb:Person [ a owl:Restriction ;\n",
|
||||
" owl:onProperty fhkb:hasSex ;\n",
|
||||
" owl:someValuesFrom fhkb:Male ] ) ] .\n",
|
||||
"\n",
|
||||
"fhkb:Woman a owl:Class ;\n",
|
||||
" owl:equivalentClass [ a owl:Class ;\n",
|
||||
" owl:intersectionOf ( fhkb:Person [ a owl:Restriction ;\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"!head -20 data/onto.ttl"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"### अनुमान के लिए ओंटोलॉजी बनाना\n",
|
||||
"\n",
|
||||
"सरलता के लिए, हम एक ओंटोलॉजी फ़ाइल बनाएंगे जिसमें परिवार ओंटोलॉजी से मूल नियम और हमारे GEDCOM फ़ाइल से व्यक्तियों के बारे में तथ्य शामिल होंगे। हम GEDCOM फ़ाइल का विश्लेषण करेंगे, परिवारों और व्यक्तियों की जानकारी निकालेंगे, और उन्हें ट्रिपलेट्स में बदलेंगे।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 7,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"!cp data/onto.ttl .\n",
|
||||
"\n",
|
||||
"gedcom_dict = g.get_element_dictionary()\n",
|
||||
"individuals, marriages = {}, {}\n",
|
||||
"\n",
|
||||
"def term2id(el):\n",
|
||||
" return \"i\" + el.get_pointer().replace('@', '').lower()\n",
|
||||
"\n",
|
||||
"out = open(\"onto.ttl\",\"a\")\n",
|
||||
"\n",
|
||||
"for k, v in gedcom_dict.items():\n",
|
||||
" if isinstance(v,IndividualElement):\n",
|
||||
" children, siblings = set(), set()\n",
|
||||
" idx = term2id(v)\n",
|
||||
"\n",
|
||||
" title = v.get_name()[0] + \" \" + v.get_name()[1]\n",
|
||||
" title = title.replace('\"', '').replace('[', '').replace(']', '').replace('(', '').replace(')', '').strip()\n",
|
||||
"\n",
|
||||
" own_families = g.get_families(v, 'FAMS')\n",
|
||||
" for fam in own_families:\n",
|
||||
" children |= set(term2id(i) for i in g.get_family_members(fam, \"CHIL\"))\n",
|
||||
"\n",
|
||||
" parent_families = g.get_families(v, 'FAMC')\n",
|
||||
" if len(parent_families):\n",
|
||||
" for member in g.get_family_members(parent_families[0], \"CHIL\"): # NB adoptive families i.e len(parent_families)>1 are not considered (TODO?)\n",
|
||||
" if member.get_pointer() == v.get_pointer():\n",
|
||||
" continue\n",
|
||||
" siblings.add(term2id(member))\n",
|
||||
"\n",
|
||||
" if idx in individuals:\n",
|
||||
" children |= individuals[idx].get('children', set())\n",
|
||||
" siblings |= individuals[idx].get('siblings', set())\n",
|
||||
" individuals[idx] = {'sex': v.get_gender().lower(), 'children': children, 'siblings': siblings, 'title': title}\n",
|
||||
"\n",
|
||||
" elif isinstance(v,FamilyElement):\n",
|
||||
" wife, husb, children = None, None, set()\n",
|
||||
" children = set(term2id(i) for i in g.get_family_members(v, \"CHIL\"))\n",
|
||||
"\n",
|
||||
" try:\n",
|
||||
" wife = g.get_family_members(v, \"WIFE\")[0]\n",
|
||||
" wife = term2id(wife)\n",
|
||||
" if wife in individuals: individuals[wife]['children'] |= children\n",
|
||||
" else: individuals[wife] = {'children': children}\n",
|
||||
" except IndexError: pass\n",
|
||||
" try:\n",
|
||||
" husb = g.get_family_members(v, \"HUSB\")[0]\n",
|
||||
" husb = term2id(husb)\n",
|
||||
" if husb in individuals: individuals[husb]['children'] |= children\n",
|
||||
" else: individuals[husb] = {'children': children}\n",
|
||||
" except IndexError: pass\n",
|
||||
"\n",
|
||||
" if wife and husb: marriages[wife + husb] = (term2id(v), wife, husb)\n",
|
||||
"\n",
|
||||
"for idx, val in individuals.items():\n",
|
||||
" added_terms = ''\n",
|
||||
" if val['sex'] == 'f':\n",
|
||||
" parent_predicate, sibl_predicate = \"isMotherOf\", \"isSisterOf\"\n",
|
||||
" else:\n",
|
||||
" parent_predicate, sibl_predicate = \"isFatherOf\", \"isBrotherOf\"\n",
|
||||
" if len(val['children']):\n",
|
||||
" added_terms += \" ;\\n fhkb:\" + parent_predicate + \" \" + \", \".join([\"fhkb:\" + i for i in val['children']])\n",
|
||||
" if len(val['siblings']):\n",
|
||||
" added_terms += \" ;\\n fhkb:\" + sibl_predicate + \" \" + \", \".join([\"fhkb:\" + i for i in val['siblings']])\n",
|
||||
" out.write(\"fhkb:%s a owl:NamedIndividual, owl:Thing%s ;\\n rdfs:label \\\"%s\\\" .\\n\" % (idx, added_terms, val['title']))\n",
|
||||
"\n",
|
||||
"for k, v in marriages.items():\n",
|
||||
" out.write(\"fhkb:%s a owl:NamedIndividual, owl:Thing ;\\n fhkb:hasFemalePartner fhkb:%s ;\\n fhkb:hasMalePartner fhkb:%s .\\n\" % v)\n",
|
||||
"\n",
|
||||
"out.write(\"[] a owl:AllDifferent ;\\n owl:distinctMembers (\")\n",
|
||||
"for idx in individuals.keys():\n",
|
||||
" out.write(\" fhkb:\" + idx)\n",
|
||||
"for k, v in marriages.items():\n",
|
||||
" out.write(\" fhkb:\" + v[0])\n",
|
||||
"out.write(\" ) .\")\n",
|
||||
"out.close()"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 8,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
" fhkb:hasFemalePartner fhkb:i34 ;\n",
|
||||
" fhkb:hasMalePartner fhkb:i33 .\n",
|
||||
"fhkb:i54 a owl:NamedIndividual, owl:Thing ;\n",
|
||||
" fhkb:hasFemalePartner fhkb:i36 ;\n",
|
||||
" fhkb:hasMalePartner fhkb:i35 .\n",
|
||||
"fhkb:i55 a owl:NamedIndividual, owl:Thing ;\n",
|
||||
" fhkb:hasFemalePartner fhkb:i38 ;\n",
|
||||
" fhkb:hasMalePartner fhkb:i37 .\n",
|
||||
"[] a owl:AllDifferent ;\n",
|
||||
" owl:distinctMembers ( fhkb:i0 fhkb:i1 fhkb:i2 fhkb:i3 fhkb:i4 fhkb:i5 fhkb:i6 fhkb:i7 fhkb:i8 fhkb:i9 fhkb:i10 fhkb:i11 fhkb:i12 fhkb:i13 fhkb:i14 fhkb:i15 fhkb:i16 fhkb:i17 fhkb:i18 fhkb:i19 fhkb:i20 fhkb:i21 fhkb:i22 fhkb:i23 fhkb:i24 fhkb:i25 fhkb:i26 fhkb:i27 fhkb:i28 fhkb:i29 fhkb:i30 fhkb:i31 fhkb:i32 fhkb:i33 fhkb:i34 fhkb:i35 fhkb:i36 fhkb:i37 fhkb:i38 fhkb:i39 fhkb:i40 fhkb:i41 fhkb:i42 fhkb:i43 fhkb:i44 fhkb:i45 fhkb:i46 fhkb:i47 fhkb:i48 fhkb:i49 fhkb:i50 fhkb:i51 fhkb:i52 fhkb:i53 fhkb:i54 fhkb:i55 ) ."
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"!tail onto.ttl"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"### अनुमान लगाना \n",
|
||||
"\n",
|
||||
"अब हम इस ऑंटोलॉजी का उपयोग अनुमान लगाने और क्वेरी करने के लिए करना चाहते हैं। हम [RDFLib](https://github.com/RDFLib) का उपयोग करेंगे, जो RDF ग्राफ को विभिन्न प्रारूपों में पढ़ने, क्वेरी करने आदि के लिए एक लाइब्रेरी है। \n",
|
||||
"\n",
|
||||
"तार्किक अनुमान के लिए, हम [OWL-RL](https://github.com/RDFLib/OWL-RL) लाइब्रेरी का उपयोग करेंगे, जो हमें RDF ग्राफ का **Closure** बनाने की अनुमति देती है, यानी सभी संभावित अवधारणाओं और संबंधों को जोड़ना जो अनुमानित किए जा सकते हैं।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 10,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Requirement already satisfied: rdflib in /home/rg/anaconda3/envs/ai4beg/lib/python3.11/site-packages (6.3.2)\n",
|
||||
"Requirement already satisfied: isodate<0.7.0,>=0.6.0 in /home/rg/anaconda3/envs/ai4beg/lib/python3.11/site-packages (from rdflib) (0.6.1)\n",
|
||||
"Requirement already satisfied: pyparsing<4,>=2.1.0 in /home/rg/anaconda3/envs/ai4beg/lib/python3.11/site-packages (from rdflib) (3.0.9)\n",
|
||||
"Requirement already satisfied: six in /home/rg/anaconda3/envs/ai4beg/lib/python3.11/site-packages (from isodate<0.7.0,>=0.6.0->rdflib) (1.16.0)\n",
|
||||
"Collecting git+https://github.com/RDFLib/OWL-RL.git\n",
|
||||
" Cloning https://github.com/RDFLib/OWL-RL.git to /tmp/pip-req-build-lbfzwi3m\n",
|
||||
" Running command git clone --filter=blob:none --quiet https://github.com/RDFLib/OWL-RL.git /tmp/pip-req-build-lbfzwi3m\n",
|
||||
" Resolved https://github.com/RDFLib/OWL-RL.git to commit a77e1791b88b54aace609bc6000aac14c7add4ff\n",
|
||||
" Preparing metadata (setup.py) ... \u001b[?25ldone\n",
|
||||
"\u001b[?25hRequirement already satisfied: rdflib>=6.0.2 in /home/rg/anaconda3/envs/ai4beg/lib/python3.11/site-packages (from owlrl==6.0.2) (6.3.2)\n",
|
||||
"Requirement already satisfied: isodate<0.7.0,>=0.6.0 in /home/rg/anaconda3/envs/ai4beg/lib/python3.11/site-packages (from rdflib>=6.0.2->owlrl==6.0.2) (0.6.1)\n",
|
||||
"Requirement already satisfied: pyparsing<4,>=2.1.0 in /home/rg/anaconda3/envs/ai4beg/lib/python3.11/site-packages (from rdflib>=6.0.2->owlrl==6.0.2) (3.0.9)\n",
|
||||
"Requirement already satisfied: six in /home/rg/anaconda3/envs/ai4beg/lib/python3.11/site-packages (from isodate<0.7.0,>=0.6.0->rdflib>=6.0.2->owlrl==6.0.2) (1.16.0)\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"!{sys.executable} -m pip install rdflib\n",
|
||||
"!{sys.executable} -m pip install git+https://github.com/RDFLib/OWL-RL.git"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"चलो ओन्टोलॉजी फ़ाइल खोलते हैं और देखते हैं कि इसमें कितने ट्रिपलेट्स हैं:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 11,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Triplets found:669\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"import rdflib\n",
|
||||
"from owlrl import DeductiveClosure, OWLRL_Extension\n",
|
||||
"\n",
|
||||
"g = rdflib.Graph()\n",
|
||||
"g.parse(\"onto.ttl\", format=\"turtle\")\n",
|
||||
"\n",
|
||||
"print(\"Triplets found:%d\" % len(g))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": []
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 12,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Triplets after inference:4246\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"DeductiveClosure(OWLRL_Extension).expand(g)\n",
|
||||
"print(\"Triplets after inference:%d\" % len(g))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"### रिश्तेदारों के लिए क्वेरी करना\n",
|
||||
"\n",
|
||||
"अब हम ग्राफ़ से क्वेरी कर सकते हैं ताकि लोगों के बीच विभिन्न संबंध देख सकें। हम **SPARQL** भाषा का उपयोग `query` मेथड के साथ कर सकते हैं। हमारे मामले में, चलिए हमारे परिवार वृक्ष में सभी **चाचा** देखते हैं:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 13,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Fedor Alekseevich Romanov is uncle of Ekaterina Ivanovna Romanova\n",
|
||||
"Aleksandr I Pavlovich Romanov is uncle of Aleksandr II Nikolaevich Romanov\n",
|
||||
"Fedor Alekseevich Romanov is uncle of Anna Ivanovna Romanova\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"qres = g.query(\n",
|
||||
" \"\"\"SELECT DISTINCT ?aname ?bname\n",
|
||||
" WHERE {\n",
|
||||
" ?a fhkb:isUncleOf ?b .\n",
|
||||
" ?a rdfs:label ?aname .\n",
|
||||
" ?b rdfs:label ?bname .\n",
|
||||
" }\"\"\")\n",
|
||||
"\n",
|
||||
"for row in qres:\n",
|
||||
" print(\"%s is uncle of %s\" % row)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"अलग-अलग पारिवारिक संबंधों के साथ प्रयोग करने में संकोच न करें। उदाहरण के लिए, आप `isAncestorOf` संबंध देख सकते हैं, जो किसी दिए गए व्यक्ति के सभी पूर्वजों को पुनरावृत्त रूप से परिभाषित करता है।\n",
|
||||
"\n",
|
||||
"अंत में, चलिए सफाई करते हैं!\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 14,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"!rm onto.ttl"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": []
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"\n---\n\n**अस्वीकरण**: \nयह दस्तावेज़ AI अनुवाद सेवा [Co-op Translator](https://github.com/Azure/co-op-translator) का उपयोग करके अनुवादित किया गया है। जबकि हम सटीकता सुनिश्चित करने का प्रयास करते हैं, कृपया ध्यान दें कि स्वचालित अनुवाद में त्रुटियां या अशुद्धियां हो सकती हैं। मूल भाषा में उपलब्ध मूल दस्तावेज़ को प्रामाणिक स्रोत माना जाना चाहिए। महत्वपूर्ण जानकारी के लिए, पेशेवर मानव अनुवाद की सिफारिश की जाती है। इस अनुवाद के उपयोग से उत्पन्न किसी भी गलतफहमी या गलत व्याख्या के लिए हम उत्तरदायी नहीं हैं।\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
"interpreter": {
|
||||
"hash": "86193a1ab0ba47eac1c69c1756090baa3b420b3eea7d4aafab8b85f8b312f0c5"
|
||||
},
|
||||
"kernelspec": {
|
||||
"display_name": "Python 3.6",
|
||||
"language": "python",
|
||||
"name": "python3"
|
||||
},
|
||||
"language_info": {
|
||||
"codemirror_mode": {
|
||||
"name": "ipython",
|
||||
"version": 3
|
||||
},
|
||||
"file_extension": ".py",
|
||||
"mimetype": "text/x-python",
|
||||
"name": "python",
|
||||
"nbconvert_exporter": "python",
|
||||
"pygments_lexer": "ipython3",
|
||||
"version": "3.11.2"
|
||||
},
|
||||
"coopTranslator": {
|
||||
"original_hash": "6537d5597320e27b6052b4377b8ff8bb",
|
||||
"translation_date": "2025-08-31T14:53:49+00:00",
|
||||
"source_file": "lessons/2-Symbolic/FamilyOntology.ipynb",
|
||||
"language_code": "hi"
|
||||
}
|
||||
},
|
||||
"nbformat": 4,
|
||||
"nbformat_minor": 2
|
||||
}
|
||||
|
|
@ -0,0 +1,548 @@
|
|||
{
|
||||
"cells": [
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {
|
||||
"collapsed": true
|
||||
},
|
||||
"source": [
|
||||
"## माइक्रोसॉफ्ट कॉन्सेप्ट ग्राफ\n",
|
||||
"\n",
|
||||
"[Microsoft Concept Graph](https://concept.research.microsoft.com/) इंटरनेट से निकाले गए शब्दों का एक बड़ा वर्गीकरण है, जिसमें अवधारणाओं के बीच `is-a` संबंध होते हैं।\n",
|
||||
"\n",
|
||||
"कॉन्सेप्ट ग्राफ दो रूपों में उपलब्ध है:\n",
|
||||
" * डाउनलोड के लिए बड़ा टेक्स्ट फाइल\n",
|
||||
" * REST API\n",
|
||||
"\n",
|
||||
"आंकड़े:\n",
|
||||
" * 5401933 अद्वितीय अवधारणाएं\n",
|
||||
" * 12551613 अद्वितीय उदाहरण\n",
|
||||
" * 87603947 `is-a` संबंध\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## वेब सेवा का उपयोग करना\n",
|
||||
"\n",
|
||||
"वेब सेवा विभिन्न समूहों में किसी अवधारणा के संबंधित होने की संभावना का अनुमान लगाने के लिए अलग-अलग कॉल प्रदान करती है। अधिक जानकारी [यहां](https://concept.research.microsoft.com/Home/Api) उपलब्ध है। \n",
|
||||
"यहां कॉल करने के लिए एक नमूना URL है: `https://concept.research.microsoft.com/api/Concept/ScoreByProb?instance=microsoft&topK=10`\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 7,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"{'company': 0.6105356614382954,\n",
|
||||
" 'vendor': 0.08858636677518003,\n",
|
||||
" 'client': 0.048239124001183784,\n",
|
||||
" 'firm': 0.045476965571668145,\n",
|
||||
" 'large company': 0.043109401203511886,\n",
|
||||
" 'organization': 0.043010752688172046,\n",
|
||||
" 'corporation': 0.035908059583703265,\n",
|
||||
" 'brand': 0.03383644076156654,\n",
|
||||
" 'software company': 0.027522935779816515,\n",
|
||||
" 'technology company': 0.023774292196902438}"
|
||||
]
|
||||
},
|
||||
"execution_count": 7,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"import urllib\n",
|
||||
"import json\n",
|
||||
"import ssl\n",
|
||||
"\n",
|
||||
"def http(x):\n",
|
||||
" ssl._create_default_https_context = ssl._create_unverified_context\n",
|
||||
" response = urllib.request.urlopen(x)\n",
|
||||
" data = response.read()\n",
|
||||
" return data.decode('utf-8')\n",
|
||||
"\n",
|
||||
"def query(x):\n",
|
||||
" return json.loads(http(\"https://concept.research.microsoft.com/api/Concept/ScoreByProb?instance={}&topK=10\".format(urllib.parse.quote(x))))\n",
|
||||
"\n",
|
||||
"query('microsoft')"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"चलो समाचार शीर्षकों को मुख्य अवधारणाओं के अनुसार वर्गीकृत करने की कोशिश करते हैं। समाचार शीर्षक प्राप्त करने के लिए, हम [NewsApi.org](http://newsapi.org) सेवा का उपयोग करेंगे। इस सेवा का उपयोग करने के लिए आपको अपना API कुंजी प्राप्त करनी होगी - वेबसाइट पर जाएं और मुफ्त डेवलपर योजना के लिए पंजीकरण करें।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 20,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"newsapi_key = '<your API key here>'\n",
|
||||
"def get_news(country='us'):\n",
|
||||
" res = json.loads(http(\"https://newsapi.org/v2/top-headlines?country={0}&apiKey={1}\".format(country,newsapi_key)))\n",
|
||||
" return res['articles']\n",
|
||||
"\n",
|
||||
"all_titles = [x['title'] for x in get_news('us')+get_news('gb')]"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 21,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"['Covid-19 Live Updates: Vaccines and Boosters News - The New York Times',\n",
|
||||
" 'Ukrainians Flee Mariupol as Russian Forces Push to Take Port City - The Wall Street Journal',\n",
|
||||
" 'Bond Yields Jump, Stock Futures Rise After Powell Says Fed Is Ready to Be More Aggressive - The Wall Street Journal',\n",
|
||||
" 'Putin critic Alexei Navalny found guilty by Russian court - New York Post ',\n",
|
||||
" \"Supreme Court nominee Ketanji Brown Jackson will face questions at confirmation hearing's second day - CNN\",\n",
|
||||
" '2 teachers killed at Swedish high school, student arrested - ABC News',\n",
|
||||
" 'Clues to Covid-19’s Next Moves Come From Sewers - The Wall Street Journal',\n",
|
||||
" 'Republicans to roll dice by grilling Jackson over child-pornography sentencing decisions | TheHill - The Hill',\n",
|
||||
" '‘Clear sign’ Putin considering using chemical weapons in Ukraine, claims President Biden - The Independent',\n",
|
||||
" 'NASA confirms there are 5,000 planets outside our solar system - Daily Mail',\n",
|
||||
" \"US stocks whipsawed overnight after Fed Chair Powell's remarks - Fox Business\",\n",
|
||||
" \"'We've learned absolutely nothing': Tests could again be in short supply if Covid surges - POLITICO\",\n",
|
||||
" \"Duchess of Cambridge swaps khaki jungle gear for Vampire's Wife dress on Belize trip - Daily Mail\",\n",
|
||||
" 'China searches for victims, flight recorders after first plane crash in 12 years - Reuters',\n",
|
||||
" 'Second superyacht linked to Russian oligarch Abramovich docks in Turkey - Reuters',\n",
|
||||
" 'Live updates: Russia stops talks with Japan over sanctions - The Associated Press - en Español',\n",
|
||||
" 'Powers Remain and Threats Lurk as Women’s Sweet 16 Is Set - The New York Times',\n",
|
||||
" 'Webb Space Telescope Begins Multi-Instrument Alignment - SciTechDaily',\n",
|
||||
" \"UConn vs UCF - NCAA women's tournament second-round highlights - March Madness\",\n",
|
||||
" 'Bucking Republican Trend, Indiana Governor Vetoes Transgender Sports Bill - The New York Times',\n",
|
||||
" \"Maggie Fox dead: Coronation Street and Shameless actress dies after 'sudden accident' - Mirror Online - The Mirror\",\n",
|
||||
" 'China plane crash – live: Search for survivors continues as witness describes moment flight fell from sky - The Independent',\n",
|
||||
" 'Daniel Morgan murder: damning report condemns Met police - The Guardian',\n",
|
||||
" 'What to expect from Rishi Sunak’s Spring Statement - BBC.com',\n",
|
||||
" 'UK and Republic of Ireland in line to host Euro 2028 after no one else bids - The Guardian',\n",
|
||||
" \"Friends beg Vladimir Putin's 'lover' to persuade him to end Ukraine invasion - The Mirror\",\n",
|
||||
" 'Brass Eye’s outtakes show the brutal TV comedy was the tip of an iceberg - The Guardian',\n",
|
||||
" \"Vladimir Putin threatens civilians to break Mariupol's spirit - The Times\",\n",
|
||||
" 'Shell U-turn on Cambo oilfield would threaten green targets, say campaigners - The Guardian',\n",
|
||||
" 'St Helens dog attack: Girl aged 17 months killed at home - BBC',\n",
|
||||
" \"PlayStation to buy 'Assassin's Creed' veteran Jade Raymond's Haven Studios - NME\",\n",
|
||||
" '‘Clear sign’ Putin considering using chemical weapons in Ukraine, claims President Biden - The Independent',\n",
|
||||
" 'NASA confirms there are 5,000 planets outside our solar system - Daily Mail',\n",
|
||||
" 'Nintendo Switch finally has folders • Eurogamer.net - Eurogamer.net',\n",
|
||||
" 'FA to “find a solution” as Liverpool fan group blasts “shambolic” Wembley travel - This Is Anfield',\n",
|
||||
" 'Manchester United transfer news LIVE Erik ten Hag latest and Man Utd manager updates - Manchester Evening News',\n",
|
||||
" 'Inflation raises cost of UK government borrowing in February; crude oil up again – business live - The Guardian',\n",
|
||||
" 'Alexei Navalny: Kremlin critic found guilty of large-scale fraud and contempt of court by Russian court - Sky News',\n",
|
||||
" \"UK prepares to nationalize Russia natural gas giant Gazprom's retail unit - Business Insider\",\n",
|
||||
" 'Zaghari-Ratcliffe: Hunt calls for inquiry into delay over Iran debt payment - The Guardian']"
|
||||
]
|
||||
},
|
||||
"execution_count": 21,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"all_titles"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"सबसे पहले, हम समाचार शीर्षकों से संज्ञा निकालने में सक्षम होना चाहते हैं। हम इसे करने के लिए `TextBlob` लाइब्रेरी का उपयोग करेंगे, जो इस तरह के सामान्य NLP कार्यों को बहुत सरल बनाती है।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 15,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Requirement already satisfied: textblob in c:\\winapp\\miniconda3\\lib\\site-packages (0.17.1)\n",
|
||||
"Requirement already satisfied: nltk>=3.1 in c:\\winapp\\miniconda3\\lib\\site-packages (from textblob) (3.5)\n",
|
||||
"Requirement already satisfied: joblib in c:\\winapp\\miniconda3\\lib\\site-packages (from nltk>=3.1->textblob) (1.0.1)\n",
|
||||
"Requirement already satisfied: regex in c:\\winapp\\miniconda3\\lib\\site-packages (from nltk>=3.1->textblob) (2021.11.10)\n",
|
||||
"Requirement already satisfied: tqdm in c:\\winapp\\miniconda3\\lib\\site-packages (from nltk>=3.1->textblob) (4.61.2)\n",
|
||||
"Requirement already satisfied: click in c:\\winapp\\miniconda3\\lib\\site-packages (from nltk>=3.1->textblob) (8.0.3)\n",
|
||||
"Requirement already satisfied: colorama in c:\\winapp\\miniconda3\\lib\\site-packages (from click->nltk>=3.1->textblob) (0.4.4)\n",
|
||||
"Finished.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"name": "stderr",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"[nltk_data] Downloading package brown to\n",
|
||||
"[nltk_data] C:\\Users\\dmitryso\\AppData\\Roaming\\nltk_data...\n",
|
||||
"[nltk_data] Package brown is already up-to-date!\n",
|
||||
"[nltk_data] Downloading package punkt to\n",
|
||||
"[nltk_data] C:\\Users\\dmitryso\\AppData\\Roaming\\nltk_data...\n",
|
||||
"[nltk_data] Package punkt is already up-to-date!\n",
|
||||
"[nltk_data] Downloading package wordnet to\n",
|
||||
"[nltk_data] C:\\Users\\dmitryso\\AppData\\Roaming\\nltk_data...\n",
|
||||
"[nltk_data] Package wordnet is already up-to-date!\n",
|
||||
"[nltk_data] Downloading package averaged_perceptron_tagger to\n",
|
||||
"[nltk_data] C:\\Users\\dmitryso\\AppData\\Roaming\\nltk_data...\n",
|
||||
"[nltk_data] Package averaged_perceptron_tagger is already up-to-\n",
|
||||
"[nltk_data] date!\n",
|
||||
"[nltk_data] Downloading package conll2000 to\n",
|
||||
"[nltk_data] C:\\Users\\dmitryso\\AppData\\Roaming\\nltk_data...\n",
|
||||
"[nltk_data] Package conll2000 is already up-to-date!\n",
|
||||
"[nltk_data] Downloading package movie_reviews to\n",
|
||||
"[nltk_data] C:\\Users\\dmitryso\\AppData\\Roaming\\nltk_data...\n",
|
||||
"[nltk_data] Package movie_reviews is already up-to-date!\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"import sys\n",
|
||||
"!{sys.executable} -m pip install textblob\n",
|
||||
"!{sys.executable} -m textblob.download_corpora\n",
|
||||
"from textblob import TextBlob"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 22,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"{'covid-19 live updates': 1,\n",
|
||||
" 'vaccines': 1,\n",
|
||||
" 'boosters': 1,\n",
|
||||
" 'york': 4,\n",
|
||||
" 'ukrainians flee mariupol': 1,\n",
|
||||
" 'forces push': 1,\n",
|
||||
" 'port city': 1,\n",
|
||||
" 'wall street journal': 3,\n",
|
||||
" 'bond yields': 1,\n",
|
||||
" 'futures rise': 1,\n",
|
||||
" 'powell says fed': 1,\n",
|
||||
" 'ready': 1,\n",
|
||||
" 'be': 1,\n",
|
||||
" 'aggressive': 1,\n",
|
||||
" 'putin': 3,\n",
|
||||
" 'alexei navalny': 2,\n",
|
||||
" 'russian': 2,\n",
|
||||
" 'supreme court nominee': 1,\n",
|
||||
" 'ketanji brown jackson': 1,\n",
|
||||
" \"confirmation hearing 's\": 1,\n",
|
||||
" 'cnn': 1,\n",
|
||||
" 'swedish': 1,\n",
|
||||
" 'high school': 1,\n",
|
||||
" 'abc': 1,\n",
|
||||
" 'clues': 1,\n",
|
||||
" 'covid-19': 1,\n",
|
||||
" '’ s': 2,\n",
|
||||
" 'moves': 1,\n",
|
||||
" 'sewers': 1,\n",
|
||||
" 'roll dice': 1,\n",
|
||||
" 'jackson': 1,\n",
|
||||
" 'decisions |': 1,\n",
|
||||
" 'thehill': 1,\n",
|
||||
" 'clear': 2,\n",
|
||||
" 'chemical weapons': 2,\n",
|
||||
" 'ukraine': 3,\n",
|
||||
" 'claims president': 2,\n",
|
||||
" 'biden': 2,\n",
|
||||
" 'nasa': 2,\n",
|
||||
" 'solar system': 2,\n",
|
||||
" 'daily mail': 3,\n",
|
||||
" 'us stocks': 1,\n",
|
||||
" 'fed chair powell': 1,\n",
|
||||
" \"'s remarks\": 1,\n",
|
||||
" 'fox': 1,\n",
|
||||
" \"'we 've\": 1,\n",
|
||||
" 'tests': 1,\n",
|
||||
" 'covid': 1,\n",
|
||||
" 'politico': 1,\n",
|
||||
" 'duchess': 1,\n",
|
||||
" 'cambridge': 1,\n",
|
||||
" 'swaps khaki jungle gear': 1,\n",
|
||||
" 'vampire': 1,\n",
|
||||
" 'wife': 1,\n",
|
||||
" 'belize': 1,\n",
|
||||
" 'china': 2,\n",
|
||||
" 'flight recorders': 1,\n",
|
||||
" 'plane crash': 1,\n",
|
||||
" 'reuters': 2,\n",
|
||||
" 'russian oligarch': 1,\n",
|
||||
" 'abramovich': 1,\n",
|
||||
" 'live': 1,\n",
|
||||
" 'russia': 2,\n",
|
||||
" 'stops talks': 1,\n",
|
||||
" 'japan': 1,\n",
|
||||
" 'español': 1,\n",
|
||||
" 'powers remain': 1,\n",
|
||||
" 'threats lurk': 1,\n",
|
||||
" 'set': 1,\n",
|
||||
" 'webb': 1,\n",
|
||||
" 'telescope begins multi-instrument alignment': 1,\n",
|
||||
" 'scitechdaily': 1,\n",
|
||||
" 'uconn': 1,\n",
|
||||
" 'ucf': 1,\n",
|
||||
" 'ncaa': 1,\n",
|
||||
" \"women 's tournament second-round highlights\": 1,\n",
|
||||
" 'march madness': 1,\n",
|
||||
" 'bucking republican trend': 1,\n",
|
||||
" 'indiana': 1,\n",
|
||||
" 'vetoes transgender': 1,\n",
|
||||
" 'bill': 1,\n",
|
||||
" 'maggie fox': 1,\n",
|
||||
" 'coronation': 1,\n",
|
||||
" 'shameless': 1,\n",
|
||||
" \"'sudden accident\": 1,\n",
|
||||
" 'mirror online': 1,\n",
|
||||
" 'mirror': 2,\n",
|
||||
" 'plane crash –': 1,\n",
|
||||
" 'search': 1,\n",
|
||||
" 'moment flight': 1,\n",
|
||||
" 'daniel morgan': 1,\n",
|
||||
" 'report condemns': 1,\n",
|
||||
" 'met': 1,\n",
|
||||
" 'guardian': 6,\n",
|
||||
" 'rishi sunak': 1,\n",
|
||||
" '’ s spring': 1,\n",
|
||||
" 'statement': 1,\n",
|
||||
" 'bbc.com': 1,\n",
|
||||
" 'uk': 3,\n",
|
||||
" 'ireland': 1,\n",
|
||||
" 'euro': 1,\n",
|
||||
" 'vladimir putin': 2,\n",
|
||||
" \"'s 'lover\": 1,\n",
|
||||
" 'brass eye': 1,\n",
|
||||
" '’ s outtakes': 1,\n",
|
||||
" 'brutal tv comedy': 1,\n",
|
||||
" 'threatens civilians': 1,\n",
|
||||
" 'mariupol': 1,\n",
|
||||
" \"'s spirit\": 1,\n",
|
||||
" 'shell u-turn': 1,\n",
|
||||
" 'cambo': 1,\n",
|
||||
" 'green targets': 1,\n",
|
||||
" 'st helens': 1,\n",
|
||||
" 'dog attack': 1,\n",
|
||||
" 'girl': 1,\n",
|
||||
" 'bbc': 1,\n",
|
||||
" 'playstation': 1,\n",
|
||||
" \"'assassin 's\": 1,\n",
|
||||
" 'creed': 1,\n",
|
||||
" 'jade raymond': 1,\n",
|
||||
" 'haven studios': 1,\n",
|
||||
" 'nme': 1,\n",
|
||||
" 'nintendo switch': 1,\n",
|
||||
" 'folders •': 1,\n",
|
||||
" 'eurogamer.net': 2,\n",
|
||||
" 'fa': 1,\n",
|
||||
" 'solution ”': 1,\n",
|
||||
" 'liverpool': 1,\n",
|
||||
" 'fan group blasts “ shambolic ”': 1,\n",
|
||||
" 'wembley': 1,\n",
|
||||
" 'anfield': 1,\n",
|
||||
" 'manchester': 1,\n",
|
||||
" 'live erik': 1,\n",
|
||||
" 'hag': 1,\n",
|
||||
" 'utd': 1,\n",
|
||||
" 'manager updates': 1,\n",
|
||||
" 'manchester evening': 1,\n",
|
||||
" 'inflation': 1,\n",
|
||||
" 'government borrowing': 1,\n",
|
||||
" 'february': 1,\n",
|
||||
" 'crude oil': 1,\n",
|
||||
" '– business': 1,\n",
|
||||
" 'kremlin': 1,\n",
|
||||
" 'large-scale fraud': 1,\n",
|
||||
" 'sky': 1,\n",
|
||||
" 'natural gas': 1,\n",
|
||||
" 'gazprom': 1,\n",
|
||||
" 'retail unit': 1,\n",
|
||||
" 'insider': 1,\n",
|
||||
" 'zaghari-ratcliffe': 1,\n",
|
||||
" 'hunt': 1,\n",
|
||||
" 'iran': 1,\n",
|
||||
" 'debt payment': 1}"
|
||||
]
|
||||
},
|
||||
"execution_count": 22,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"w = {}\n",
|
||||
"for x in all_titles:\n",
|
||||
" for n in TextBlob(x).noun_phrases:\n",
|
||||
" if n in w:\n",
|
||||
" w[n].append(x)\n",
|
||||
" else:\n",
|
||||
" w[n]=[x]\n",
|
||||
"{ x:len(w[x]) for x in w.keys()}"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"हम देख सकते हैं कि संज्ञाएं हमें बड़े विषयगत समूह नहीं देती हैं। आइए संज्ञाओं को अवधारणा ग्राफ से प्राप्त अधिक सामान्य शब्दों से बदलें। इसमें कुछ समय लगेगा, क्योंकि हम प्रत्येक संज्ञा वाक्यांश के लिए REST कॉल कर रहे हैं।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 23,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"w = {}\n",
|
||||
"for x in all_titles:\n",
|
||||
" for noun in TextBlob(x).noun_phrases:\n",
|
||||
" terms = query(noun.replace(' ','%20'))\n",
|
||||
" for term in [u for u in terms.keys() if terms[u]>0.1]:\n",
|
||||
" if term in w:\n",
|
||||
" w[term].append(x)\n",
|
||||
" else:\n",
|
||||
" w[term]=[x]"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 24,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"{'city': 9,\n",
|
||||
" 'brand': 4,\n",
|
||||
" 'place': 9,\n",
|
||||
" 'town': 4,\n",
|
||||
" 'factor': 4,\n",
|
||||
" 'film': 4,\n",
|
||||
" 'nation': 11,\n",
|
||||
" 'state': 5,\n",
|
||||
" 'person': 4,\n",
|
||||
" 'organization': 5,\n",
|
||||
" 'publication': 10,\n",
|
||||
" 'market': 5,\n",
|
||||
" 'economy': 4,\n",
|
||||
" 'company': 6,\n",
|
||||
" 'newspaper': 6,\n",
|
||||
" 'relationship': 6}"
|
||||
]
|
||||
},
|
||||
"execution_count": 24,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"{ x:len(w[x]) for x in w.keys() if len(w[x])>3}"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 27,
|
||||
"metadata": {
|
||||
"trusted": true
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"\n",
|
||||
"ECONOMY:\n",
|
||||
"China searches for victims, flight recorders after first plane crash in 12 years - Reuters\n",
|
||||
"Live updates: Russia stops talks with Japan over sanctions - The Associated Press - en Español\n",
|
||||
"China plane crash – live: Search for survivors continues as witness describes moment flight fell from sky - The Independent\n",
|
||||
"UK prepares to nationalize Russia natural gas giant Gazprom's retail unit - Business Insider\n",
|
||||
"\n",
|
||||
"NATION:\n",
|
||||
"‘Clear sign’ Putin considering using chemical weapons in Ukraine, claims President Biden - The Independent\n",
|
||||
"Duchess of Cambridge swaps khaki jungle gear for Vampire's Wife dress on Belize trip - Daily Mail\n",
|
||||
"China searches for victims, flight recorders after first plane crash in 12 years - Reuters\n",
|
||||
"Live updates: Russia stops talks with Japan over sanctions - The Associated Press - en Español\n",
|
||||
"Live updates: Russia stops talks with Japan over sanctions - The Associated Press - en Español\n",
|
||||
"China plane crash – live: Search for survivors continues as witness describes moment flight fell from sky - The Independent\n",
|
||||
"UK and Republic of Ireland in line to host Euro 2028 after no one else bids - The Guardian\n",
|
||||
"Friends beg Vladimir Putin's 'lover' to persuade him to end Ukraine invasion - The Mirror\n",
|
||||
"‘Clear sign’ Putin considering using chemical weapons in Ukraine, claims President Biden - The Independent\n",
|
||||
"UK prepares to nationalize Russia natural gas giant Gazprom's retail unit - Business Insider\n",
|
||||
"Zaghari-Ratcliffe: Hunt calls for inquiry into delay over Iran debt payment - The Guardian\n",
|
||||
"\n",
|
||||
"PERSON:\n",
|
||||
"‘Clear sign’ Putin considering using chemical weapons in Ukraine, claims President Biden - The Independent\n",
|
||||
"Duchess of Cambridge swaps khaki jungle gear for Vampire's Wife dress on Belize trip - Daily Mail\n",
|
||||
"Second superyacht linked to Russian oligarch Abramovich docks in Turkey - Reuters\n",
|
||||
"‘Clear sign’ Putin considering using chemical weapons in Ukraine, claims President Biden - The Independent\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"print('\\nECONOMY:\\n'+'\\n'.join(w['economy']))\n",
|
||||
"print('\\nNATION:\\n'+'\\n'.join(w['nation']))\n",
|
||||
"print('\\nPERSON:\\n'+'\\n'.join(w['person']))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"\n---\n\n**अस्वीकरण**: \nयह दस्तावेज़ AI अनुवाद सेवा [Co-op Translator](https://github.com/Azure/co-op-translator) का उपयोग करके अनुवादित किया गया है। जबकि हम सटीकता सुनिश्चित करने का प्रयास करते हैं, कृपया ध्यान दें कि स्वचालित अनुवाद में त्रुटियां या अशुद्धियां हो सकती हैं। मूल भाषा में उपलब्ध मूल दस्तावेज़ को प्रामाणिक स्रोत माना जाना चाहिए। महत्वपूर्ण जानकारी के लिए, पेशेवर मानव अनुवाद की सिफारिश की जाती है। इस अनुवाद के उपयोग से उत्पन्न किसी भी गलतफहमी या गलत व्याख्या के लिए हम उत्तरदायी नहीं हैं।\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
"kernelspec": {
|
||||
"display_name": "Python 3.7.4 64-bit (conda)",
|
||||
"metadata": {
|
||||
"interpreter": {
|
||||
"hash": "86193a1ab0ba47eac1c69c1756090baa3b420b3eea7d4aafab8b85f8b312f0c5"
|
||||
}
|
||||
},
|
||||
"name": "python3"
|
||||
},
|
||||
"language_info": {
|
||||
"codemirror_mode": {
|
||||
"name": "ipython",
|
||||
"version": 3
|
||||
},
|
||||
"file_extension": ".py",
|
||||
"mimetype": "text/x-python",
|
||||
"name": "python",
|
||||
"nbconvert_exporter": "python",
|
||||
"pygments_lexer": "ipython3",
|
||||
"version": "3.9.5"
|
||||
},
|
||||
"coopTranslator": {
|
||||
"original_hash": "4087f998407d06ceb2947016ba4605d0",
|
||||
"translation_date": "2025-08-31T14:54:16+00:00",
|
||||
"source_file": "lessons/2-Symbolic/MSConceptGraph.ipynb",
|
||||
"language_code": "hi"
|
||||
}
|
||||
},
|
||||
"nbformat": 4,
|
||||
"nbformat_minor": 2
|
||||
}
|
||||
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
|
|
@ -1,31 +1,33 @@
|
|||
<!--
|
||||
CO_OP_TRANSLATOR_METADATA:
|
||||
{
|
||||
"original_hash": "7336583e4630220c835335da640016db",
|
||||
"translation_date": "2025-08-24T10:01:06+00:00",
|
||||
"original_hash": "ba5d1eb353d20d3e7181066b3c424b99",
|
||||
"translation_date": "2025-08-31T14:21:50+00:00",
|
||||
"source_file": "lessons/3-NeuralNetworks/03-Perceptron/lab/README.md",
|
||||
"language_code": "hi"
|
||||
}
|
||||
-->
|
||||
# मल्टी-क्लास वर्गीकरण परसेप्ट्रॉन के साथ
|
||||
# परसेप्ट्रॉन के साथ मल्टी-क्लास वर्गीकरण
|
||||
|
||||
[AI for Beginners Curriculum](https://github.com/microsoft/ai-for-beginners) से लैब असाइनमेंट।
|
||||
|
||||
## कार्य
|
||||
|
||||
इस पाठ में हमने MNIST हस्तलिखित अंकों के द्विआधारी वर्गीकरण के लिए जो कोड विकसित किया है, उसका उपयोग करके एक मल्टी-क्लास वर्गीकृत बनाएं जो किसी भी अंक को पहचान सके। ट्रेन और टेस्ट डेटासेट पर वर्गीकरण सटीकता की गणना करें, और भ्रम मैट्रिक्स को प्रिंट करें।
|
||||
इस पाठ में विकसित कोड का उपयोग करते हुए, जो MNIST हस्तलिखित अंकों के बाइनरी वर्गीकरण के लिए है, एक मल्टी-क्लास वर्गीकर्ता बनाएं जो किसी भी अंक को पहचान सके। ट्रेन और टेस्ट डेटा सेट पर वर्गीकरण की सटीकता की गणना करें, और कन्फ्यूजन मैट्रिक्स को प्रिंट करें।
|
||||
|
||||
## संकेत
|
||||
## सुझाव
|
||||
|
||||
1. प्रत्येक अंक के लिए, "इस अंक बनाम अन्य सभी अंकों" के द्विआधारी वर्गीकरण के लिए एक डेटासेट बनाएं।
|
||||
1. द्विआधारी वर्गीकरण के लिए 10 अलग-अलग परसेप्ट्रॉन प्रशिक्षित करें (प्रत्येक अंक के लिए एक)।
|
||||
1. प्रत्येक अंक के लिए, "यह अंक बनाम अन्य सभी अंक" के बाइनरी वर्गीकरण के लिए एक डेटा सेट बनाएं।
|
||||
1. बाइनरी वर्गीकरण के लिए 10 अलग-अलग परसेप्ट्रॉन प्रशिक्षित करें (प्रत्येक अंक के लिए एक)।
|
||||
1. एक फ़ंक्शन परिभाषित करें जो इनपुट अंक को वर्गीकृत करेगा।
|
||||
|
||||
> **संकेत**: यदि हम सभी 10 परसेप्ट्रॉन के वज़न को एक मैट्रिक्स में जोड़ते हैं, तो हम इनपुट अंकों पर एक मैट्रिक्स गुणा द्वारा सभी 10 परसेप्ट्रॉन लागू कर सकते हैं। सबसे संभावित अंक को `argmax` ऑपरेशन लागू करके आउटपुट से पाया जा सकता है।
|
||||
> **सुझाव**: यदि हम सभी 10 परसेप्ट्रॉनों के वज़न को एक मैट्रिक्स में संयोजित करें, तो हम एक ही मैट्रिक्स गुणा द्वारा इनपुट अंकों पर सभी 10 परसेप्ट्रॉनों को लागू कर सकते हैं। सबसे संभावित अंक को `argmax` ऑपरेशन को आउटपुट पर लागू करके आसानी से पाया जा सकता है।
|
||||
|
||||
## प्रारंभिक नोटबुक
|
||||
|
||||
लैब शुरू करने के लिए [PerceptronMultiClass.ipynb](../../../../../../lessons/3-NeuralNetworks/03-Perceptron/lab/PerceptronMultiClass.ipynb) खोलें।
|
||||
लैब शुरू करने के लिए [PerceptronMultiClass.ipynb](PerceptronMultiClass.ipynb) खोलें।
|
||||
|
||||
---
|
||||
|
||||
**अस्वीकरण**:
|
||||
यह दस्तावेज़ AI अनुवाद सेवा [Co-op Translator](https://github.com/Azure/co-op-translator) का उपयोग करके अनुवादित किया गया है। जबकि हम सटीकता के लिए प्रयासरत हैं, कृपया ध्यान दें कि स्वचालित अनुवाद में त्रुटियां या अशुद्धियां हो सकती हैं। मूल भाषा में उपलब्ध मूल दस्तावेज़ को प्रामाणिक स्रोत माना जाना चाहिए। महत्वपूर्ण जानकारी के लिए, पेशेवर मानव अनुवाद की सिफारिश की जाती है। इस अनुवाद के उपयोग से उत्पन्न किसी भी गलतफहमी या गलत व्याख्या के लिए हम उत्तरदायी नहीं हैं।
|
||||
File diff suppressed because one or more lines are too long
|
|
@ -0,0 +1,183 @@
|
|||
{
|
||||
"cells": [
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"# हमारे फ्रेमवर्क के साथ MNIST अंकों का वर्गीकरण\n",
|
||||
"\n",
|
||||
"[AI for Beginners Curriculum](https://github.com/microsoft/ai-for-beginners) से लैब असाइनमेंट।\n",
|
||||
"\n",
|
||||
"### डेटासेट पढ़ना\n",
|
||||
"\n",
|
||||
"यह कोड इंटरनेट पर रिपॉजिटरी से डेटासेट डाउनलोड करता है। आप AI Curriculum रिपो की `/data` डायरेक्टरी से डेटासेट को मैन्युअली भी कॉपी कर सकते हैं।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 4,
|
||||
"metadata": {
|
||||
"tags": []
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stderr",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
" % Total % Received % Xferd Average Speed Time Time Time Current\n",
|
||||
" Dload Upload Total Spent Left Speed\n",
|
||||
"\n",
|
||||
" 0 0 0 0 0 0 0 0 --:--:-- --:--:-- --:--:-- 0\n",
|
||||
"100 9.9M 100 9.9M 0 0 9.9M 0 0:00:01 --:--:-- 0:00:01 15.8M\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"!rm *.pkl\n",
|
||||
"!wget https://raw.githubusercontent.com/microsoft/AI-For-Beginners/main/data/mnist.pkl.gz\n",
|
||||
"!gzip -d mnist.pkl.gz"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 3,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"import pickle\n",
|
||||
"with open('mnist.pkl','rb') as f:\n",
|
||||
" MNIST = pickle.load(f)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"labels = MNIST['Train']['Labels']\n",
|
||||
"data = MNIST['Train']['Features']"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"आइए देखें कि हमारे पास डेटा का आकार क्या है:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 5,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"(42000, 784)"
|
||||
]
|
||||
},
|
||||
"execution_count": 5,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"data.shape"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"### डेटा को विभाजित करना\n",
|
||||
"\n",
|
||||
"हम प्रशिक्षण और परीक्षण डेटा सेट के बीच डेटा को विभाजित करने के लिए Scikit Learn का उपयोग करेंगे:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 6,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Train samples: 33600, test samples: 8400\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"from sklearn.model_selection import train_test_split\n",
|
||||
"\n",
|
||||
"features_train, features_test, labels_train, labels_test = train_test_split(data,labels,test_size=0.2)\n",
|
||||
"\n",
|
||||
"print(f\"Train samples: {len(features_train)}, test samples: {len(features_test)}\")"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"### निर्देश\n",
|
||||
"\n",
|
||||
"1. पाठ से फ्रेमवर्क कोड लें और इसे इस नोटबुक में पेस्ट करें, या (और बेहतर) इसे एक अलग Python मॉड्यूल में रखें।\n",
|
||||
"1. एक-स्तरीय परसेप्ट्रॉन को परिभाषित करें और प्रशिक्षित करें, प्रशिक्षण और सत्यापन सटीकता को प्रशिक्षण के दौरान देखें।\n",
|
||||
"1. समझने की कोशिश करें कि क्या ओवरफिटिंग हुई है, और सटीकता सुधारने के लिए लेयर पैरामीटर समायोजित करें।\n",
|
||||
"1. पिछले चरणों को 2- और 3-स्तरीय परसेप्ट्रॉन के लिए दोहराएं। लेयर्स के बीच विभिन्न सक्रियण फ़ंक्शनों के साथ प्रयोग करने की कोशिश करें।\n",
|
||||
"1. निम्नलिखित प्रश्नों का उत्तर देने की कोशिश करें:\n",
|
||||
" - क्या इंटर-लेयर सक्रियण फ़ंक्शन नेटवर्क प्रदर्शन को प्रभावित करता है?\n",
|
||||
" - क्या इस कार्य के लिए हमें 2- या 3-स्तरीय नेटवर्क की आवश्यकता है?\n",
|
||||
" - क्या आपने नेटवर्क को प्रशिक्षित करने में कोई समस्या अनुभव की? विशेष रूप से जब लेयर्स की संख्या बढ़ गई।\n",
|
||||
" - प्रशिक्षण के दौरान नेटवर्क के वज़न कैसे व्यवहार करते हैं? आप वज़न के अधिकतम abs मान को epochs के मुकाबले प्लॉट कर सकते हैं ताकि संबंध को समझा जा सके।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": []
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"\n---\n\n**अस्वीकरण**: \nयह दस्तावेज़ AI अनुवाद सेवा [Co-op Translator](https://github.com/Azure/co-op-translator) का उपयोग करके अनुवादित किया गया है। जबकि हम सटीकता सुनिश्चित करने का प्रयास करते हैं, कृपया ध्यान दें कि स्वचालित अनुवाद में त्रुटियां या अशुद्धियां हो सकती हैं। मूल भाषा में उपलब्ध मूल दस्तावेज़ को प्रामाणिक स्रोत माना जाना चाहिए। महत्वपूर्ण जानकारी के लिए, पेशेवर मानव अनुवाद की सिफारिश की जाती है। इस अनुवाद के उपयोग से उत्पन्न किसी भी गलतफहमी या गलत व्याख्या के लिए हम उत्तरदायी नहीं हैं।\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
"kernelspec": {
|
||||
"display_name": "Python 3.7.4 64-bit (conda)",
|
||||
"metadata": {
|
||||
"interpreter": {
|
||||
"hash": "86193a1ab0ba47eac1c69c1756090baa3b420b3eea7d4aafab8b85f8b312f0c5"
|
||||
}
|
||||
},
|
||||
"name": "python3"
|
||||
},
|
||||
"language_info": {
|
||||
"codemirror_mode": {
|
||||
"name": "ipython",
|
||||
"version": 3
|
||||
},
|
||||
"file_extension": ".py",
|
||||
"mimetype": "text/x-python",
|
||||
"name": "python",
|
||||
"nbconvert_exporter": "python",
|
||||
"pygments_lexer": "ipython3",
|
||||
"version": "3.9.5"
|
||||
},
|
||||
"orig_nbformat": 2,
|
||||
"coopTranslator": {
|
||||
"original_hash": "6fa055f484eb5d6bdf41166a356d3abf",
|
||||
"translation_date": "2025-08-31T14:58:44+00:00",
|
||||
"source_file": "lessons/3-NeuralNetworks/04-OwnFramework/lab/MyFW_MNIST.ipynb",
|
||||
"language_code": "hi"
|
||||
}
|
||||
},
|
||||
"nbformat": 4,
|
||||
"nbformat_minor": 2
|
||||
}
|
||||
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
|
|
@ -0,0 +1,108 @@
|
|||
{
|
||||
"cells": [
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## ऑप्टिकल फ्लो का उपयोग करके हथेली की गति का पता लगाना\n",
|
||||
"\n",
|
||||
"यह लैब [AI for Beginners Curriculum](http://aka.ms/ai-beginners) का हिस्सा है।\n",
|
||||
"\n",
|
||||
"[इस वीडियो](../../../../../../lessons/4-ComputerVision/06-IntroCV/lab/palm-movement.mp4) पर विचार करें, जिसमें एक व्यक्ति की हथेली स्थिर पृष्ठभूमि पर बाईं/दाईं/ऊपर/नीचे की ओर हिलती है।\n",
|
||||
"\n",
|
||||
"**आपका लक्ष्य** ऑप्टिकल फ्लो का उपयोग करके यह निर्धारित करना होगा कि वीडियो के कौन से हिस्से में ऊपर/नीचे/बाईं/दाईं ओर की गति है।\n",
|
||||
"\n",
|
||||
"लेक्चर में बताए गए अनुसार वीडियो फ्रेम प्राप्त करके शुरू करें:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"# Code here"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"अब, व्याख्यान में वर्णित अनुसार सघन ऑप्टिकल प्रवाह फ्रेम्स की गणना करें, और सघन ऑप्टिकल प्रवाह को ध्रुवीय निर्देशांकों में परिवर्तित करें:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"# Code here"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"प्रत्येक ऑप्टिकल फ्लो फ्रेम के लिए दिशाओं का हिस्टोग्राम बनाएं। एक हिस्टोग्राम दिखाता है कि कितने वेक्टर एक निश्चित बिन के अंतर्गत आते हैं, और यह फ्रेम पर विभिन्न दिशाओं की गति को अलग करना चाहिए।\n",
|
||||
"\n",
|
||||
"> आप उन सभी वेक्टर को भी शून्य कर सकते हैं जिनका परिमाण एक निश्चित सीमा से नीचे है। यह वीडियो में छोटी अतिरिक्त गतियों, जैसे आंखें और सिर, को हटा देगा।\n",
|
||||
"\n",
|
||||
"कुछ फ्रेम्स के लिए हिस्टोग्राम प्लॉट करें।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"# Code here"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"हिस्टोग्राम को देखते हुए, यह समझना काफी आसान होना चाहिए कि गति की दिशा कैसे निर्धारित करें। आपको उन बिन्स को चुनना होगा जो ऊपर/नीचे/बाएँ/दाएँ दिशाओं से संबंधित हैं, और जो एक निश्चित सीमा से ऊपर हैं।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"# Code here"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"बधाई हो! यदि आपने ऊपर दिए गए सभी चरण पूरे कर लिए हैं, तो आपने प्रयोगशाला पूरी कर ली है!\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"\n---\n\n**अस्वीकरण**: \nयह दस्तावेज़ AI अनुवाद सेवा [Co-op Translator](https://github.com/Azure/co-op-translator) का उपयोग करके अनुवादित किया गया है। जबकि हम सटीकता सुनिश्चित करने का प्रयास करते हैं, कृपया ध्यान दें कि स्वचालित अनुवाद में त्रुटियां या अशुद्धियां हो सकती हैं। मूल भाषा में उपलब्ध मूल दस्तावेज़ को प्रामाणिक स्रोत माना जाना चाहिए। महत्वपूर्ण जानकारी के लिए, पेशेवर मानव अनुवाद की सिफारिश की जाती है। इस अनुवाद के उपयोग से उत्पन्न किसी भी गलतफहमी या गलत व्याख्या के लिए हम उत्तरदायी नहीं हैं।\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
"language_info": {
|
||||
"name": "python"
|
||||
},
|
||||
"orig_nbformat": 4,
|
||||
"coopTranslator": {
|
||||
"original_hash": "153d9e417e079bf62f8f693002d0deaf",
|
||||
"translation_date": "2025-08-31T14:42:52+00:00",
|
||||
"source_file": "lessons/4-ComputerVision/06-IntroCV/lab/MovementDetection.ipynb",
|
||||
"language_code": "hi"
|
||||
}
|
||||
},
|
||||
"nbformat": 4,
|
||||
"nbformat_minor": 2
|
||||
}
|
||||
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
|
|
@ -0,0 +1,577 @@
|
|||
{
|
||||
"cells": [
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"# टेक्स्ट वर्गीकरण कार्य\n",
|
||||
"\n",
|
||||
"जैसा कि हमने उल्लेख किया है, हम **AG_NEWS** डेटासेट पर आधारित एक सरल टेक्स्ट वर्गीकरण कार्य पर ध्यान केंद्रित करेंगे, जिसमें समाचार सुर्खियों को 4 श्रेणियों में वर्गीकृत करना है: विश्व, खेल, व्यवसाय और विज्ञान/तकनीक।\n",
|
||||
"\n",
|
||||
"## डेटासेट\n",
|
||||
"\n",
|
||||
"यह डेटासेट [`torchtext`](https://github.com/pytorch/text) मॉड्यूल में शामिल है, इसलिए हम इसे आसानी से एक्सेस कर सकते हैं।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 1,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"import torch\n",
|
||||
"import torchtext\n",
|
||||
"import os\n",
|
||||
"import collections\n",
|
||||
"os.makedirs('./data',exist_ok=True)\n",
|
||||
"train_dataset, test_dataset = torchtext.datasets.AG_NEWS(root='./data')\n",
|
||||
"classes = ['World', 'Sports', 'Business', 'Sci/Tech']"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"यहाँ, `train_dataset` और `test_dataset` में संग्रह होते हैं जो क्रमशः वर्ग (कक्षा की संख्या) और पाठ के जोड़े लौटाते हैं, उदाहरण के लिए:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 2,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"(3,\n",
|
||||
" \"Wall St. Bears Claw Back Into the Black (Reuters) Reuters - Short-sellers, Wall Street's dwindling\\\\band of ultra-cynics, are seeing green again.\")"
|
||||
]
|
||||
},
|
||||
"execution_count": 2,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"list(train_dataset)[0]"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"तो, चलिए हमारे डेटासेट से पहले 10 नई सुर्खियाँ प्रिंट करते हैं:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 5,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"**Sci/Tech** -> Wall St. Bears Claw Back Into the Black (Reuters) Reuters - Short-sellers, Wall Street's dwindling\\band of ultra-cynics, are seeing green again.\n",
|
||||
"**Sci/Tech** -> Carlyle Looks Toward Commercial Aerospace (Reuters) Reuters - Private investment firm Carlyle Group,\\which has a reputation for making well-timed and occasionally\\controversial plays in the defense industry, has quietly placed\\its bets on another part of the market.\n",
|
||||
"**Sci/Tech** -> Oil and Economy Cloud Stocks' Outlook (Reuters) Reuters - Soaring crude prices plus worries\\about the economy and the outlook for earnings are expected to\\hang over the stock market next week during the depth of the\\summer doldrums.\n",
|
||||
"**Sci/Tech** -> Iraq Halts Oil Exports from Main Southern Pipeline (Reuters) Reuters - Authorities have halted oil export\\flows from the main pipeline in southern Iraq after\\intelligence showed a rebel militia could strike\\infrastructure, an oil official said on Saturday.\n",
|
||||
"**Sci/Tech** -> Oil prices soar to all-time record, posing new menace to US economy (AFP) AFP - Tearaway world oil prices, toppling records and straining wallets, present a new economic menace barely three months before the US presidential elections.\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"for i,x in zip(range(5),train_dataset):\n",
|
||||
" print(f\"**{classes[x[0]]}** -> {x[1]}\")\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"क्योंकि डेटासेट इटरेटर होते हैं, यदि हम डेटा का उपयोग कई बार करना चाहते हैं तो हमें इसे सूची में बदलना होगा:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 3,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"train_dataset, test_dataset = torchtext.datasets.AG_NEWS(root='./data')\n",
|
||||
"train_dataset = list(train_dataset)\n",
|
||||
"test_dataset = list(test_dataset)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## टोकनाइज़ेशन\n",
|
||||
"\n",
|
||||
"अब हमें टेक्स्ट को **संख्याओं** में बदलने की आवश्यकता है, जिन्हें टेन्सर के रूप में प्रस्तुत किया जा सके। यदि हम शब्द-स्तरीय प्रतिनिधित्व चाहते हैं, तो हमें दो चीजें करनी होंगी:\n",
|
||||
"* **टोकनाइज़र** का उपयोग करके टेक्स्ट को **टोकन** में विभाजित करना\n",
|
||||
"* उन टोकन का एक **शब्दकोश** बनाना।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"['he', 'said', 'hello']"
|
||||
]
|
||||
},
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"tokenizer = torchtext.data.utils.get_tokenizer('basic_english')\n",
|
||||
"tokenizer('He said: hello')"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 5,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"counter = collections.Counter()\n",
|
||||
"for (label, line) in train_dataset:\n",
|
||||
" counter.update(tokenizer(line))\n",
|
||||
"vocab = torchtext.vocab.vocab(counter, min_freq=1)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"शब्दावली का उपयोग करके, हम आसानी से अपने टोकनयुक्त स्ट्रिंग को संख्याओं के सेट में एन्कोड कर सकते हैं:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 19,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Vocab size if 95810\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"[599, 3279, 97, 1220, 329, 225, 7368]"
|
||||
]
|
||||
},
|
||||
"execution_count": 19,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"vocab_size = len(vocab)\n",
|
||||
"print(f\"Vocab size if {vocab_size}\")\n",
|
||||
"\n",
|
||||
"stoi = vocab.get_stoi() # dict to convert tokens to indices\n",
|
||||
"\n",
|
||||
"def encode(x):\n",
|
||||
" return [stoi[s] for s in tokenizer(x)]\n",
|
||||
"\n",
|
||||
"encode('I love to play with my words')"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## शब्दों का थैला (Bag of Words) टेक्स्ट प्रतिनिधित्व\n",
|
||||
"\n",
|
||||
"क्योंकि शब्द अर्थ को दर्शाते हैं, कभी-कभी हम केवल व्यक्तिगत शब्दों को देखकर, उनके वाक्य में क्रम की परवाह किए बिना, टेक्स्ट का अर्थ समझ सकते हैं। उदाहरण के लिए, जब समाचार वर्गीकृत कर रहे हों, तो *मौसम*, *बर्फ* जैसे शब्द *मौसम पूर्वानुमान* का संकेत दे सकते हैं, जबकि *शेयर*, *डॉलर* जैसे शब्द *वित्तीय समाचार* की ओर इशारा करेंगे।\n",
|
||||
"\n",
|
||||
"**शब्दों का थैला** (BoW) वेक्टर प्रतिनिधित्व सबसे सामान्य रूप से उपयोग किया जाने वाला पारंपरिक वेक्टर प्रतिनिधित्व है। प्रत्येक शब्द को एक वेक्टर इंडेक्स से जोड़ा जाता है, और वेक्टर तत्व में किसी दिए गए दस्तावेज़ में शब्द की घटनाओं की संख्या होती है।\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"> **Note**: आप BoW को टेक्स्ट में व्यक्तिगत शब्दों के लिए सभी वन-हॉट-एन्कोडेड वेक्टर का योग भी मान सकते हैं।\n",
|
||||
"\n",
|
||||
"नीचे Scikit Learn पायथन लाइब्रेरी का उपयोग करके शब्दों के थैले का प्रतिनिधित्व बनाने का एक उदाहरण दिया गया है:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 7,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"array([[1, 1, 0, 2, 0, 0, 0, 0, 0]], dtype=int64)"
|
||||
]
|
||||
},
|
||||
"execution_count": 7,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"from sklearn.feature_extraction.text import CountVectorizer\n",
|
||||
"vectorizer = CountVectorizer()\n",
|
||||
"corpus = [\n",
|
||||
" 'I like hot dogs.',\n",
|
||||
" 'The dog ran fast.',\n",
|
||||
" 'Its hot outside.',\n",
|
||||
" ]\n",
|
||||
"vectorizer.fit_transform(corpus)\n",
|
||||
"vectorizer.transform(['My dog likes hot dogs on a hot day.']).toarray()"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"AG_NEWS डेटासेट के वेक्टर प्रतिनिधित्व से बैग-ऑफ-वर्ड्स वेक्टर की गणना करने के लिए, हम निम्नलिखित फ़ंक्शन का उपयोग कर सकते हैं:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 20,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"tensor([2., 1., 2., ..., 0., 0., 0.])\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"vocab_size = len(vocab)\n",
|
||||
"\n",
|
||||
"def to_bow(text,bow_vocab_size=vocab_size):\n",
|
||||
" res = torch.zeros(bow_vocab_size,dtype=torch.float32)\n",
|
||||
" for i in encode(text):\n",
|
||||
" if i<bow_vocab_size:\n",
|
||||
" res[i] += 1\n",
|
||||
" return res\n",
|
||||
"\n",
|
||||
"print(to_bow(train_dataset[0][1]))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"> **नोट:** यहां हम वैश्विक `vocab_size` चर का उपयोग कर रहे हैं ताकि शब्दावली के डिफ़ॉल्ट आकार को निर्दिष्ट किया जा सके। चूंकि अक्सर शब्दावली का आकार काफी बड़ा होता है, हम सबसे अधिक बार उपयोग किए जाने वाले शब्दों तक शब्दावली के आकार को सीमित कर सकते हैं। `vocab_size` मान को कम करने और नीचे दिए गए कोड को चलाने का प्रयास करें, और देखें कि यह सटीकता को कैसे प्रभावित करता है। आपको कुछ सटीकता में गिरावट की उम्मीद करनी चाहिए, लेकिन प्रदर्शन में सुधार के बदले यह नाटकीय नहीं होगा।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## BoW क्लासिफायर का प्रशिक्षण\n",
|
||||
"\n",
|
||||
"अब जब हमने अपने टेक्स्ट का बैग-ऑफ-वर्ड्स प्रतिनिधित्व बनाना सीख लिया है, तो आइए इसके ऊपर एक क्लासिफायर को प्रशिक्षित करें। सबसे पहले, हमें अपने डेटासेट को इस तरह से प्रशिक्षण के लिए बदलने की आवश्यकता है कि सभी पोजिशनल वेक्टर प्रतिनिधित्व को बैग-ऑफ-वर्ड्स प्रतिनिधित्व में परिवर्तित किया जा सके। इसे `bowify` फ़ंक्शन को मानक टॉर्च `DataLoader` में `collate_fn` पैरामीटर के रूप में पास करके प्राप्त किया जा सकता है:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 21,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"from torch.utils.data import DataLoader\n",
|
||||
"import numpy as np \n",
|
||||
"\n",
|
||||
"# this collate function gets list of batch_size tuples, and needs to \n",
|
||||
"# return a pair of label-feature tensors for the whole minibatch\n",
|
||||
"def bowify(b):\n",
|
||||
" return (\n",
|
||||
" torch.LongTensor([t[0]-1 for t in b]),\n",
|
||||
" torch.stack([to_bow(t[1]) for t in b])\n",
|
||||
" )\n",
|
||||
"\n",
|
||||
"train_loader = DataLoader(train_dataset, batch_size=16, collate_fn=bowify, shuffle=True)\n",
|
||||
"test_loader = DataLoader(test_dataset, batch_size=16, collate_fn=bowify, shuffle=True)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"अब चलिए एक साधारण वर्गीकरण न्यूरल नेटवर्क को परिभाषित करते हैं जिसमें एक रैखिक परत होती है। इनपुट वेक्टर का आकार `vocab_size` के बराबर है, और आउटपुट आकार वर्गों की संख्या (4) के अनुरूप है। क्योंकि हम वर्गीकरण कार्य को हल कर रहे हैं, अंतिम सक्रियता फ़ंक्शन `LogSoftmax()` है।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 22,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"net = torch.nn.Sequential(torch.nn.Linear(vocab_size,4),torch.nn.LogSoftmax(dim=1))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"अब हम मानक PyTorch प्रशिक्षण लूप को परिभाषित करेंगे। क्योंकि हमारा डेटासेट काफी बड़ा है, शिक्षण उद्देश्य के लिए हम केवल एक युग के लिए प्रशिक्षण करेंगे, और कभी-कभी एक युग से भी कम (प्रशिक्षण को सीमित करने के लिए `epoch_size` पैरामीटर निर्दिष्ट किया जाता है)। हम प्रशिक्षण के दौरान संचित प्रशिक्षण सटीकता की भी रिपोर्ट करेंगे; रिपोर्टिंग की आवृत्ति `report_freq` पैरामीटर का उपयोग करके निर्दिष्ट की जाती है।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 24,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"def train_epoch(net,dataloader,lr=0.01,optimizer=None,loss_fn = torch.nn.NLLLoss(),epoch_size=None, report_freq=200):\n",
|
||||
" optimizer = optimizer or torch.optim.Adam(net.parameters(),lr=lr)\n",
|
||||
" net.train()\n",
|
||||
" total_loss,acc,count,i = 0,0,0,0\n",
|
||||
" for labels,features in dataloader:\n",
|
||||
" optimizer.zero_grad()\n",
|
||||
" out = net(features)\n",
|
||||
" loss = loss_fn(out,labels) #cross_entropy(out,labels)\n",
|
||||
" loss.backward()\n",
|
||||
" optimizer.step()\n",
|
||||
" total_loss+=loss\n",
|
||||
" _,predicted = torch.max(out,1)\n",
|
||||
" acc+=(predicted==labels).sum()\n",
|
||||
" count+=len(labels)\n",
|
||||
" i+=1\n",
|
||||
" if i%report_freq==0:\n",
|
||||
" print(f\"{count}: acc={acc.item()/count}\")\n",
|
||||
" if epoch_size and count>epoch_size:\n",
|
||||
" break\n",
|
||||
" return total_loss.item()/count, acc.item()/count"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 25,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"3200: acc=0.8028125\n",
|
||||
"6400: acc=0.8371875\n",
|
||||
"9600: acc=0.8534375\n",
|
||||
"12800: acc=0.85765625\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"(0.026090790722161722, 0.8620069296375267)"
|
||||
]
|
||||
},
|
||||
"execution_count": 25,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"train_epoch(net,train_loader,epoch_size=15000)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## बाईग्राम्स, ट्राईग्राम्स और एन-ग्राम्स\n",
|
||||
"\n",
|
||||
"बैग ऑफ वर्ड्स दृष्टिकोण की एक सीमा यह है कि कुछ शब्द बहु-शब्द अभिव्यक्तियों का हिस्सा होते हैं। उदाहरण के लिए, 'हॉट डॉग' शब्द का अर्थ पूरी तरह से अलग होता है, जबकि 'हॉट' और 'डॉग' शब्द अन्य संदर्भों में अलग-अलग अर्थ रखते हैं। यदि हम हमेशा 'हॉट' और 'डॉग' शब्दों को एक ही वेक्टर द्वारा प्रदर्शित करें, तो यह हमारे मॉडल को भ्रमित कर सकता है।\n",
|
||||
"\n",
|
||||
"इस समस्या को हल करने के लिए, **एन-ग्राम प्रतिनिधित्व** का उपयोग अक्सर दस्तावेज़ वर्गीकरण की विधियों में किया जाता है, जहां प्रत्येक शब्द, द्वि-शब्द या त्रि-शब्द की आवृत्ति वर्गीकरणकर्ताओं को प्रशिक्षित करने के लिए एक उपयोगी विशेषता होती है। उदाहरण के लिए, बाईग्राम प्रतिनिधित्व में, हम मूल शब्दों के अलावा सभी शब्द युग्मों को शब्दावली में जोड़ेंगे।\n",
|
||||
"\n",
|
||||
"नीचे एक उदाहरण दिया गया है कि कैसे Scikit Learn का उपयोग करके बाईग्राम बैग ऑफ वर्ड्स प्रतिनिधित्व बनाया जा सकता है:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 26,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Vocabulary:\n",
|
||||
" {'i': 7, 'like': 11, 'hot': 4, 'dogs': 2, 'i like': 8, 'like hot': 12, 'hot dogs': 5, 'the': 16, 'dog': 0, 'ran': 14, 'fast': 3, 'the dog': 17, 'dog ran': 1, 'ran fast': 15, 'its': 9, 'outside': 13, 'its hot': 10, 'hot outside': 6}\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"array([[1, 0, 1, 0, 2, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]],\n",
|
||||
" dtype=int64)"
|
||||
]
|
||||
},
|
||||
"execution_count": 26,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"bigram_vectorizer = CountVectorizer(ngram_range=(1, 2), token_pattern=r'\\b\\w+\\b', min_df=1)\n",
|
||||
"corpus = [\n",
|
||||
" 'I like hot dogs.',\n",
|
||||
" 'The dog ran fast.',\n",
|
||||
" 'Its hot outside.',\n",
|
||||
" ]\n",
|
||||
"bigram_vectorizer.fit_transform(corpus)\n",
|
||||
"print(\"Vocabulary:\\n\",bigram_vectorizer.vocabulary_)\n",
|
||||
"bigram_vectorizer.transform(['My dog likes hot dogs on a hot day.']).toarray()\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"N-gram दृष्टिकोण की मुख्य कमी यह है कि शब्दावली का आकार बहुत तेजी से बढ़ने लगता है। व्यवहार में, हमें N-gram प्रतिनिधित्व को कुछ आयाम-घटाने की तकनीकों, जैसे *embeddings*, के साथ संयोजित करने की आवश्यकता होती है, जिन पर हम अगले यूनिट में चर्चा करेंगे।\n",
|
||||
"\n",
|
||||
"हमारे **AG News** डेटासेट में N-gram प्रतिनिधित्व का उपयोग करने के लिए, हमें एक विशेष ngram शब्दावली बनानी होगी:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 27,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Bigram vocabulary length = 1308842\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"counter = collections.Counter()\n",
|
||||
"for (label, line) in train_dataset:\n",
|
||||
" l = tokenizer(line)\n",
|
||||
" counter.update(torchtext.data.utils.ngrams_iterator(l,ngrams=2))\n",
|
||||
" \n",
|
||||
"bi_vocab = torchtext.vocab.vocab(counter, min_freq=1)\n",
|
||||
"\n",
|
||||
"print(\"Bigram vocabulary length = \",len(bi_vocab))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"हम ऊपर दिए गए कोड का उपयोग करके क्लासिफायर को ट्रेन कर सकते हैं, लेकिन यह मेमोरी के लिहाज से बहुत अक्षम होगा। अगले यूनिट में, हम एम्बेडिंग्स का उपयोग करके बिग्राम क्लासिफायर को ट्रेन करेंगे।\n",
|
||||
"\n",
|
||||
"> **नोट:** आप केवल उन्हीं ngrams को छोड़ सकते हैं जो टेक्स्ट में निर्दिष्ट संख्या से अधिक बार आते हैं। यह सुनिश्चित करेगा कि कम बार आने वाले बिग्राम्स को हटा दिया जाए, और डाइमेंशनलिटी को काफी हद तक कम किया जा सके। ऐसा करने के लिए, `min_freq` पैरामीटर को उच्च मान पर सेट करें, और शब्दावली की लंबाई में बदलाव को देखें।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## टर्म फ्रीक्वेंसी इनवर्स डॉक्यूमेंट फ्रीक्वेंसी (TF-IDF)\n",
|
||||
"\n",
|
||||
"BoW (बैग ऑफ वर्ड्स) प्रतिनिधित्व में, शब्दों की उपस्थिति को समान रूप से महत्व दिया जाता है, चाहे वह शब्द कोई भी हो। हालांकि, यह स्पष्ट है कि सामान्य शब्द जैसे *a*, *in* आदि वर्गीकरण के लिए उतने महत्वपूर्ण नहीं होते जितने कि विशेष शब्द। वास्तव में, अधिकांश NLP कार्यों में कुछ शब्द अन्य शब्दों की तुलना में अधिक प्रासंगिक होते हैं।\n",
|
||||
"\n",
|
||||
"**TF-IDF** का मतलब है **टर्म फ्रीक्वेंसी–इनवर्स डॉक्यूमेंट फ्रीक्वेंसी**। यह बैग ऑफ वर्ड्स का एक प्रकार है, जिसमें किसी दस्तावेज़ में शब्द की उपस्थिति को दर्शाने वाले 0/1 बाइनरी मान के बजाय एक फ्लोटिंग-पॉइंट मान का उपयोग किया जाता है, जो कॉर्पस में शब्द की उपस्थिति की आवृत्ति से संबंधित होता है।\n",
|
||||
"\n",
|
||||
"औपचारिक रूप से, किसी शब्द $i$ का वजन $w_{ij}$ किसी दस्तावेज़ $j$ में इस प्रकार परिभाषित किया जाता है:\n",
|
||||
"$$\n",
|
||||
"w_{ij} = tf_{ij}\\times\\log({N\\over df_i})\n",
|
||||
"$$\n",
|
||||
"जहां:\n",
|
||||
"* $tf_{ij}$ किसी दस्तावेज़ $j$ में शब्द $i$ की उपस्थिति की संख्या है, यानी वह BoW मान जिसे हमने पहले देखा था\n",
|
||||
"* $N$ संग्रह में दस्तावेज़ों की कुल संख्या है\n",
|
||||
"* $df_i$ पूरे संग्रह में शब्द $i$ को शामिल करने वाले दस्तावेज़ों की संख्या है\n",
|
||||
"\n",
|
||||
"TF-IDF मान $w_{ij}$ किसी दस्तावेज़ में शब्द की उपस्थिति की संख्या के अनुपात में बढ़ता है और कॉर्पस में उस शब्द को शामिल करने वाले दस्तावेज़ों की संख्या से समायोजित होता है। यह इस तथ्य को संतुलित करने में मदद करता है कि कुछ शब्द अन्य शब्दों की तुलना में अधिक बार दिखाई देते हैं। उदाहरण के लिए, यदि कोई शब्द *हर* दस्तावेज़ में दिखाई देता है, तो $df_i=N$, और $w_{ij}=0$, और ऐसे शब्दों को पूरी तरह से नजरअंदाज कर दिया जाएगा।\n",
|
||||
"\n",
|
||||
"आप आसानी से Scikit Learn का उपयोग करके टेक्स्ट का TF-IDF वेक्टराइज़ेशन बना सकते हैं:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 28,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"array([[0.43381609, 0. , 0.43381609, 0. , 0.65985664,\n",
|
||||
" 0.43381609, 0. , 0. , 0. , 0. ,\n",
|
||||
" 0. , 0. , 0. , 0. , 0. ,\n",
|
||||
" 0. ]])"
|
||||
]
|
||||
},
|
||||
"execution_count": 28,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"from sklearn.feature_extraction.text import TfidfVectorizer\n",
|
||||
"vectorizer = TfidfVectorizer(ngram_range=(1,2))\n",
|
||||
"vectorizer.fit_transform(corpus)\n",
|
||||
"vectorizer.transform(['My dog likes hot dogs on a hot day.']).toarray()"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## निष्कर्ष\n",
|
||||
"\n",
|
||||
"हालांकि TF-IDF प्रतिनिधित्व विभिन्न शब्दों को आवृत्ति भार प्रदान करते हैं, वे अर्थ या क्रम को व्यक्त करने में असमर्थ होते हैं। जैसा कि प्रसिद्ध भाषाविद् जे. आर. फर्थ ने 1935 में कहा था, \"शब्द का पूर्ण अर्थ हमेशा संदर्भात्मक होता है, और संदर्भ से अलग अर्थ का कोई भी अध्ययन गंभीरता से नहीं लिया जा सकता।\" हम इस पाठ्यक्रम में आगे सीखेंगे कि भाषा मॉडलिंग का उपयोग करके पाठ से संदर्भात्मक जानकारी कैसे प्राप्त करें।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"\n---\n\n**अस्वीकरण**: \nयह दस्तावेज़ AI अनुवाद सेवा [Co-op Translator](https://github.com/Azure/co-op-translator) का उपयोग करके अनुवादित किया गया है। जबकि हम सटीकता सुनिश्चित करने का प्रयास करते हैं, कृपया ध्यान दें कि स्वचालित अनुवाद में त्रुटियां या अशुद्धियां हो सकती हैं। मूल भाषा में उपलब्ध मूल दस्तावेज़ को प्रामाणिक स्रोत माना जाना चाहिए। महत्वपूर्ण जानकारी के लिए, पेशेवर मानव अनुवाद की सिफारिश की जाती है। इस अनुवाद के उपयोग से उत्पन्न किसी भी गलतफहमी या गलत व्याख्या के लिए हम उत्तरदायी नहीं हैं।\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
"interpreter": {
|
||||
"hash": "16af2a8bbb083ea23e5e41c7f5787656b2ce26968575d8763f2c4b17f9cd711f"
|
||||
},
|
||||
"kernelspec": {
|
||||
"display_name": "Python 3.8.12 ('py38')",
|
||||
"language": "python",
|
||||
"name": "python3"
|
||||
},
|
||||
"language_info": {
|
||||
"codemirror_mode": {
|
||||
"name": "ipython",
|
||||
"version": 3
|
||||
},
|
||||
"file_extension": ".py",
|
||||
"mimetype": "text/x-python",
|
||||
"name": "python",
|
||||
"nbconvert_exporter": "python",
|
||||
"pygments_lexer": "ipython3",
|
||||
"version": "3.8.12"
|
||||
},
|
||||
"coopTranslator": {
|
||||
"original_hash": "7b9040985e748e4e2d4c689892456ad7",
|
||||
"translation_date": "2025-08-31T15:29:56+00:00",
|
||||
"source_file": "lessons/5-NLP/13-TextRep/TextRepresentationPyTorch.ipynb",
|
||||
"language_code": "hi"
|
||||
}
|
||||
},
|
||||
"nbformat": 4,
|
||||
"nbformat_minor": 2
|
||||
}
|
||||
|
|
@ -0,0 +1,647 @@
|
|||
{
|
||||
"cells": [
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"# टेक्स्ट वर्गीकरण कार्य\n",
|
||||
"\n",
|
||||
"इस मॉड्यूल में, हम **[AG_NEWS](http://www.di.unipi.it/~gulli/AG_corpus_of_news_articles.html)** डेटासेट पर आधारित एक सरल टेक्स्ट वर्गीकरण कार्य से शुरुआत करेंगे: हम समाचार शीर्षकों को चार श्रेणियों में वर्गीकृत करेंगे: वर्ल्ड, स्पोर्ट्स, बिज़नेस और साइ/टेक।\n",
|
||||
"\n",
|
||||
"## डेटासेट\n",
|
||||
"\n",
|
||||
"डेटासेट को लोड करने के लिए, हम **[TensorFlow Datasets](https://www.tensorflow.org/datasets)** API का उपयोग करेंगे।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 1,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"import tensorflow as tf\n",
|
||||
"from tensorflow import keras\n",
|
||||
"import tensorflow_datasets as tfds\n",
|
||||
"\n",
|
||||
"# In this tutorial, we will be training a lot of models. In order to use GPU memory cautiously,\n",
|
||||
"# we will set tensorflow option to grow GPU memory allocation when required.\n",
|
||||
"physical_devices = tf.config.list_physical_devices('GPU') \n",
|
||||
"if len(physical_devices)>0:\n",
|
||||
" tf.config.experimental.set_memory_growth(physical_devices[0], True)\n",
|
||||
"\n",
|
||||
"dataset = tfds.load('ag_news_subset')"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"हम अब `dataset['train']` और `dataset['test']` का उपयोग करके डेटासेट के प्रशिक्षण और परीक्षण भागों तक पहुंच सकते हैं:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 3,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Length of train dataset = 120000\n",
|
||||
"Length of test dataset = 7600\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"ds_train = dataset['train']\n",
|
||||
"ds_test = dataset['test']\n",
|
||||
"\n",
|
||||
"print(f\"Length of train dataset = {len(ds_train)}\")\n",
|
||||
"print(f\"Length of test dataset = {len(ds_test)}\")"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"चलो हमारे डेटा सेट से पहले 10 नई सुर्खियाँ प्रिंट करें:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"3 (Sci/Tech) -> b'AMD Debuts Dual-Core Opteron Processor' b'AMD #39;s new dual-core Opteron chip is designed mainly for corporate computing applications, including databases, Web services, and financial transactions.'\n",
|
||||
"1 (Sports) -> b\"Wood's Suspension Upheld (Reuters)\" b'Reuters - Major League Baseball\\\\Monday announced a decision on the appeal filed by Chicago Cubs\\\\pitcher Kerry Wood regarding a suspension stemming from an\\\\incident earlier this season.'\n",
|
||||
"2 (Business) -> b'Bush reform may have blue states seeing red' b'President Bush #39;s quot;revenue-neutral quot; tax reform needs losers to balance its winners, and people claiming the federal deduction for state and local taxes may be in administration planners #39; sights, news reports say.'\n",
|
||||
"3 (Sci/Tech) -> b\"'Halt science decline in schools'\" b'Britain will run out of leading scientists unless science education is improved, says Professor Colin Pillinger.'\n",
|
||||
"1 (Sports) -> b'Gerrard leaves practice' b'London, England (Sports Network) - England midfielder Steven Gerrard injured his groin late in Thursday #39;s training session, but is hopeful he will be ready for Saturday #39;s World Cup qualifier against Austria.'\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"classes = ['World', 'Sports', 'Business', 'Sci/Tech']\n",
|
||||
"\n",
|
||||
"for i,x in zip(range(5),ds_train):\n",
|
||||
" print(f\"{x['label']} ({classes[x['label']]}) -> {x['title']} {x['description']}\")"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## टेक्स्ट वेक्टराइजेशन\n",
|
||||
"\n",
|
||||
"अब हमें टेक्स्ट को **संख्याओं** में बदलना होगा, जिन्हें टेन्सर्स के रूप में प्रस्तुत किया जा सके। अगर हमें शब्द-स्तरीय प्रतिनिधित्व चाहिए, तो हमें दो चीजें करनी होंगी:\n",
|
||||
"\n",
|
||||
"* एक **टोकनाइज़र** का उपयोग करके टेक्स्ट को **टोकन्स** में विभाजित करें।\n",
|
||||
"* उन टोकन्स का एक **शब्दकोश** (वोकैबुलरी) बनाएं।\n",
|
||||
"\n",
|
||||
"### शब्दकोश का आकार सीमित करना\n",
|
||||
"\n",
|
||||
"AG News डेटासेट के उदाहरण में, शब्दकोश का आकार काफी बड़ा है, 100k से अधिक शब्द। सामान्य तौर पर, हमें उन शब्दों की आवश्यकता नहीं होती जो टेक्स्ट में बहुत कम बार आते हैं — केवल कुछ वाक्यों में ही वे मौजूद होंगे, और मॉडल उनसे कुछ सीख नहीं पाएगा। इसलिए, शब्दकोश के आकार को छोटा करने के लिए इसे सीमित करना समझदारी है। यह वेक्टराइज़र कंस्ट्रक्टर में एक आर्ग्युमेंट पास करके किया जा सकता है:\n",
|
||||
"\n",
|
||||
"इन दोनों चरणों को **TextVectorization** लेयर का उपयोग करके संभाला जा सकता है। आइए वेक्टराइज़र ऑब्जेक्ट को इंस्टैंसिएट करें, और फिर `adapt` मेथड को कॉल करें ताकि सभी टेक्स्ट को पार किया जा सके और एक शब्दकोश बनाया जा सके:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 5,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"vocab_size = 50000\n",
|
||||
"vectorizer = keras.layers.experimental.preprocessing.TextVectorization(max_tokens=vocab_size)\n",
|
||||
"vectorizer.adapt(ds_train.take(500).map(lambda x: x['title']+' '+x['description']))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"> **नोट** हम पूरे डेटासेट का केवल एक छोटा हिस्सा उपयोग कर रहे हैं ताकि शब्दावली बनाई जा सके। ऐसा हम निष्पादन समय को तेज करने और आपको इंतजार न कराने के लिए कर रहे हैं। हालांकि, हम यह जोखिम उठा रहे हैं कि पूरे डेटासेट के कुछ शब्द शब्दावली में शामिल नहीं होंगे और प्रशिक्षण के दौरान अनदेखा कर दिए जाएंगे। इसलिए, पूरे शब्दावली आकार का उपयोग करना और `adapt` के दौरान पूरे डेटासेट से गुजरना अंतिम सटीकता को बढ़ा सकता है, लेकिन बहुत अधिक नहीं।\n",
|
||||
"\n",
|
||||
"अब हम वास्तविक शब्दावली तक पहुंच सकते हैं:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 6,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"['', '[UNK]', 'the', 'to', 'a', 'in', 'of', 'and', 'on', 'for']\n",
|
||||
"Length of vocabulary: 5335\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"vocab = vectorizer.get_vocabulary()\n",
|
||||
"vocab_size = len(vocab)\n",
|
||||
"print(vocab[:10])\n",
|
||||
"print(f\"Length of vocabulary: {vocab_size}\")"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"वेक्टराइज़र का उपयोग करके, हम आसानी से किसी भी पाठ को संख्याओं के एक सेट में एन्कोड कर सकते हैं:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 7,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"<tf.Tensor: shape=(7,), dtype=int64, numpy=array([ 112, 3695, 3, 304, 11, 1041, 1], dtype=int64)>"
|
||||
]
|
||||
},
|
||||
"execution_count": 7,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"vectorizer('I love to play with my words')"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## बैग-ऑफ-वर्ड्स टेक्स्ट प्रतिनिधित्व\n",
|
||||
"\n",
|
||||
"क्योंकि शब्द अर्थ को दर्शाते हैं, कभी-कभी हम केवल व्यक्तिगत शब्दों को देखकर किसी टेक्स्ट के अर्थ का पता लगा सकते हैं, भले ही वे वाक्य में किस क्रम में हों। उदाहरण के लिए, जब समाचार वर्गीकृत कर रहे हों, तो *मौसम* और *बर्फ* जैसे शब्द *मौसम पूर्वानुमान* का संकेत दे सकते हैं, जबकि *शेयर* और *डॉलर* जैसे शब्द *वित्तीय समाचार* की ओर इशारा करेंगे।\n",
|
||||
"\n",
|
||||
"**बैग-ऑफ-वर्ड्स** (BoW) वेक्टर प्रतिनिधित्व सबसे सरल और पारंपरिक वेक्टर प्रतिनिधित्व है जिसे समझा जा सकता है। प्रत्येक शब्द को एक वेक्टर इंडेक्स से जोड़ा जाता है, और एक वेक्टर तत्व में दिए गए दस्तावेज़ में प्रत्येक शब्द की घटनाओं की संख्या होती है।\n",
|
||||
"\n",
|
||||
" \n",
|
||||
"\n",
|
||||
"> **Note**: आप BoW को टेक्स्ट में व्यक्तिगत शब्दों के लिए सभी वन-हॉट-एनकोडेड वेक्टर का योग भी मान सकते हैं।\n",
|
||||
"\n",
|
||||
"नीचे Scikit Learn पायथन लाइब्रेरी का उपयोग करके बैग-ऑफ-वर्ड्स प्रतिनिधित्व उत्पन्न करने का एक उदाहरण दिया गया है:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 8,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"array([[1, 1, 0, 2, 0, 0, 0, 0, 0]], dtype=int64)"
|
||||
]
|
||||
},
|
||||
"execution_count": 8,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"from sklearn.feature_extraction.text import CountVectorizer\n",
|
||||
"sc_vectorizer = CountVectorizer()\n",
|
||||
"corpus = [\n",
|
||||
" 'I like hot dogs.',\n",
|
||||
" 'The dog ran fast.',\n",
|
||||
" 'Its hot outside.',\n",
|
||||
" ]\n",
|
||||
"sc_vectorizer.fit_transform(corpus)\n",
|
||||
"sc_vectorizer.transform(['My dog likes hot dogs on a hot day.']).toarray()"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"हम ऊपर परिभाषित किए गए Keras वेक्टराइज़र का भी उपयोग कर सकते हैं, प्रत्येक शब्द संख्या को एक वन-हॉट एन्कोडिंग में परिवर्तित करके और उन सभी वेक्टरों को जोड़ सकते हैं:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 9,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"array([0., 5., 0., ..., 0., 0., 0.], dtype=float32)"
|
||||
]
|
||||
},
|
||||
"execution_count": 9,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"def to_bow(text):\n",
|
||||
" return tf.reduce_sum(tf.one_hot(vectorizer(text),vocab_size),axis=0)\n",
|
||||
"\n",
|
||||
"to_bow('My dog likes hot dogs on a hot day.').numpy()"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"> **नोट**: आपको यह देखकर आश्चर्य हो सकता है कि परिणाम पिछले उदाहरण से अलग है। इसका कारण यह है कि Keras उदाहरण में वेक्टर की लंबाई शब्दावली के आकार के अनुरूप होती है, जो पूरे AG News डेटासेट से बनाई गई थी, जबकि Scikit Learn उदाहरण में हमने नमूना पाठ से तुरंत शब्दावली बनाई।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## BoW क्लासिफायर को प्रशिक्षित करना\n",
|
||||
"\n",
|
||||
"अब जब हमने अपने टेक्स्ट का बैग-ऑफ-वर्ड्स प्रतिनिधित्व बनाना सीख लिया है, तो चलिए एक क्लासिफायर को प्रशिक्षित करते हैं जो इसका उपयोग करता है। सबसे पहले, हमें अपने डेटासेट को बैग-ऑफ-वर्ड्स प्रतिनिधित्व में बदलना होगा। इसे निम्नलिखित तरीके से `map` फ़ंक्शन का उपयोग करके प्राप्त किया जा सकता है:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 11,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"batch_size = 128\n",
|
||||
"\n",
|
||||
"ds_train_bow = ds_train.map(lambda x: (to_bow(x['title']+x['description']),x['label'])).batch(batch_size)\n",
|
||||
"ds_test_bow = ds_test.map(lambda x: (to_bow(x['title']+x['description']),x['label'])).batch(batch_size)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"अब चलिए एक साधारण वर्गीकरण न्यूरल नेटवर्क को परिभाषित करते हैं जिसमें एक रैखिक परत होती है। इनपुट आकार `vocab_size` है, और आउटपुट आकार वर्गों की संख्या (4) के अनुरूप है। क्योंकि हम एक वर्गीकरण कार्य को हल कर रहे हैं, अंतिम सक्रियण फ़ंक्शन **softmax** है:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 12,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"938/938 [==============================] - 66s 70ms/step - loss: 0.6144 - acc: 0.8427 - val_loss: 0.4416 - val_acc: 0.8697\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"<keras.callbacks.History at 0x20c70a947f0>"
|
||||
]
|
||||
},
|
||||
"execution_count": 12,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"model = keras.models.Sequential([\n",
|
||||
" keras.layers.Dense(4,activation='softmax',input_shape=(vocab_size,))\n",
|
||||
"])\n",
|
||||
"model.compile(loss='sparse_categorical_crossentropy',optimizer='adam',metrics=['acc'])\n",
|
||||
"model.fit(ds_train_bow,validation_data=ds_test_bow)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"चूंकि हमारे पास 4 क्लासेस हैं, 80% से अधिक की सटीकता एक अच्छा परिणाम है।\n",
|
||||
"\n",
|
||||
"## एक नेटवर्क के रूप में क्लासिफायर को ट्रेन करना\n",
|
||||
"\n",
|
||||
"क्योंकि वेक्टराइज़र भी एक Keras लेयर है, हम एक ऐसा नेटवर्क परिभाषित कर सकते हैं जिसमें यह शामिल हो, और इसे एंड-टू-एंड ट्रेन कर सकते हैं। इस तरीके से हमें `map` का उपयोग करके डेटासेट को वेक्टराइज़ करने की आवश्यकता नहीं होगी, हम बस मूल डेटासेट को नेटवर्क के इनपुट में पास कर सकते हैं।\n",
|
||||
"\n",
|
||||
"> **Note**: फिर भी हमें अपने डेटासेट पर `map` लागू करना होगा ताकि डिक्शनरी (जैसे `title`, `description` और `label`) से फील्ड्स को ट्यूपल्स में बदल सकें। हालांकि, जब डिस्क से डेटा लोड कर रहे हों, तो हम शुरुआत में ही आवश्यक संरचना के साथ एक डेटासेट बना सकते हैं।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 13,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Model: \"model\"\n",
|
||||
"_________________________________________________________________\n",
|
||||
" Layer (type) Output Shape Param # \n",
|
||||
"=================================================================\n",
|
||||
" input_1 (InputLayer) [(None, 1)] 0 \n",
|
||||
" \n",
|
||||
" text_vectorization (TextVec (None, None) 0 \n",
|
||||
" torization) \n",
|
||||
" \n",
|
||||
" tf.one_hot (TFOpLambda) (None, None, 5335) 0 \n",
|
||||
" \n",
|
||||
" tf.math.reduce_sum (TFOpLam (None, 5335) 0 \n",
|
||||
" bda) \n",
|
||||
" \n",
|
||||
" dense_2 (Dense) (None, 4) 21344 \n",
|
||||
" \n",
|
||||
"=================================================================\n",
|
||||
"Total params: 21,344\n",
|
||||
"Trainable params: 21,344\n",
|
||||
"Non-trainable params: 0\n",
|
||||
"_________________________________________________________________\n",
|
||||
"938/938 [==============================] - 73s 77ms/step - loss: 0.6057 - acc: 0.8414 - val_loss: 0.4202 - val_acc: 0.8736\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"<keras.callbacks.History at 0x20c721521f0>"
|
||||
]
|
||||
},
|
||||
"execution_count": 13,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"def extract_text(x):\n",
|
||||
" return x['title']+' '+x['description']\n",
|
||||
"\n",
|
||||
"def tupelize(x):\n",
|
||||
" return (extract_text(x),x['label'])\n",
|
||||
"\n",
|
||||
"inp = keras.Input(shape=(1,),dtype=tf.string)\n",
|
||||
"x = vectorizer(inp)\n",
|
||||
"x = tf.reduce_sum(tf.one_hot(x,vocab_size),axis=1)\n",
|
||||
"out = keras.layers.Dense(4,activation='softmax')(x)\n",
|
||||
"model = keras.models.Model(inp,out)\n",
|
||||
"model.summary()\n",
|
||||
"\n",
|
||||
"model.compile(loss='sparse_categorical_crossentropy',optimizer='adam',metrics=['acc'])\n",
|
||||
"model.fit(ds_train.map(tupelize).batch(batch_size),validation_data=ds_test.map(tupelize).batch(batch_size))\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## बाइग्राम, ट्राइग्राम और एन-ग्राम\n",
|
||||
"\n",
|
||||
"बैग-ऑफ-वर्ड्स दृष्टिकोण की एक सीमा यह है कि कुछ शब्द बहु-शब्द अभिव्यक्तियों का हिस्सा होते हैं। उदाहरण के लिए, 'हॉट डॉग' शब्द का अर्थ 'हॉट' और 'डॉग' शब्दों से बिल्कुल अलग होता है। यदि हम हमेशा 'हॉट' और 'डॉग' शब्दों को एक ही वेक्टर का उपयोग करके दर्शाते हैं, तो यह हमारे मॉडल को भ्रमित कर सकता है।\n",
|
||||
"\n",
|
||||
"इस समस्या को हल करने के लिए, **एन-ग्राम प्रतिनिधित्व** का उपयोग अक्सर दस्तावेज़ वर्गीकरण की विधियों में किया जाता है, जहां प्रत्येक शब्द, द्वि-शब्द या त्रि-शब्द की आवृत्ति वर्गीकरण मॉडल को प्रशिक्षित करने के लिए एक उपयोगी विशेषता होती है। उदाहरण के लिए, बाइग्राम प्रतिनिधित्व में, हम मूल शब्दों के अलावा सभी शब्द युग्मों को शब्दावली में जोड़ देंगे।\n",
|
||||
"\n",
|
||||
"नीचे यह दिखाने के लिए एक उदाहरण दिया गया है कि स्कikit Learn का उपयोग करके बाइग्राम बैग-ऑफ-वर्ड्स प्रतिनिधित्व कैसे बनाया जा सकता है:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 14,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Vocabulary:\n",
|
||||
" {'i': 7, 'like': 11, 'hot': 4, 'dogs': 2, 'i like': 8, 'like hot': 12, 'hot dogs': 5, 'the': 16, 'dog': 0, 'ran': 14, 'fast': 3, 'the dog': 17, 'dog ran': 1, 'ran fast': 15, 'its': 9, 'outside': 13, 'its hot': 10, 'hot outside': 6}\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"array([[1, 0, 1, 0, 2, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]],\n",
|
||||
" dtype=int64)"
|
||||
]
|
||||
},
|
||||
"execution_count": 14,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"bigram_vectorizer = CountVectorizer(ngram_range=(1, 2), token_pattern=r'\\b\\w+\\b', min_df=1)\n",
|
||||
"corpus = [\n",
|
||||
" 'I like hot dogs.',\n",
|
||||
" 'The dog ran fast.',\n",
|
||||
" 'Its hot outside.',\n",
|
||||
" ]\n",
|
||||
"bigram_vectorizer.fit_transform(corpus)\n",
|
||||
"print(\"Vocabulary:\\n\",bigram_vectorizer.vocabulary_)\n",
|
||||
"bigram_vectorizer.transform(['My dog likes hot dogs on a hot day.']).toarray()\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"n-gram दृष्टिकोण की मुख्य कमी यह है कि शब्दावली का आकार बहुत तेजी से बढ़ने लगता है। व्यवहार में, हमें n-gram प्रतिनिधित्व को एक आयामीय कमी तकनीक, जैसे *embeddings*, के साथ संयोजित करने की आवश्यकता होती है, जिसे हम अगले यूनिट में चर्चा करेंगे।\n",
|
||||
"\n",
|
||||
"हमारे **AG News** डेटासेट में n-gram प्रतिनिधित्व का उपयोग करने के लिए, हमें `TextVectorization` कंस्ट्रक्टर में `ngrams` पैरामीटर पास करना होगा। एक bigram शब्दावली की लंबाई **काफी बड़ी** होती है, हमारे मामले में यह 1.3 मिलियन से अधिक टोकन है! इसलिए, यह समझदारी होगी कि bigram टोकन को भी किसी उचित संख्या तक सीमित किया जाए।\n",
|
||||
"\n",
|
||||
"हम ऊपर दिए गए कोड का उपयोग करके क्लासिफायर को प्रशिक्षित कर सकते हैं, लेकिन यह मेमोरी के लिहाज से बहुत अक्षम होगा। अगले यूनिट में, हम embeddings का उपयोग करके bigram क्लासिफायर को प्रशिक्षित करेंगे। इस बीच, आप इस नोटबुक में bigram क्लासिफायर प्रशिक्षण के साथ प्रयोग कर सकते हैं और देख सकते हैं कि क्या आप उच्च सटीकता प्राप्त कर सकते हैं।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## BoW वेक्टर स्वचालित रूप से गणना करना\n",
|
||||
"\n",
|
||||
"ऊपर दिए गए उदाहरण में, हमने व्यक्तिगत शब्दों के वन-हॉट एन्कोडिंग को जोड़कर BoW वेक्टर को हाथ से गणना की थी। हालांकि, TensorFlow के नवीनतम संस्करण में, हम BoW वेक्टर को स्वचालित रूप से गणना कर सकते हैं, बस `output_mode='count` पैरामीटर को वेक्टराइज़र कंस्ट्रक्टर में पास करके। यह हमारे मॉडल को परिभाषित और प्रशिक्षित करना काफी आसान बना देता है:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 15,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Training vectorizer\n",
|
||||
"938/938 [==============================] - 7s 7ms/step - loss: 0.5929 - acc: 0.8486 - val_loss: 0.4168 - val_acc: 0.8772\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"<keras.callbacks.History at 0x20c725217c0>"
|
||||
]
|
||||
},
|
||||
"execution_count": 15,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"model = keras.models.Sequential([\n",
|
||||
" keras.layers.experimental.preprocessing.TextVectorization(max_tokens=vocab_size,output_mode='count'),\n",
|
||||
" keras.layers.Dense(4,input_shape=(vocab_size,), activation='softmax')\n",
|
||||
"])\n",
|
||||
"print(\"Training vectorizer\")\n",
|
||||
"model.layers[0].adapt(ds_train.take(500).map(extract_text))\n",
|
||||
"model.compile(loss='sparse_categorical_crossentropy',optimizer='adam',metrics=['acc'])\n",
|
||||
"model.fit(ds_train.map(tupelize).batch(batch_size),validation_data=ds_test.map(tupelize).batch(batch_size))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## टर्म फ्रीक्वेंसी - इनवर्स डॉक्यूमेंट फ्रीक्वेंसी (TF-IDF)\n",
|
||||
"\n",
|
||||
"BoW प्रतिनिधित्व में, शब्दों की उपस्थिति को एक ही तकनीक का उपयोग करके वेट किया जाता है, चाहे वह शब्द कोई भी हो। हालांकि, यह स्पष्ट है कि *a* और *in* जैसे सामान्य शब्द वर्गीकरण के लिए उतने महत्वपूर्ण नहीं होते जितने कि विशेष शब्द। अधिकांश NLP कार्यों में कुछ शब्द दूसरों की तुलना में अधिक प्रासंगिक होते हैं।\n",
|
||||
"\n",
|
||||
"**TF-IDF** का मतलब है **टर्म फ्रीक्वेंसी - इनवर्स डॉक्यूमेंट फ्रीक्वेंसी**। यह बैग-ऑफ-वर्ड्स का एक प्रकार है, जिसमें किसी दस्तावेज़ में शब्द की उपस्थिति को दर्शाने वाले बाइनरी 0/1 मान के बजाय, एक फ्लोटिंग-पॉइंट मान का उपयोग किया जाता है, जो कॉर्पस में शब्द की उपस्थिति की आवृत्ति से संबंधित होता है।\n",
|
||||
"\n",
|
||||
"औपचारिक रूप से, किसी शब्द $i$ का वजन $w_{ij}$ दस्तावेज़ $j$ में इस प्रकार परिभाषित किया गया है:\n",
|
||||
"$$\n",
|
||||
"w_{ij} = tf_{ij}\\times\\log({N\\over df_i})\n",
|
||||
"$$\n",
|
||||
"जहां\n",
|
||||
"* $tf_{ij}$ दस्तावेज़ $j$ में $i$ की उपस्थिति की संख्या है, यानी वह BoW मान जिसे हमने पहले देखा था\n",
|
||||
"* $N$ संग्रह में दस्तावेज़ों की संख्या है\n",
|
||||
"* $df_i$ पूरे संग्रह में शब्द $i$ को शामिल करने वाले दस्तावेज़ों की संख्या है\n",
|
||||
"\n",
|
||||
"TF-IDF मान $w_{ij}$ किसी दस्तावेज़ में शब्द की उपस्थिति की संख्या के अनुपात में बढ़ता है और कॉर्पस में उन दस्तावेज़ों की संख्या से ऑफसेट होता है जिसमें वह शब्द शामिल है। यह इस तथ्य को समायोजित करने में मदद करता है कि कुछ शब्द दूसरों की तुलना में अधिक बार दिखाई देते हैं। उदाहरण के लिए, यदि कोई शब्द *हर* दस्तावेज़ में दिखाई देता है, तो $df_i=N$, और $w_{ij}=0$, और उन शब्दों को पूरी तरह से नजरअंदाज कर दिया जाएगा।\n",
|
||||
"\n",
|
||||
"आप आसानी से Scikit Learn का उपयोग करके टेक्स्ट का TF-IDF वेक्टराइज़ेशन बना सकते हैं:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 16,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"array([[0.43381609, 0. , 0.43381609, 0. , 0.65985664,\n",
|
||||
" 0.43381609, 0. , 0. , 0. , 0. ,\n",
|
||||
" 0. , 0. , 0. , 0. , 0. ,\n",
|
||||
" 0. ]])"
|
||||
]
|
||||
},
|
||||
"execution_count": 16,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"from sklearn.feature_extraction.text import TfidfVectorizer\n",
|
||||
"vectorizer = TfidfVectorizer(ngram_range=(1,2))\n",
|
||||
"vectorizer.fit_transform(corpus)\n",
|
||||
"vectorizer.transform(['My dog likes hot dogs on a hot day.']).toarray()"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"Keras में, `TextVectorization` लेयर `output_mode='tf-idf'` पैरामीटर पास करके स्वचालित रूप से TF-IDF आवृत्तियों की गणना कर सकती है। आइए ऊपर उपयोग किए गए कोड को दोहराते हैं ताकि देख सकें कि TF-IDF का उपयोग करने से सटीकता बढ़ती है या नहीं:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 17,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Training vectorizer\n",
|
||||
"938/938 [==============================] - 12s 12ms/step - loss: 0.4197 - acc: 0.8662 - val_loss: 0.3432 - val_acc: 0.8849\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"<keras.callbacks.History at 0x20c729dfd30>"
|
||||
]
|
||||
},
|
||||
"execution_count": 17,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"model = keras.models.Sequential([\n",
|
||||
" keras.layers.experimental.preprocessing.TextVectorization(max_tokens=vocab_size,output_mode='tf-idf'),\n",
|
||||
" keras.layers.Dense(4,input_shape=(vocab_size,), activation='softmax')\n",
|
||||
"])\n",
|
||||
"print(\"Training vectorizer\")\n",
|
||||
"model.layers[0].adapt(ds_train.take(500).map(extract_text))\n",
|
||||
"model.compile(loss='sparse_categorical_crossentropy',optimizer='adam',metrics=['acc'])\n",
|
||||
"model.fit(ds_train.map(tupelize).batch(batch_size),validation_data=ds_test.map(tupelize).batch(batch_size))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## निष्कर्ष\n",
|
||||
"\n",
|
||||
"हालांकि TF-IDF प्रतिनिधित्व विभिन्न शब्दों को आवृत्ति भार प्रदान करते हैं, वे न तो अर्थ को व्यक्त कर सकते हैं और न ही क्रम को। जैसा कि प्रसिद्ध भाषाविद् जे. आर. फर्थ ने 1935 में कहा था, \"शब्द का पूर्ण अर्थ हमेशा संदर्भात्मक होता है, और संदर्भ से अलग अर्थ का कोई भी अध्ययन गंभीरता से नहीं लिया जा सकता।\" हम इस पाठ्यक्रम में आगे भाषा मॉडलिंग का उपयोग करके पाठ से संदर्भात्मक जानकारी को कैप्चर करना सीखेंगे।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"\n---\n\n**अस्वीकरण**: \nयह दस्तावेज़ AI अनुवाद सेवा [Co-op Translator](https://github.com/Azure/co-op-translator) का उपयोग करके अनुवादित किया गया है। जबकि हम सटीकता सुनिश्चित करने का प्रयास करते हैं, कृपया ध्यान दें कि स्वचालित अनुवाद में त्रुटियां या अशुद्धियां हो सकती हैं। मूल भाषा में उपलब्ध मूल दस्तावेज़ को प्रामाणिक स्रोत माना जाना चाहिए। महत्वपूर्ण जानकारी के लिए, पेशेवर मानव अनुवाद की सिफारिश की जाती है। इस अनुवाद के उपयोग से उत्पन्न किसी भी गलतफहमी या गलत व्याख्या के लिए हम उत्तरदायी नहीं हैं।\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
"interpreter": {
|
||||
"hash": "0cb620c6d4b9f7a635928804c26cf22403d89d98d79684e4529119355ee6d5a5"
|
||||
},
|
||||
"kernel_info": {
|
||||
"name": "conda-env-py37_tensorflow-py"
|
||||
},
|
||||
"kernelspec": {
|
||||
"display_name": "py37_tensorflow",
|
||||
"language": "python",
|
||||
"name": "python3"
|
||||
},
|
||||
"language_info": {
|
||||
"codemirror_mode": {
|
||||
"name": "ipython",
|
||||
"version": 3
|
||||
},
|
||||
"file_extension": ".py",
|
||||
"mimetype": "text/x-python",
|
||||
"name": "python",
|
||||
"nbconvert_exporter": "python",
|
||||
"pygments_lexer": "ipython3",
|
||||
"version": "3.8.12"
|
||||
},
|
||||
"nteract": {
|
||||
"version": "nteract-front-end@1.0.0"
|
||||
},
|
||||
"coopTranslator": {
|
||||
"original_hash": "19b43951d55b377a76209c24c1f017e4",
|
||||
"translation_date": "2025-08-31T15:31:44+00:00",
|
||||
"source_file": "lessons/5-NLP/13-TextRep/TextRepresentationTF.ipynb",
|
||||
"language_code": "hi"
|
||||
}
|
||||
},
|
||||
"nbformat": 4,
|
||||
"nbformat_minor": 4
|
||||
}
|
||||
|
|
@ -0,0 +1,720 @@
|
|||
{
|
||||
"cells": [
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## एम्बेडिंग्स\n",
|
||||
"\n",
|
||||
"हमारे पिछले उदाहरण में, हमने उच्च-आयामी बैग-ऑफ-वर्ड्स वेक्टर पर काम किया था, जिसकी लंबाई `vocab_size` थी, और हम स्पष्ट रूप से निम्न-आयामी स्थानिक प्रतिनिधित्व वेक्टर को विरल वन-हॉट प्रतिनिधित्व में परिवर्तित कर रहे थे। यह वन-हॉट प्रतिनिधित्व मेमोरी-कुशल नहीं है, इसके अलावा, प्रत्येक शब्द को एक-दूसरे से स्वतंत्र रूप से माना जाता है, यानी वन-हॉट एन्कोडेड वेक्टर शब्दों के बीच किसी भी अर्थपूर्ण समानता को व्यक्त नहीं करते हैं।\n",
|
||||
"\n",
|
||||
"इस इकाई में, हम **News AG** डेटासेट का अन्वेषण जारी रखेंगे। शुरू करने के लिए, आइए डेटा लोड करें और पिछले नोटबुक से कुछ परिभाषाएँ प्राप्त करें।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 1,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Loading dataset...\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"name": "stderr",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"d:\\WORK\\ai-for-beginners\\5-NLP\\14-Embeddings\\data\\train.csv: 29.5MB [00:01, 18.8MB/s] \n",
|
||||
"d:\\WORK\\ai-for-beginners\\5-NLP\\14-Embeddings\\data\\test.csv: 1.86MB [00:00, 11.2MB/s] \n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Building vocab...\n",
|
||||
"Vocab size = 95812\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"import torch\n",
|
||||
"import torchtext\n",
|
||||
"import numpy as np\n",
|
||||
"from torchnlp import *\n",
|
||||
"train_dataset, test_dataset, classes, vocab = load_dataset()\n",
|
||||
"vocab_size = len(vocab)\n",
|
||||
"print(\"Vocab size = \",vocab_size)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## एम्बेडिंग क्या है?\n",
|
||||
"\n",
|
||||
"**एम्बेडिंग** का विचार यह है कि शब्दों को निम्न-आयामी घने वेक्टरों द्वारा दर्शाया जाए, जो किसी तरह शब्द के अर्थ को प्रतिबिंबित करते हैं। हम बाद में चर्चा करेंगे कि सार्थक शब्द एम्बेडिंग कैसे बनाई जाए, लेकिन फिलहाल, एम्बेडिंग को शब्द वेक्टर की आयाम संख्या को कम करने के तरीके के रूप में सोचें।\n",
|
||||
"\n",
|
||||
"इस प्रकार, एम्बेडिंग लेयर एक शब्द को इनपुट के रूप में लेगी और निर्दिष्ट `embedding_size` का आउटपुट वेक्टर उत्पन्न करेगी। एक अर्थ में, यह `Linear` लेयर के समान है, लेकिन एक-हॉट एनकोडेड वेक्टर लेने के बजाय, यह एक शब्द संख्या को इनपुट के रूप में ले सकेगी।\n",
|
||||
"\n",
|
||||
"हमारे नेटवर्क में एम्बेडिंग लेयर को पहली लेयर के रूप में उपयोग करके, हम बैग-ऑफ-वर्ड्स से **एम्बेडिंग बैग** मॉडल में स्विच कर सकते हैं, जहां हम पहले अपने टेक्स्ट के प्रत्येक शब्द को संबंधित एम्बेडिंग में बदलते हैं, और फिर उन सभी एम्बेडिंग पर कुछ समग्र फ़ंक्शन की गणना करते हैं, जैसे `sum`, `average` या `max`।\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"हमारा क्लासिफायर न्यूरल नेटवर्क एम्बेडिंग लेयर से शुरू होगा, फिर एग्रीगेशन लेयर, और उसके ऊपर एक लीनियर क्लासिफायर:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 2,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"class EmbedClassifier(torch.nn.Module):\n",
|
||||
" def __init__(self, vocab_size, embed_dim, num_class):\n",
|
||||
" super().__init__()\n",
|
||||
" self.embedding = torch.nn.Embedding(vocab_size, embed_dim)\n",
|
||||
" self.fc = torch.nn.Linear(embed_dim, num_class)\n",
|
||||
"\n",
|
||||
" def forward(self, x):\n",
|
||||
" x = self.embedding(x)\n",
|
||||
" x = torch.mean(x,dim=1)\n",
|
||||
" return self.fc(x)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"### चर अनुक्रम आकार से निपटना\n",
|
||||
"\n",
|
||||
"इस आर्किटेक्चर के परिणामस्वरूप, हमारे नेटवर्क के लिए मिनीबैच को एक विशेष तरीके से बनाना होगा। पिछले यूनिट में, जब बैग-ऑफ-वर्ड्स (BoW) का उपयोग कर रहे थे, तो मिनीबैच में सभी BoW टेंसर का आकार `vocab_size` के बराबर होता था, चाहे हमारे टेक्स्ट अनुक्रम की वास्तविक लंबाई कुछ भी हो। जब हम वर्ड एम्बेडिंग्स पर जाते हैं, तो प्रत्येक टेक्स्ट सैंपल में शब्दों की संख्या अलग-अलग हो सकती है, और इन सैंपल्स को मिनीबैच में मिलाने के दौरान हमें कुछ पैडिंग लागू करनी होगी।\n",
|
||||
"\n",
|
||||
"यह `collate_fn` फ़ंक्शन को डेटा स्रोत में प्रदान करने की उसी तकनीक का उपयोग करके किया जा सकता है:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 3,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"def padify(b):\n",
|
||||
" # b is the list of tuples of length batch_size\n",
|
||||
" # - first element of a tuple = label, \n",
|
||||
" # - second = feature (text sequence)\n",
|
||||
" # build vectorized sequence\n",
|
||||
" v = [encode(x[1]) for x in b]\n",
|
||||
" # first, compute max length of a sequence in this minibatch\n",
|
||||
" l = max(map(len,v))\n",
|
||||
" return ( # tuple of two tensors - labels and features\n",
|
||||
" torch.LongTensor([t[0]-1 for t in b]),\n",
|
||||
" torch.stack([torch.nn.functional.pad(torch.tensor(t),(0,l-len(t)),mode='constant',value=0) for t in v])\n",
|
||||
" )\n",
|
||||
"\n",
|
||||
"train_loader = torch.utils.data.DataLoader(train_dataset, batch_size=16, collate_fn=padify, shuffle=True)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"### एम्बेडिंग क्लासिफायर का प्रशिक्षण\n",
|
||||
"\n",
|
||||
"अब जब हमने सही डाटालोडर परिभाषित कर लिया है, तो हम पिछले यूनिट में परिभाषित प्रशिक्षण फ़ंक्शन का उपयोग करके मॉडल को प्रशिक्षित कर सकते हैं:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"3200: acc=0.6415625\n",
|
||||
"6400: acc=0.6865625\n",
|
||||
"9600: acc=0.7103125\n",
|
||||
"12800: acc=0.726953125\n",
|
||||
"16000: acc=0.739375\n",
|
||||
"19200: acc=0.75046875\n",
|
||||
"22400: acc=0.7572321428571429\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"(0.889799795315499, 0.7623160588611644)"
|
||||
]
|
||||
},
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"net = EmbedClassifier(vocab_size,32,len(classes)).to(device)\n",
|
||||
"train_epoch(net,train_loader, lr=1, epoch_size=25000)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"> **नोट**: हम यहां केवल 25k रिकॉर्ड्स (एक पूर्ण युग से कम) के लिए प्रशिक्षण कर रहे हैं समय बचाने के लिए, लेकिन आप प्रशिक्षण जारी रख सकते हैं, कई युगों के लिए प्रशिक्षण के लिए एक फ़ंक्शन लिख सकते हैं, और उच्च सटीकता प्राप्त करने के लिए लर्निंग रेट पैरामीटर के साथ प्रयोग कर सकते हैं। आपको लगभग 90% सटीकता तक पहुंचने में सक्षम होना चाहिए।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"### EmbeddingBag लेयर और चर लंबाई अनुक्रम प्रतिनिधित्व\n",
|
||||
"\n",
|
||||
"पिछली संरचना में, हमें सभी अनुक्रमों को एक ही लंबाई तक बढ़ाना पड़ता था ताकि उन्हें मिनीबैच में फिट किया जा सके। यह चर लंबाई अनुक्रमों को प्रस्तुत करने का सबसे प्रभावी तरीका नहीं है - एक अन्य दृष्टिकोण **ऑफसेट** वेक्टर का उपयोग करना होगा, जो एक बड़े वेक्टर में संग्रहीत सभी अनुक्रमों के ऑफसेट को रखेगा।\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"> **Note**: ऊपर दी गई तस्वीर में, हमने अक्षरों के अनुक्रम को दिखाया है, लेकिन हमारे उदाहरण में हम शब्दों के अनुक्रमों के साथ काम कर रहे हैं। हालांकि, ऑफसेट वेक्टर के साथ अनुक्रमों को प्रस्तुत करने का सामान्य सिद्धांत वही रहता है।\n",
|
||||
"\n",
|
||||
"ऑफसेट प्रतिनिधित्व के साथ काम करने के लिए, हम [`EmbeddingBag`](https://pytorch.org/docs/stable/generated/torch.nn.EmbeddingBag.html) लेयर का उपयोग करते हैं। यह `Embedding` के समान है, लेकिन यह सामग्री वेक्टर और ऑफसेट वेक्टर को इनपुट के रूप में लेता है, और इसमें औसत लेयर भी शामिल होती है, जो `mean`, `sum` या `max` हो सकती है।\n",
|
||||
"\n",
|
||||
"यहाँ एक संशोधित नेटवर्क है जो `EmbeddingBag` का उपयोग करता है:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 5,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"class EmbedClassifier(torch.nn.Module):\n",
|
||||
" def __init__(self, vocab_size, embed_dim, num_class):\n",
|
||||
" super().__init__()\n",
|
||||
" self.embedding = torch.nn.EmbeddingBag(vocab_size, embed_dim)\n",
|
||||
" self.fc = torch.nn.Linear(embed_dim, num_class)\n",
|
||||
"\n",
|
||||
" def forward(self, text, off):\n",
|
||||
" x = self.embedding(text, off)\n",
|
||||
" return self.fc(x)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"डेटासेट को प्रशिक्षण के लिए तैयार करने के लिए, हमें एक रूपांतरण फ़ंक्शन प्रदान करना होगा जो ऑफ़सेट वेक्टर तैयार करेगा:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 6,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"def offsetify(b):\n",
|
||||
" # first, compute data tensor from all sequences\n",
|
||||
" x = [torch.tensor(encode(t[1])) for t in b]\n",
|
||||
" # now, compute the offsets by accumulating the tensor of sequence lengths\n",
|
||||
" o = [0] + [len(t) for t in x]\n",
|
||||
" o = torch.tensor(o[:-1]).cumsum(dim=0)\n",
|
||||
" return ( \n",
|
||||
" torch.LongTensor([t[0]-1 for t in b]), # labels\n",
|
||||
" torch.cat(x), # text \n",
|
||||
" o\n",
|
||||
" )\n",
|
||||
"\n",
|
||||
"train_loader = torch.utils.data.DataLoader(train_dataset, batch_size=16, collate_fn=offsetify, shuffle=True)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"ध्यान दें कि सभी पिछले उदाहरणों के विपरीत, हमारा नेटवर्क अब दो पैरामीटर स्वीकार करता है: डेटा वेक्टर और ऑफसेट वेक्टर, जो अलग-अलग आकार के होते हैं। इसी तरह, हमारा डेटा लोडर भी हमें 2 के बजाय 3 मान प्रदान करता है: टेक्स्ट और ऑफसेट वेक्टर दोनों को फीचर्स के रूप में प्रदान किया जाता है। इसलिए, हमें अपने प्रशिक्षण फ़ंक्शन को थोड़ा समायोजित करने की आवश्यकता है ताकि इसका ध्यान रखा जा सके:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 7,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"3200: acc=0.6153125\n",
|
||||
"6400: acc=0.6615625\n",
|
||||
"9600: acc=0.6932291666666667\n",
|
||||
"12800: acc=0.715078125\n",
|
||||
"16000: acc=0.7270625\n",
|
||||
"19200: acc=0.7382291666666667\n",
|
||||
"22400: acc=0.7486160714285715\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"(22.771553103007037, 0.7551983365323096)"
|
||||
]
|
||||
},
|
||||
"execution_count": 7,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"net = EmbedClassifier(vocab_size,32,len(classes)).to(device)\n",
|
||||
"\n",
|
||||
"def train_epoch_emb(net,dataloader,lr=0.01,optimizer=None,loss_fn = torch.nn.CrossEntropyLoss(),epoch_size=None, report_freq=200):\n",
|
||||
" optimizer = optimizer or torch.optim.Adam(net.parameters(),lr=lr)\n",
|
||||
" loss_fn = loss_fn.to(device)\n",
|
||||
" net.train()\n",
|
||||
" total_loss,acc,count,i = 0,0,0,0\n",
|
||||
" for labels,text,off in dataloader:\n",
|
||||
" optimizer.zero_grad()\n",
|
||||
" labels,text,off = labels.to(device), text.to(device), off.to(device)\n",
|
||||
" out = net(text, off)\n",
|
||||
" loss = loss_fn(out,labels) #cross_entropy(out,labels)\n",
|
||||
" loss.backward()\n",
|
||||
" optimizer.step()\n",
|
||||
" total_loss+=loss\n",
|
||||
" _,predicted = torch.max(out,1)\n",
|
||||
" acc+=(predicted==labels).sum()\n",
|
||||
" count+=len(labels)\n",
|
||||
" i+=1\n",
|
||||
" if i%report_freq==0:\n",
|
||||
" print(f\"{count}: acc={acc.item()/count}\")\n",
|
||||
" if epoch_size and count>epoch_size:\n",
|
||||
" break\n",
|
||||
" return total_loss.item()/count, acc.item()/count\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"train_epoch_emb(net,train_loader, lr=4, epoch_size=25000)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## सेमांटिक एम्बेडिंग्स: वर्ड2वेक\n",
|
||||
"\n",
|
||||
"पिछले उदाहरण में, मॉडल एम्बेडिंग लेयर ने शब्दों को वेक्टर प्रतिनिधित्व में मैप करना सीखा, लेकिन इस प्रतिनिधित्व में ज्यादा सेमांटिक अर्थ नहीं था। यह अच्छा होगा कि ऐसा वेक्टर प्रतिनिधित्व सीखा जाए, जिसमें समान शब्द या पर्यायवाची शब्द ऐसे वेक्टर से मेल खाएं जो किसी वेक्टर दूरी (जैसे, यूक्लिडियन दूरी) के संदर्भ में एक-दूसरे के करीब हों।\n",
|
||||
"\n",
|
||||
"इसके लिए, हमें अपने एम्बेडिंग मॉडल को एक बड़े टेक्स्ट संग्रह पर एक विशेष तरीके से प्री-ट्रेन करना होगा। सेमांटिक एम्बेडिंग्स को ट्रेन करने के शुरुआती तरीकों में से एक को [वर्ड2वेक](https://en.wikipedia.org/wiki/Word2vec) कहा जाता है। यह शब्दों के वितरित प्रतिनिधित्व को उत्पन्न करने के लिए दो मुख्य आर्किटेक्चर पर आधारित है:\n",
|
||||
"\n",
|
||||
" - **कंटीन्युअस बैग-ऑफ-वर्ड्स** (CBoW) — इस आर्किटेक्चर में, हम मॉडल को आस-पास के संदर्भ से एक शब्द की भविष्यवाणी करने के लिए ट्रेन करते हैं। दिए गए ngram $(W_{-2},W_{-1},W_0,W_1,W_2)$ में, मॉडल का लक्ष्य $(W_{-2},W_{-1},W_1,W_2)$ से $W_0$ की भविष्यवाणी करना है।\n",
|
||||
" - **कंटीन्युअस स्किप-ग्राम** CBoW के विपरीत है। मॉडल संदर्भ शब्दों की आस-पास की विंडो का उपयोग करके वर्तमान शब्द की भविष्यवाणी करता है।\n",
|
||||
"\n",
|
||||
"CBoW तेज है, जबकि स्किप-ग्राम धीमा है, लेकिन यह कम बार उपयोग होने वाले शब्दों का बेहतर प्रतिनिधित्व करता है।\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"Google News डेटासेट पर प्री-ट्रेन किए गए वर्ड2वेक एम्बेडिंग के साथ प्रयोग करने के लिए, हम **gensim** लाइब्रेरी का उपयोग कर सकते हैं। नीचे हम 'neural' के सबसे समान शब्दों को ढूंढते हैं।\n",
|
||||
"\n",
|
||||
"> **Note:** जब आप पहली बार शब्द वेक्टर बनाते हैं, तो उन्हें डाउनलोड करने में कुछ समय लग सकता है!\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 8,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"import gensim.downloader as api\n",
|
||||
"w2v = api.load('word2vec-google-news-300')"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 9,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"neuronal -> 0.7804799675941467\n",
|
||||
"neurons -> 0.7326500415802002\n",
|
||||
"neural_circuits -> 0.7252851724624634\n",
|
||||
"neuron -> 0.7174385190010071\n",
|
||||
"cortical -> 0.6941086649894714\n",
|
||||
"brain_circuitry -> 0.6923246383666992\n",
|
||||
"synaptic -> 0.6699118614196777\n",
|
||||
"neural_circuitry -> 0.6638563275337219\n",
|
||||
"neurochemical -> 0.6555314064025879\n",
|
||||
"neuronal_activity -> 0.6531826257705688\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"for w,p in w2v.most_similar('neural'):\n",
|
||||
" print(f\"{w} -> {p}\")"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"हम शब्द से वेक्टर एम्बेडिंग भी गणना कर सकते हैं, जिसे वर्गीकरण मॉडल के प्रशिक्षण में उपयोग किया जा सकता है (स्पष्टता के लिए हम केवल वेक्टर के पहले 20 घटक दिखाते हैं):\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 10,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"array([ 0.01226807, 0.06225586, 0.10693359, 0.05810547, 0.23828125,\n",
|
||||
" 0.03686523, 0.05151367, -0.20703125, 0.01989746, 0.10058594,\n",
|
||||
" -0.03759766, -0.1015625 , -0.15820312, -0.08105469, -0.0390625 ,\n",
|
||||
" -0.05053711, 0.16015625, 0.2578125 , 0.10058594, -0.25976562],\n",
|
||||
" dtype=float32)"
|
||||
]
|
||||
},
|
||||
"execution_count": 10,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"w2v.word_vec('play')[:20]"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": []
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 10,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"('queen', 0.7118192911148071)"
|
||||
]
|
||||
},
|
||||
"execution_count": 10,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"w2v.most_similar(positive=['king','woman'],negative=['man'])[0]"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"CBoW और Skip-Grams दोनों \"predictive\" embeddings हैं, क्योंकि ये केवल स्थानीय संदर्भों को ध्यान में रखते हैं। Word2Vec वैश्विक संदर्भ का लाभ नहीं उठाता है।\n",
|
||||
"\n",
|
||||
"**FastText**, Word2Vec पर आधारित है और प्रत्येक शब्द और उसमें पाए जाने वाले अक्षर n-grams के लिए वेक्टर प्रतिनिधित्व सीखता है। इन प्रतिनिधित्वों के मानों को प्रत्येक प्रशिक्षण चरण में एक वेक्टर में औसत किया जाता है। हालांकि यह प्री-ट्रेनिंग में काफी अतिरिक्त गणना जोड़ता है, लेकिन यह वर्ड एम्बेडिंग्स को सब-वर्ड जानकारी को एन्कोड करने में सक्षम बनाता है।\n",
|
||||
"\n",
|
||||
"एक और विधि, **GloVe**, सह-अस्तित्व मैट्रिक्स (co-occurrence matrix) के विचार का उपयोग करती है और सह-अस्तित्व मैट्रिक्स को अधिक अभिव्यक्तिपूर्ण और गैर-रेखीय (non-linear) वर्ड वेक्टर में विभाजित करने के लिए न्यूरल विधियों का उपयोग करती है।\n",
|
||||
"\n",
|
||||
"आप उदाहरण के साथ खेल सकते हैं और embeddings को FastText और GloVe में बदल सकते हैं, क्योंकि gensim कई अलग-अलग वर्ड एम्बेडिंग मॉडल का समर्थन करता है।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## PyTorch में प्री-ट्रेंड एम्बेडिंग्स का उपयोग करना\n",
|
||||
"\n",
|
||||
"हम ऊपर दिए गए उदाहरण को संशोधित कर सकते हैं ताकि हमारे एम्बेडिंग लेयर के मैट्रिक्स को पहले से तैयार किए गए सेमांटिकल एम्बेडिंग्स, जैसे Word2Vec, से प्री-पॉप्युलेट किया जा सके। हमें यह ध्यान रखना होगा कि प्री-ट्रेंड एम्बेडिंग और हमारे टेक्स्ट कॉर्पस की वोकैब्युलरी शायद मेल नहीं खाएगी, इसलिए हम उन शब्दों के लिए वेट्स को रैंडम वैल्यू से इनिशियलाइज़ करेंगे जो गायब हैं:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 11,
|
||||
"metadata": {
|
||||
"tags": []
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Embedding size: 300\n",
|
||||
"Populating matrix, this will take some time...Done, found 41080 words, 54732 words missing\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"embed_size = len(w2v.get_vector('hello'))\n",
|
||||
"print(f'Embedding size: {embed_size}')\n",
|
||||
"\n",
|
||||
"net = EmbedClassifier(vocab_size,embed_size,len(classes))\n",
|
||||
"\n",
|
||||
"print('Populating matrix, this will take some time...',end='')\n",
|
||||
"found, not_found = 0,0\n",
|
||||
"for i,w in enumerate(vocab.get_itos()):\n",
|
||||
" try:\n",
|
||||
" net.embedding.weight[i].data = torch.tensor(w2v.get_vector(w))\n",
|
||||
" found+=1\n",
|
||||
" except:\n",
|
||||
" net.embedding.weight[i].data = torch.normal(0.0,1.0,(embed_size,))\n",
|
||||
" not_found+=1\n",
|
||||
"\n",
|
||||
"print(f\"Done, found {found} words, {not_found} words missing\")\n",
|
||||
"net = net.to(device)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": []
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 12,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"3200: acc=0.6359375\n",
|
||||
"6400: acc=0.68109375\n",
|
||||
"9600: acc=0.7067708333333333\n",
|
||||
"12800: acc=0.723671875\n",
|
||||
"16000: acc=0.73625\n",
|
||||
"19200: acc=0.7463541666666667\n",
|
||||
"22400: acc=0.7560714285714286\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"(214.1013875559821, 0.7626759436980166)"
|
||||
]
|
||||
},
|
||||
"execution_count": 12,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"train_epoch_emb(net,train_loader, lr=4, epoch_size=25000)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"हमारे मामले में, हमें सटीकता में बहुत अधिक वृद्धि नहीं दिखती है, जो संभवतः अलग-अलग शब्दावली के कारण है। \n",
|
||||
"अलग-अलग शब्दावली की समस्या को हल करने के लिए, हम निम्नलिखित समाधानों में से एक का उपयोग कर सकते हैं: \n",
|
||||
"* हमारे शब्दावली पर word2vec मॉडल को पुनः प्रशिक्षित करें \n",
|
||||
"* प्री-ट्रेंड word2vec मॉडल की शब्दावली के साथ हमारा डेटासेट लोड करें। डेटासेट को लोड करने के लिए उपयोग की जाने वाली शब्दावली को लोडिंग के दौरान निर्दिष्ट किया जा सकता है। \n",
|
||||
"\n",
|
||||
"दूसरा तरीका अधिक आसान लगता है, खासकर क्योंकि PyTorch `torchtext` फ्रेमवर्क में एम्बेडिंग के लिए बिल्ट-इन सपोर्ट है। उदाहरण के लिए, हम निम्नलिखित तरीके से GloVe-आधारित शब्दावली को इंस्टैंशिएट कर सकते हैं: \n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 14,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stderr",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"100%|█████████▉| 399999/400000 [00:15<00:00, 25411.14it/s]\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"vocab = torchtext.vocab.GloVe(name='6B', dim=50)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"लोडेड शब्दावली में निम्नलिखित बुनियादी ऑपरेशन्स होते हैं:\n",
|
||||
"* `vocab.stoi` डिक्शनरी हमें किसी शब्द को उसके डिक्शनरी इंडेक्स में बदलने की अनुमति देती है\n",
|
||||
"* `vocab.itos` इसका उल्टा करता है - नंबर को शब्द में बदलता है\n",
|
||||
"* `vocab.vectors` एम्बेडिंग वेक्टर का ऐरे है, इसलिए किसी शब्द `s` की एम्बेडिंग प्राप्त करने के लिए हमें `vocab.vectors[vocab.stoi[s]]` का उपयोग करना होगा\n",
|
||||
"\n",
|
||||
"यहाँ एम्बेडिंग्स के साथ छेड़छाड़ का एक उदाहरण है, जो समीकरण **kind-man+woman = queen** को प्रदर्शित करता है (मुझे इसे काम करने के लिए कोएफिशिएंट को थोड़ा समायोजित करना पड़ा):\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 15,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"'queen'"
|
||||
]
|
||||
},
|
||||
"execution_count": 15,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"# get the vector corresponding to kind-man+woman\n",
|
||||
"qvec = vocab.vectors[vocab.stoi['king']]-vocab.vectors[vocab.stoi['man']]+1.3*vocab.vectors[vocab.stoi['woman']]\n",
|
||||
"# find the index of the closest embedding vector \n",
|
||||
"d = torch.sum((vocab.vectors-qvec)**2,dim=1)\n",
|
||||
"min_idx = torch.argmin(d)\n",
|
||||
"# find the corresponding word\n",
|
||||
"vocab.itos[min_idx]"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"उन एम्बेडिंग्स का उपयोग करके वर्गीकरणकर्ता को प्रशिक्षित करने के लिए, हमें पहले अपने डेटासेट को GloVe शब्दावली का उपयोग करके एन्कोड करना होगा:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 16,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"def offsetify(b):\n",
|
||||
" # first, compute data tensor from all sequences\n",
|
||||
" x = [torch.tensor(encode(t[1],voc=vocab)) for t in b] # pass the instance of vocab to encode function!\n",
|
||||
" # now, compute the offsets by accumulating the tensor of sequence lengths\n",
|
||||
" o = [0] + [len(t) for t in x]\n",
|
||||
" o = torch.tensor(o[:-1]).cumsum(dim=0)\n",
|
||||
" return ( \n",
|
||||
" torch.LongTensor([t[0]-1 for t in b]), # labels\n",
|
||||
" torch.cat(x), # text \n",
|
||||
" o\n",
|
||||
" )"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"जैसा कि हमने ऊपर देखा, सभी वेक्टर एम्बेडिंग्स `vocab.vectors` मैट्रिक्स में संग्रहीत होती हैं। इसे एम्बेडिंग लेयर के वेट्स में सरल कॉपीिंग के माध्यम से लोड करना बहुत आसान बनाता है:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 17,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"net = EmbedClassifier(len(vocab),len(vocab.vectors[0]),len(classes))\n",
|
||||
"net.embedding.weight.data = vocab.vectors\n",
|
||||
"net = net.to(device)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": []
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 18,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"3200: acc=0.6271875\n",
|
||||
"6400: acc=0.68078125\n",
|
||||
"9600: acc=0.7030208333333333\n",
|
||||
"12800: acc=0.71984375\n",
|
||||
"16000: acc=0.7346875\n",
|
||||
"19200: acc=0.7455729166666667\n",
|
||||
"22400: acc=0.7529464285714286\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"(35.53972978646833, 0.7575175943698017)"
|
||||
]
|
||||
},
|
||||
"execution_count": 18,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"train_loader = torch.utils.data.DataLoader(train_dataset, batch_size=16, collate_fn=offsetify, shuffle=True)\n",
|
||||
"train_epoch_emb(net,train_loader, lr=4, epoch_size=25000)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"हमारी सटीकता में महत्वपूर्ण वृद्धि न देखने के कारणों में से एक यह है कि हमारे डेटासेट के कुछ शब्द प्री-ट्रेंड GloVe शब्दावली में नहीं हैं, और इसलिए उन्हें अनदेखा कर दिया जाता है। इस तथ्य को दूर करने के लिए, हम अपने डेटासेट पर अपने स्वयं के एम्बेडिंग्स को प्रशिक्षित कर सकते हैं।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## संदर्भात्मक एम्बेडिंग्स\n",
|
||||
"\n",
|
||||
"पारंपरिक प्रीट्रेंड एम्बेडिंग प्रतिनिधित्व, जैसे Word2Vec, की एक मुख्य सीमा शब्दार्थ अस्पष्टता (word sense disambiguation) की समस्या है। जबकि प्रीट्रेंड एम्बेडिंग्स शब्दों के कुछ अर्थ को संदर्भ में पकड़ सकते हैं, एक शब्द के हर संभावित अर्थ को एक ही एम्बेडिंग में एन्कोड किया जाता है। यह डाउनस्ट्रीम मॉडल्स में समस्याएं पैदा कर सकता है, क्योंकि कई शब्दों, जैसे 'play', के अलग-अलग संदर्भों में अलग-अलग अर्थ हो सकते हैं।\n",
|
||||
"\n",
|
||||
"उदाहरण के लिए, 'play' शब्द का इन दो वाक्यों में काफी अलग अर्थ है:\n",
|
||||
"- मैं थिएटर में एक **play** देखने गया।\n",
|
||||
"- जॉन अपने दोस्तों के साथ **play** करना चाहता है।\n",
|
||||
"\n",
|
||||
"ऊपर दिए गए प्रीट्रेंड एम्बेडिंग्स 'play' शब्द के इन दोनों अर्थों को एक ही एम्बेडिंग में दर्शाते हैं। इस सीमा को दूर करने के लिए, हमें **भाषा मॉडल** पर आधारित एम्बेडिंग्स बनानी होंगी, जो एक बड़े टेक्स्ट कॉर्पस पर प्रशिक्षित होता है और *जानता है* कि शब्दों को विभिन्न संदर्भों में कैसे जोड़ा जा सकता है। संदर्भात्मक एम्बेडिंग्स पर चर्चा करना इस ट्यूटोरियल के दायरे से बाहर है, लेकिन हम अगले यूनिट में भाषा मॉडलों पर बात करते समय इस पर वापस आएंगे।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"\n---\n\n**अस्वीकरण**: \nयह दस्तावेज़ AI अनुवाद सेवा [Co-op Translator](https://github.com/Azure/co-op-translator) का उपयोग करके अनुवादित किया गया है। जबकि हम सटीकता सुनिश्चित करने का प्रयास करते हैं, कृपया ध्यान दें कि स्वचालित अनुवाद में त्रुटियां या अशुद्धियां हो सकती हैं। मूल भाषा में उपलब्ध मूल दस्तावेज़ को प्रामाणिक स्रोत माना जाना चाहिए। महत्वपूर्ण जानकारी के लिए, पेशेवर मानव अनुवाद की सिफारिश की जाती है। इस अनुवाद के उपयोग से उत्पन्न किसी भी गलतफहमी या गलत व्याख्या के लिए हम उत्तरदायी नहीं हैं।\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
"interpreter": {
|
||||
"hash": "0cb620c6d4b9f7a635928804c26cf22403d89d98d79684e4529119355ee6d5a5"
|
||||
},
|
||||
"kernelspec": {
|
||||
"display_name": "py37_pytorch",
|
||||
"language": "python",
|
||||
"name": "python3"
|
||||
},
|
||||
"language_info": {
|
||||
"codemirror_mode": {
|
||||
"name": "ipython",
|
||||
"version": 3
|
||||
},
|
||||
"file_extension": ".py",
|
||||
"mimetype": "text/x-python",
|
||||
"name": "python",
|
||||
"nbconvert_exporter": "python",
|
||||
"pygments_lexer": "ipython3",
|
||||
"version": "3.8.12"
|
||||
},
|
||||
"coopTranslator": {
|
||||
"original_hash": "f50b026abce5cf36783a560ea72cb9b1",
|
||||
"translation_date": "2025-08-31T15:28:25+00:00",
|
||||
"source_file": "lessons/5-NLP/14-Embeddings/EmbeddingsPyTorch.ipynb",
|
||||
"language_code": "hi"
|
||||
}
|
||||
},
|
||||
"nbformat": 4,
|
||||
"nbformat_minor": 2
|
||||
}
|
||||
|
|
@ -0,0 +1,695 @@
|
|||
{
|
||||
"cells": [
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## एम्बेडिंग्स\n",
|
||||
"\n",
|
||||
"हमारे पिछले उदाहरण में, हमने `vocab_size` लंबाई वाले उच्च-आयामी बैग-ऑफ-वर्ड्स वेक्टर पर काम किया था, और हमने निम्न-आयामी पोज़िशनल रिप्रेज़ेंटेशन वेक्टर को स्पष्ट रूप से स्पार्स वन-हॉट रिप्रेज़ेंटेशन में परिवर्तित किया था। यह वन-हॉट रिप्रेज़ेंटेशन मेमोरी-कुशल नहीं है। इसके अलावा, प्रत्येक शब्द को एक-दूसरे से स्वतंत्र रूप से माना जाता है, इसलिए वन-हॉट एन्कोडेड वेक्टर शब्दों के बीच के अर्थपूर्ण समानताओं को व्यक्त नहीं करते हैं।\n",
|
||||
"\n",
|
||||
"इस यूनिट में, हम **News AG** डेटासेट का और अधिक अन्वेषण करेंगे। शुरू करने के लिए, आइए डेटा लोड करें और पिछले यूनिट से कुछ परिभाषाएँ प्राप्त करें।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 2,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"import tensorflow as tf\n",
|
||||
"from tensorflow import keras\n",
|
||||
"import tensorflow_datasets as tfds\n",
|
||||
"import numpy as np\n",
|
||||
"\n",
|
||||
"ds_train, ds_test = tfds.load('ag_news_subset').values()"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"### एम्बेडिंग क्या है?\n",
|
||||
"\n",
|
||||
"**एम्बेडिंग** का विचार यह है कि शब्दों को निम्न-आयामी घने वेक्टरों के रूप में प्रस्तुत किया जाए, जो शब्द के अर्थपूर्ण अर्थ को प्रतिबिंबित करते हैं। हम बाद में चर्चा करेंगे कि अर्थपूर्ण शब्द एम्बेडिंग कैसे बनाई जाए, लेकिन फिलहाल, एम्बेडिंग को शब्द वेक्टर की आयामीयता को कम करने के एक तरीके के रूप में सोचें। \n",
|
||||
"\n",
|
||||
"इस प्रकार, एक एम्बेडिंग लेयर एक शब्द को इनपुट के रूप में लेती है और निर्दिष्ट `embedding_size` का आउटपुट वेक्टर उत्पन्न करती है। एक तरह से, यह `Dense` लेयर के समान है, लेकिन यह एक-हॉट एन्कोडेड वेक्टर को इनपुट के रूप में लेने के बजाय, शब्द संख्या को इनपुट के रूप में ले सकती है।\n",
|
||||
"\n",
|
||||
"हमारे नेटवर्क में पहली लेयर के रूप में एम्बेडिंग लेयर का उपयोग करके, हम बैग-ऑफ-वर्ड्स से **एम्बेडिंग बैग** मॉडल में स्विच कर सकते हैं, जहां हम पहले अपने टेक्स्ट के प्रत्येक शब्द को संबंधित एम्बेडिंग में परिवर्तित करते हैं, और फिर उन सभी एम्बेडिंग पर कुछ समग्र फ़ंक्शन की गणना करते हैं, जैसे `sum`, `average` या `max`। \n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"हमारे क्लासिफायर न्यूरल नेटवर्क में निम्नलिखित लेयर शामिल हैं:\n",
|
||||
"\n",
|
||||
"* `TextVectorization` लेयर, जो एक स्ट्रिंग को इनपुट के रूप में लेती है और टोकन नंबरों का एक टेन्सर उत्पन्न करती है। हम एक उचित शब्दावली आकार `vocab_size` निर्दिष्ट करेंगे और कम बार उपयोग किए जाने वाले शब्दों को अनदेखा करेंगे। इनपुट आकार 1 होगा, और आउटपुट आकार $n$ होगा, क्योंकि हमें $n$ टोकन प्राप्त होंगे, जिनमें से प्रत्येक में 0 से `vocab_size` तक की संख्या होगी।\n",
|
||||
"* `Embedding` लेयर, जो $n$ नंबर लेती है और प्रत्येक नंबर को एक निर्दिष्ट लंबाई (हमारे उदाहरण में 100) के घने वेक्टर में बदल देती है। इस प्रकार, $n$ आकार के इनपुट टेन्सर को $n\\times 100$ आकार के टेन्सर में परिवर्तित किया जाएगा। \n",
|
||||
"* एग्रीगेशन लेयर, जो इस टेन्सर का पहले अक्ष के साथ औसत लेती है, यानी यह विभिन्न शब्दों से संबंधित सभी $n$ इनपुट टेन्सर का औसत गणना करेगी। इस लेयर को लागू करने के लिए, हम एक `Lambda` लेयर का उपयोग करेंगे और उसमें औसत गणना करने के लिए फ़ंक्शन पास करेंगे। आउटपुट का आकार 100 होगा, और यह पूरे इनपुट अनुक्रम का संख्यात्मक प्रतिनिधित्व होगा।\n",
|
||||
"* अंतिम `Dense` रैखिक क्लासिफायर।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 3,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Model: \"sequential\"\n",
|
||||
"_________________________________________________________________\n",
|
||||
" Layer (type) Output Shape Param # \n",
|
||||
"=================================================================\n",
|
||||
" text_vectorization (TextVec (None, None) 0 \n",
|
||||
" torization) \n",
|
||||
" \n",
|
||||
" embedding (Embedding) (None, None, 100) 3000000 \n",
|
||||
" \n",
|
||||
" lambda (Lambda) (None, 100) 0 \n",
|
||||
" \n",
|
||||
" dense (Dense) (None, 4) 404 \n",
|
||||
" \n",
|
||||
"=================================================================\n",
|
||||
"Total params: 3,000,404\n",
|
||||
"Trainable params: 3,000,404\n",
|
||||
"Non-trainable params: 0\n",
|
||||
"_________________________________________________________________\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"vocab_size = 30000\n",
|
||||
"batch_size = 128\n",
|
||||
"\n",
|
||||
"vectorizer = keras.layers.experimental.preprocessing.TextVectorization(max_tokens=vocab_size,input_shape=(1,))\n",
|
||||
"\n",
|
||||
"model = keras.models.Sequential([\n",
|
||||
" vectorizer, \n",
|
||||
" keras.layers.Embedding(vocab_size,100),\n",
|
||||
" keras.layers.Lambda(lambda x: tf.reduce_mean(x,axis=1)),\n",
|
||||
" keras.layers.Dense(4, activation='softmax')\n",
|
||||
"])\n",
|
||||
"model.summary()"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"`summary` प्रिंटआउट में, **output shape** कॉलम में पहला टेंसर डायमेंशन `None` मिनीबैच साइज को दर्शाता है, और दूसरा टोकन अनुक्रम की लंबाई को। मिनीबैच में सभी टोकन अनुक्रमों की लंबाई अलग-अलग होती है। हम अगले सेक्शन में इसे संभालने के तरीके पर चर्चा करेंगे।\n",
|
||||
"\n",
|
||||
"अब चलिए नेटवर्क को ट्रेन करते हैं:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Training vectorizer\n",
|
||||
"938/938 [==============================] - 20s 20ms/step - loss: 0.7891 - acc: 0.8155 - val_loss: 0.4470 - val_acc: 0.8642\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"<keras.callbacks.History at 0x22255515100>"
|
||||
]
|
||||
},
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"def extract_text(x):\n",
|
||||
" return x['title']+' '+x['description']\n",
|
||||
"\n",
|
||||
"def tupelize(x):\n",
|
||||
" return (extract_text(x),x['label'])\n",
|
||||
"\n",
|
||||
"print(\"Training vectorizer\")\n",
|
||||
"vectorizer.adapt(ds_train.take(500).map(extract_text))\n",
|
||||
"\n",
|
||||
"model.compile(loss='sparse_categorical_crossentropy',metrics=['acc'])\n",
|
||||
"model.fit(ds_train.map(tupelize).batch(batch_size),validation_data=ds_test.map(tupelize).batch(batch_size))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {
|
||||
"nteract": {
|
||||
"transient": {
|
||||
"deleting": false
|
||||
}
|
||||
}
|
||||
},
|
||||
"source": [
|
||||
"> **नोट** कि हम डेटा के एक उपसमुच्चय के आधार पर वेक्टराइज़र बना रहे हैं। यह प्रक्रिया को तेज करने के लिए किया जाता है, और इससे ऐसी स्थिति उत्पन्न हो सकती है जब हमारे पाठ के सभी टोकन शब्दावली में मौजूद न हों। इस स्थिति में, उन टोकनों को अनदेखा कर दिया जाएगा, जिससे सटीकता में थोड़ी कमी हो सकती है। हालांकि, वास्तविक जीवन में पाठ का एक उपसमुच्चय अक्सर शब्दावली का अच्छा अनुमान प्रदान करता है।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"### परिवर्तनीय अनुक्रम आकारों से निपटना\n",
|
||||
"\n",
|
||||
"आइए समझते हैं कि मिनीबैच में प्रशिक्षण कैसे होता है। ऊपर दिए गए उदाहरण में, इनपुट टेन्सर का आयाम 1 है, और हम 128-लंबे मिनीबैच का उपयोग करते हैं, जिससे टेन्सर का वास्तविक आकार $128 \\times 1$ हो जाता है। हालांकि, प्रत्येक वाक्य में टोकन की संख्या अलग-अलग होती है। यदि हम `TextVectorization` लेयर को एकल इनपुट पर लागू करते हैं, तो लौटाए गए टोकन की संख्या अलग-अलग होती है, यह इस बात पर निर्भर करता है कि टेक्स्ट को कैसे टोकनाइज़ किया गया है:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 5,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"tf.Tensor([ 1 45], shape=(2,), dtype=int64)\n",
|
||||
"tf.Tensor([ 112 1271 1 3 1747 158], shape=(6,), dtype=int64)\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"print(vectorizer('Hello, world!'))\n",
|
||||
"print(vectorizer('I am glad to meet you!'))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"हालांकि, जब हम वेक्टराइज़र को कई अनुक्रमों पर लागू करते हैं, तो इसे आयताकार आकार का एक टेंसर उत्पन्न करना होता है, इसलिए यह अप्रयुक्त तत्वों को PAD टोकन (जो हमारे मामले में शून्य है) से भरता है:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 6,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"<tf.Tensor: shape=(2, 6), dtype=int64, numpy=\n",
|
||||
"array([[ 1, 45, 0, 0, 0, 0],\n",
|
||||
" [ 112, 1271, 1, 3, 1747, 158]], dtype=int64)>"
|
||||
]
|
||||
},
|
||||
"execution_count": 6,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"vectorizer(['Hello, world!','I am glad to meet you!'])"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"यहां हम एम्बेडिंग्स देख सकते हैं:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 7,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"array([[[ 1.53059261e-02, 6.80514947e-02, 3.14026810e-02, ...,\n",
|
||||
" -8.92002955e-02, 1.52911525e-04, -5.65562584e-02],\n",
|
||||
" [ 2.57456154e-01, 2.79364467e-01, -2.03605562e-01, ...,\n",
|
||||
" -2.07474351e-01, 8.31158683e-02, -2.03911960e-01],\n",
|
||||
" [ 3.98201384e-02, -8.03454965e-03, 2.39790026e-02, ...,\n",
|
||||
" -7.18549127e-04, 2.66963355e-02, -4.30646613e-02],\n",
|
||||
" [ 3.98201384e-02, -8.03454965e-03, 2.39790026e-02, ...,\n",
|
||||
" -7.18549127e-04, 2.66963355e-02, -4.30646613e-02],\n",
|
||||
" [ 3.98201384e-02, -8.03454965e-03, 2.39790026e-02, ...,\n",
|
||||
" -7.18549127e-04, 2.66963355e-02, -4.30646613e-02],\n",
|
||||
" [ 3.98201384e-02, -8.03454965e-03, 2.39790026e-02, ...,\n",
|
||||
" -7.18549127e-04, 2.66963355e-02, -4.30646613e-02]],\n",
|
||||
"\n",
|
||||
" [[ 1.89674050e-01, 2.61548996e-01, -3.67433839e-02, ...,\n",
|
||||
" -2.07366899e-01, -1.05442435e-01, -2.36952081e-01],\n",
|
||||
" [ 6.16133213e-02, 1.80511594e-01, 9.77298319e-02, ...,\n",
|
||||
" -5.46628237e-02, -1.07340455e-01, -1.06589928e-01],\n",
|
||||
" [ 1.53059261e-02, 6.80514947e-02, 3.14026810e-02, ...,\n",
|
||||
" -8.92002955e-02, 1.52911525e-04, -5.65562584e-02],\n",
|
||||
" [-4.84890305e-02, -8.41715634e-02, 1.51529670e-01, ...,\n",
|
||||
" 1.28192469e-01, -7.77286515e-02, 1.26041949e-01],\n",
|
||||
" [-4.17212099e-02, -5.60694858e-02, 4.08860669e-02, ...,\n",
|
||||
" 8.70475471e-02, 8.92383084e-02, 1.67974353e-01],\n",
|
||||
" [ 2.85779923e-01, 4.57767487e-01, 4.52292450e-02, ...,\n",
|
||||
" -1.97419018e-01, -2.04659685e-01, -2.79758364e-01]]],\n",
|
||||
" dtype=float32)"
|
||||
]
|
||||
},
|
||||
"execution_count": 7,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"model.layers[1](vectorizer(['Hello, world!','I am glad to meet you!'])).numpy()"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"> **नोट**: पैडिंग की मात्रा को कम करने के लिए, कुछ मामलों में यह समझदारी होती है कि डेटासेट में सभी अनुक्रमों को उनकी लंबाई बढ़ने के क्रम में (या अधिक सटीक रूप से, टोकन की संख्या के अनुसार) क्रमबद्ध किया जाए। इससे यह सुनिश्चित होगा कि प्रत्येक मिनीबैच में समान लंबाई के अनुक्रम शामिल हों।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## सेमांटिक एम्बेडिंग्स: वर्ड2वेक\n",
|
||||
"\n",
|
||||
"हमारे पिछले उदाहरण में, एम्बेडिंग लेयर ने शब्दों को वेक्टर प्रतिनिधित्व में मैप करना सीखा, लेकिन इन प्रतिनिधित्वों में सेमांटिक अर्थ नहीं था। यह अच्छा होगा कि हम एक ऐसा वेक्टर प्रतिनिधित्व सीखें जिसमें समान शब्द या पर्यायवाची शब्द कुछ वेक्टर दूरी (जैसे यूक्लिडियन दूरी) के संदर्भ में एक-दूसरे के करीब हों।\n",
|
||||
"\n",
|
||||
"इसके लिए, हमें अपने एम्बेडिंग मॉडल को [Word2Vec](https://en.wikipedia.org/wiki/Word2vec) जैसी तकनीक का उपयोग करके बड़े टेक्स्ट संग्रह पर प्रीट्रेन करना होगा। यह दो मुख्य आर्किटेक्चर पर आधारित है जो शब्दों का वितरित प्रतिनिधित्व उत्पन्न करने के लिए उपयोग किए जाते हैं:\n",
|
||||
"\n",
|
||||
" - **कंटीन्युअस बैग-ऑफ-वर्ड्स** (CBoW), जिसमें हम मॉडल को आस-पास के संदर्भ से एक शब्द की भविष्यवाणी करने के लिए प्रशिक्षित करते हैं। दिए गए ngram $(W_{-2},W_{-1},W_0,W_1,W_2)$ में, मॉडल का लक्ष्य $(W_{-2},W_{-1},W_1,W_2)$ से $W_0$ की भविष्यवाणी करना है।\n",
|
||||
" - **कंटीन्युअस स्किप-ग्राम** CBoW के विपरीत है। यह मॉडल संदर्भ शब्दों की आस-पास की विंडो का उपयोग करके वर्तमान शब्द की भविष्यवाणी करता है।\n",
|
||||
"\n",
|
||||
"CBoW तेज है, जबकि स्किप-ग्राम धीमा है, लेकिन यह कम बार उपयोग होने वाले शब्दों का बेहतर प्रतिनिधित्व करता है।\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"Google News डेटासेट पर प्रीट्रेन किए गए Word2Vec एम्बेडिंग के साथ प्रयोग करने के लिए, हम **gensim** लाइब्रेरी का उपयोग कर सकते हैं। नीचे हम 'neural' के सबसे समान शब्दों को ढूंढते हैं।\n",
|
||||
"\n",
|
||||
"> **Note:** जब आप पहली बार शब्द वेक्टर बनाते हैं, तो उन्हें डाउनलोड करने में कुछ समय लग सकता है!\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 8,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"import gensim.downloader as api\n",
|
||||
"w2v = api.load('word2vec-google-news-300')"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 12,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"neuronal -> 0.7804799675941467\n",
|
||||
"neurons -> 0.7326500415802002\n",
|
||||
"neural_circuits -> 0.7252851724624634\n",
|
||||
"neuron -> 0.7174385190010071\n",
|
||||
"cortical -> 0.6941086649894714\n",
|
||||
"brain_circuitry -> 0.6923246383666992\n",
|
||||
"synaptic -> 0.6699118614196777\n",
|
||||
"neural_circuitry -> 0.6638563275337219\n",
|
||||
"neurochemical -> 0.6555314064025879\n",
|
||||
"neuronal_activity -> 0.6531826257705688\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"for w,p in w2v.most_similar('neural'):\n",
|
||||
" print(f\"{w} -> {p}\")"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"हम शब्द से वेक्टर एम्बेडिंग भी निकाल सकते हैं, जिसे वर्गीकरण मॉडल के प्रशिक्षण में उपयोग किया जा सकता है। एम्बेडिंग में 300 घटक होते हैं, लेकिन यहां स्पष्टता के लिए हम केवल वेक्टर के पहले 20 घटक दिखा रहे हैं:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 13,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"array([ 0.01226807, 0.06225586, 0.10693359, 0.05810547, 0.23828125,\n",
|
||||
" 0.03686523, 0.05151367, -0.20703125, 0.01989746, 0.10058594,\n",
|
||||
" -0.03759766, -0.1015625 , -0.15820312, -0.08105469, -0.0390625 ,\n",
|
||||
" -0.05053711, 0.16015625, 0.2578125 , 0.10058594, -0.25976562],\n",
|
||||
" dtype=float32)"
|
||||
]
|
||||
},
|
||||
"execution_count": 13,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"w2v['play'][:20]"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"सार्थक एम्बेडिंग की महान बात यह है कि आप अर्थ के आधार पर वेक्टर एन्कोडिंग को संशोधित कर सकते हैं। उदाहरण के लिए, हम ऐसा शब्द खोजने के लिए कह सकते हैं जिसका वेक्टर प्रतिनिधित्व *राजा* और *महिला* शब्दों के जितना करीब हो सके, और *पुरुष* शब्द से जितना दूर हो सके:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 14,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"('queen', 0.7118192911148071)"
|
||||
]
|
||||
},
|
||||
"execution_count": 14,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"w2v.most_similar(positive=['king','woman'],negative=['man'])[0]"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {
|
||||
"tags": []
|
||||
},
|
||||
"source": [
|
||||
"ऊपर दिए गए उदाहरण में कुछ आंतरिक GenSym जादू का उपयोग किया गया है, लेकिन मूल तर्क वास्तव में काफी सरल है। एम्बेडिंग्स के बारे में एक दिलचस्प बात यह है कि आप एम्बेडिंग वेक्टर पर सामान्य वेक्टर संचालन कर सकते हैं, और वह शब्दों के **अर्थों** पर संचालन को प्रतिबिंबित करेगा। ऊपर दिए गए उदाहरण को वेक्टर संचालन के रूप में व्यक्त किया जा सकता है: हम **KING-MAN+WOMAN** के अनुरूप वेक्टर की गणना करते हैं (संबंधित शब्दों के वेक्टर प्रतिनिधित्व पर `+` और `-` संचालन किए जाते हैं), और फिर उस वेक्टर के सबसे निकटतम शब्द को शब्दकोश में खोजते हैं:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 15,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"'queen'"
|
||||
]
|
||||
},
|
||||
"execution_count": 15,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"# get the vector corresponding to kind-man+woman\n",
|
||||
"qvec = w2v['king']-1.7*w2v['man']+1.7*w2v['woman']\n",
|
||||
"# find the index of the closest embedding vector \n",
|
||||
"d = np.sum((w2v.vectors-qvec)**2,axis=1)\n",
|
||||
"min_idx = np.argmin(d)\n",
|
||||
"# find the corresponding word\n",
|
||||
"w2v.index_to_key[min_idx]"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"> **NOTE**: हमने *man* और *woman* वेक्टर में एक छोटा गुणांक जोड़ना पड़ा - इसे हटाकर देखें कि क्या होता है।\n",
|
||||
"\n",
|
||||
"सबसे नज़दीकी वेक्टर खोजने के लिए, हम TensorFlow की तकनीक का उपयोग करते हैं ताकि हमारे वेक्टर और शब्दावली में सभी वेक्टर के बीच की दूरी का वेक्टर प्राप्त किया जा सके, और फिर `argmin` का उपयोग करके न्यूनतम शब्द का इंडेक्स खोजा जा सके।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"हालांकि Word2Vec शब्दार्थ को व्यक्त करने का एक शानदार तरीका लगता है, इसके कई नुकसान भी हैं, जिनमें निम्नलिखित शामिल हैं:\n",
|
||||
"\n",
|
||||
"* CBoW और skip-gram मॉडल दोनों **पूर्वानुमानात्मक एम्बेडिंग्स** हैं, और ये केवल स्थानीय संदर्भ को ध्यान में रखते हैं। Word2Vec वैश्विक संदर्भ का लाभ नहीं उठाता।\n",
|
||||
"* Word2Vec शब्द की **रूप-रचना** (morphology) को ध्यान में नहीं रखता, यानी इस तथ्य को कि शब्द का अर्थ उसके विभिन्न भागों, जैसे मूल (root), पर निर्भर कर सकता है। \n",
|
||||
"\n",
|
||||
"**FastText** दूसरे प्रतिबंध को दूर करने की कोशिश करता है और Word2Vec पर आधारित होकर प्रत्येक शब्द और उसमें पाए जाने वाले अक्षर n-grams के लिए वेक्टर प्रतिनिधित्व सीखता है। इन प्रतिनिधित्वों के मानों को प्रत्येक प्रशिक्षण चरण में एक वेक्टर में औसतित किया जाता है। हालांकि यह पूर्व-प्रशिक्षण में अतिरिक्त गणना जोड़ता है, यह शब्द एम्बेडिंग्स को उप-शब्द जानकारी को एन्कोड करने में सक्षम बनाता है।\n",
|
||||
"\n",
|
||||
"एक अन्य विधि, **GloVe**, शब्द एम्बेडिंग्स के लिए एक अलग दृष्टिकोण अपनाती है, जो शब्द-संदर्भ मैट्रिक्स के गुणनखंडन (factorization) पर आधारित है। सबसे पहले, यह एक बड़ा मैट्रिक्स बनाता है जो विभिन्न संदर्भों में शब्दों की घटनाओं की संख्या को गिनता है, और फिर यह इस मैट्रिक्स को निम्न आयामों में इस तरह से प्रस्तुत करने की कोशिश करता है कि पुनर्निर्माण हानि (reconstruction loss) न्यूनतम हो।\n",
|
||||
"\n",
|
||||
"gensim लाइब्रेरी इन शब्द एम्बेडिंग्स का समर्थन करती है, और आप ऊपर दिए गए मॉडल लोडिंग कोड को बदलकर इनके साथ प्रयोग कर सकते हैं।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Keras में प्रीट्रेंड एम्बेडिंग्स का उपयोग करना\n",
|
||||
"\n",
|
||||
"हम ऊपर दिए गए उदाहरण को संशोधित कर सकते हैं ताकि हमारे एम्बेडिंग लेयर की मैट्रिक्स को वर्ड2वेक जैसे सेमांटिक एम्बेडिंग्स से पहले से भर सकें। प्रीट्रेंड एम्बेडिंग और टेक्स्ट कॉर्पस की शब्दावली संभवतः मेल नहीं खाएगी, इसलिए हमें एक को चुनना होगा। यहां हम दो संभावित विकल्पों का पता लगाते हैं: टोकनाइज़र शब्दावली का उपयोग करना, और वर्ड2वेक एम्बेडिंग्स की शब्दावली का उपयोग करना।\n",
|
||||
"\n",
|
||||
"### टोकनाइज़र शब्दावली का उपयोग करना\n",
|
||||
"\n",
|
||||
"जब टोकनाइज़र शब्दावली का उपयोग करते हैं, तो शब्दावली के कुछ शब्दों के लिए वर्ड2वेक एम्बेडिंग्स उपलब्ध होंगे, और कुछ गायब होंगे। मान लें कि हमारी शब्दावली का आकार `vocab_size` है, और वर्ड2वेक एम्बेडिंग वेक्टर की लंबाई `embed_size` है, तो एम्बेडिंग लेयर को `vocab_size`$\\times$`embed_size` आकार की वेट मैट्रिक्स द्वारा दर्शाया जाएगा। हम इस मैट्रिक्स को शब्दावली के माध्यम से जाकर भरेंगे:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 9,
|
||||
"metadata": {
|
||||
"tags": []
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Embedding size: 300\n",
|
||||
"Populating matrix, this will take some time...Done, found 4551 words, 784 words missing\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"embed_size = len(w2v.get_vector('hello'))\n",
|
||||
"print(f'Embedding size: {embed_size}')\n",
|
||||
"\n",
|
||||
"vocab = vectorizer.get_vocabulary()\n",
|
||||
"W = np.zeros((vocab_size,embed_size))\n",
|
||||
"print('Populating matrix, this will take some time...',end='')\n",
|
||||
"found, not_found = 0,0\n",
|
||||
"for i,w in enumerate(vocab):\n",
|
||||
" try:\n",
|
||||
" W[i] = w2v.get_vector(w)\n",
|
||||
" found+=1\n",
|
||||
" except:\n",
|
||||
" # W[i] = np.random.normal(0.0,0.3,size=(embed_size,))\n",
|
||||
" not_found+=1\n",
|
||||
"\n",
|
||||
"print(f\"Done, found {found} words, {not_found} words missing\")"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"शब्दों के लिए जो Word2Vec शब्दावली में मौजूद नहीं हैं, हम उन्हें शून्य के रूप में छोड़ सकते हैं, या एक रैंडम वेक्टर उत्पन्न कर सकते हैं।\n",
|
||||
"\n",
|
||||
"अब हम प्रीट्रेंड वेट्स के साथ एक एम्बेडिंग लेयर को परिभाषित कर सकते हैं:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 10,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"emb = keras.layers.Embedding(vocab_size,embed_size,weights=[W],trainable=False)\n",
|
||||
"model = keras.models.Sequential([\n",
|
||||
" vectorizer, emb,\n",
|
||||
" keras.layers.Lambda(lambda x: tf.reduce_mean(x,axis=1)),\n",
|
||||
" keras.layers.Dense(4, activation='softmax')\n",
|
||||
"])"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": []
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 11,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"938/938 [==============================] - 10s 10ms/step - loss: 1.1075 - acc: 0.7822 - val_loss: 0.9134 - val_acc: 0.8175\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"<keras.callbacks.History at 0x2220226ef10>"
|
||||
]
|
||||
},
|
||||
"execution_count": 11,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"model.compile(loss='sparse_categorical_crossentropy',metrics=['acc'])\n",
|
||||
"model.fit(ds_train.map(tupelize).batch(batch_size),\n",
|
||||
" validation_data=ds_test.map(tupelize).batch(batch_size))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"> **नोट**: ध्यान दें कि जब हम `Embedding` बनाते समय `trainable=False` सेट करते हैं, तो इसका मतलब है कि हम Embedding लेयर को पुनः प्रशिक्षित नहीं कर रहे हैं। इससे सटीकता थोड़ी कम हो सकती है, लेकिन यह प्रशिक्षण को तेज कर देता है।\n",
|
||||
"\n",
|
||||
"### एम्बेडिंग शब्दावली का उपयोग करना\n",
|
||||
"\n",
|
||||
"पिछले दृष्टिकोण के साथ एक समस्या यह है कि TextVectorization और Embedding में उपयोग की गई शब्दावलियां अलग-अलग हैं। इस समस्या को हल करने के लिए, हम निम्नलिखित समाधानों में से एक का उपयोग कर सकते हैं:\n",
|
||||
"* हमारे शब्दावली पर Word2Vec मॉडल को पुनः प्रशिक्षित करें।\n",
|
||||
"* प्रीट्रेंड Word2Vec मॉडल की शब्दावली के साथ हमारा डेटासेट लोड करें। डेटासेट को लोड करते समय उपयोग की जाने वाली शब्दावलियां निर्दिष्ट की जा सकती हैं।\n",
|
||||
"\n",
|
||||
"दूसरा दृष्टिकोण आसान लगता है, तो चलिए इसे लागू करते हैं। सबसे पहले, हम Word2Vec एम्बेडिंग से ली गई निर्दिष्ट शब्दावली के साथ एक `TextVectorization` लेयर बनाएंगे:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 12,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"vocab = list(w2v.vocab.keys())\n",
|
||||
"vectorizer = keras.layers.experimental.preprocessing.TextVectorization(input_shape=(1,))\n",
|
||||
"vectorizer.set_vocabulary(vocab)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"जेनसिम वर्ड एम्बेडिंग्स लाइब्रेरी में एक सुविधाजनक फ़ंक्शन, `get_keras_embeddings`, होता है, जो आपके लिए स्वचालित रूप से संबंधित Keras एम्बेडिंग्स लेयर बना देगा।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 13,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Epoch 1/5\n",
|
||||
"938/938 [==============================] - 20s 14ms/step - loss: 1.3377 - acc: 0.4978 - val_loss: 1.2995 - val_acc: 0.5647\n",
|
||||
"Epoch 2/5\n",
|
||||
"938/938 [==============================] - 10s 10ms/step - loss: 1.2587 - acc: 0.5722 - val_loss: 1.2339 - val_acc: 0.5842\n",
|
||||
"Epoch 3/5\n",
|
||||
"938/938 [==============================] - 10s 10ms/step - loss: 1.1980 - acc: 0.5884 - val_loss: 1.1826 - val_acc: 0.5954\n",
|
||||
"Epoch 4/5\n",
|
||||
"938/938 [==============================] - 12s 13ms/step - loss: 1.1503 - acc: 0.6002 - val_loss: 1.1417 - val_acc: 0.6018\n",
|
||||
"Epoch 5/5\n",
|
||||
"938/938 [==============================] - 11s 12ms/step - loss: 1.1120 - acc: 0.6097 - val_loss: 1.1083 - val_acc: 0.6104\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"<keras.callbacks.History at 0x2220ccb81c0>"
|
||||
]
|
||||
},
|
||||
"execution_count": 13,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"model = keras.models.Sequential([\n",
|
||||
" vectorizer, \n",
|
||||
" w2v.get_keras_embedding(train_embeddings=False),\n",
|
||||
" keras.layers.Lambda(lambda x: tf.reduce_mean(x,axis=1)),\n",
|
||||
" keras.layers.Dense(4, activation='softmax')\n",
|
||||
"])\n",
|
||||
"model.compile(loss='sparse_categorical_crossentropy',metrics=['acc'])\n",
|
||||
"model.fit(ds_train.map(tupelize).batch(128),validation_data=ds_test.map(tupelize).batch(128),epochs=5)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"हम जो उच्च सटीकता नहीं देख रहे हैं, उसके कारणों में से एक यह है कि हमारे डेटासेट के कुछ शब्द प्रीट्रेंड GloVe शब्दावली में नहीं हैं, और इसलिए उन्हें अनदेखा कर दिया जाता है। इसे दूर करने के लिए, हम अपने डेटासेट के आधार पर अपने स्वयं के एम्बेडिंग्स को प्रशिक्षित कर सकते हैं।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## संदर्भात्मक एम्बेडिंग\n",
|
||||
"\n",
|
||||
"पारंपरिक प्रीट्रेंड एम्बेडिंग जैसे Word2Vec की एक मुख्य सीमा यह है कि, भले ही वे किसी शब्द का कुछ अर्थ पकड़ सकते हैं, वे विभिन्न अर्थों के बीच अंतर नहीं कर सकते। यह डाउनस्ट्रीम मॉडल्स में समस्याएं पैदा कर सकता है।\n",
|
||||
"\n",
|
||||
"उदाहरण के लिए, शब्द 'play' का इन दो वाक्यों में अलग-अलग अर्थ है:\n",
|
||||
"- मैं थिएटर में एक **play** देखने गया।\n",
|
||||
"- जॉन अपने दोस्तों के साथ **play** करना चाहता है।\n",
|
||||
"\n",
|
||||
"हमने जिन प्रीट्रेंड एम्बेडिंग की बात की, वे शब्द 'play' के दोनों अर्थों को एक ही एम्बेडिंग में दर्शाते हैं। इस सीमा को दूर करने के लिए, हमें **भाषा मॉडल** पर आधारित एम्बेडिंग बनानी होगी, जो बड़े टेक्स्ट कॉर्पस पर प्रशिक्षित होता है और *जानता है* कि शब्दों को विभिन्न संदर्भों में कैसे जोड़ा जा सकता है। संदर्भात्मक एम्बेडिंग पर चर्चा करना इस ट्यूटोरियल के दायरे से बाहर है, लेकिन हम अगले यूनिट में भाषा मॉडल्स पर चर्चा करते समय इस पर वापस आएंगे।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"\n---\n\n**अस्वीकरण**: \nयह दस्तावेज़ AI अनुवाद सेवा [Co-op Translator](https://github.com/Azure/co-op-translator) का उपयोग करके अनुवादित किया गया है। जबकि हम सटीकता सुनिश्चित करने का प्रयास करते हैं, कृपया ध्यान दें कि स्वचालित अनुवाद में त्रुटियां या अशुद्धियां हो सकती हैं। मूल भाषा में उपलब्ध मूल दस्तावेज़ को प्रामाणिक स्रोत माना जाना चाहिए। महत्वपूर्ण जानकारी के लिए, पेशेवर मानव अनुवाद की सिफारिश की जाती है। इस अनुवाद के उपयोग से उत्पन्न किसी भी गलतफहमी या गलत व्याख्या के लिए हम उत्तरदायी नहीं हैं।\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
"interpreter": {
|
||||
"hash": "0cb620c6d4b9f7a635928804c26cf22403d89d98d79684e4529119355ee6d5a5"
|
||||
},
|
||||
"kernel_info": {
|
||||
"name": "conda-env-py37_tensorflow-py"
|
||||
},
|
||||
"kernelspec": {
|
||||
"display_name": "py37_tensorflow",
|
||||
"language": "python",
|
||||
"name": "python3"
|
||||
},
|
||||
"language_info": {
|
||||
"codemirror_mode": {
|
||||
"name": "ipython",
|
||||
"version": 3
|
||||
},
|
||||
"file_extension": ".py",
|
||||
"mimetype": "text/x-python",
|
||||
"name": "python",
|
||||
"nbconvert_exporter": "python",
|
||||
"pygments_lexer": "ipython3",
|
||||
"version": "3.8.12"
|
||||
},
|
||||
"nteract": {
|
||||
"version": "nteract-front-end@1.0.0"
|
||||
},
|
||||
"coopTranslator": {
|
||||
"original_hash": "b859482be7f61d1eadc2c6a2720a37e4",
|
||||
"translation_date": "2025-08-31T15:26:22+00:00",
|
||||
"source_file": "lessons/5-NLP/14-Embeddings/EmbeddingsTF.ipynb",
|
||||
"language_code": "hi"
|
||||
}
|
||||
},
|
||||
"nbformat": 4,
|
||||
"nbformat_minor": 4
|
||||
}
|
||||
|
|
@ -0,0 +1,574 @@
|
|||
{
|
||||
"cells": [
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {
|
||||
"id": "NXTSugt6ieXh"
|
||||
},
|
||||
"source": [
|
||||
"## CBoW मॉडल का प्रशिक्षण\n",
|
||||
"\n",
|
||||
"यह नोटबुक [AI for Beginners Curriculum](http://aka.ms/ai-beginners) का हिस्सा है।\n",
|
||||
"\n",
|
||||
"इस उदाहरण में, हम CBoW भाषा मॉडल को प्रशिक्षित करने पर ध्यान देंगे ताकि हम अपना Word2Vec एम्बेडिंग स्पेस प्राप्त कर सकें। हम AG News डेटासेट को टेक्स्ट के स्रोत के रूप में उपयोग करेंगे।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"source": [
|
||||
"import torch\n",
|
||||
"import torchtext\n",
|
||||
"import os\n",
|
||||
"import collections\n",
|
||||
"import builtins\n",
|
||||
"import random\n",
|
||||
"import numpy as np"
|
||||
],
|
||||
"metadata": {
|
||||
"id": "q-UiiJUKaxHj"
|
||||
},
|
||||
"execution_count": null,
|
||||
"outputs": []
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"source": [
|
||||
"device = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")"
|
||||
],
|
||||
"metadata": {
|
||||
"id": "TFbR8CZaTZ1q"
|
||||
},
|
||||
"execution_count": null,
|
||||
"outputs": []
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"source": [
|
||||
"पहले चलिए हमारा डेटासेट लोड करते हैं और टोकनाइज़र और शब्दावली को परिभाषित करते हैं। हम `vocab_size` को 5000 पर सेट करेंगे ताकि गणनाओं को थोड़ा सीमित किया जा सके।\n"
|
||||
],
|
||||
"metadata": {
|
||||
"id": "HIwC7lI5T-ov"
|
||||
}
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"source": [
|
||||
"def load_dataset(ngrams = 1, min_freq = 1, vocab_size = 5000 , lines_cnt = 500):\n",
|
||||
" tokenizer = torchtext.data.utils.get_tokenizer('basic_english')\n",
|
||||
" print(\"Loading dataset...\")\n",
|
||||
" test_dataset, train_dataset = torchtext.datasets.AG_NEWS(root='./data')\n",
|
||||
" train_dataset = list(train_dataset)\n",
|
||||
" test_dataset = list(test_dataset)\n",
|
||||
" classes = ['World', 'Sports', 'Business', 'Sci/Tech']\n",
|
||||
" print('Building vocab...')\n",
|
||||
" counter = collections.Counter()\n",
|
||||
" for i, (_, line) in enumerate(train_dataset):\n",
|
||||
" counter.update(torchtext.data.utils.ngrams_iterator(tokenizer(line),ngrams=ngrams))\n",
|
||||
" if i == lines_cnt:\n",
|
||||
" break\n",
|
||||
" vocab = torchtext.vocab.Vocab(collections.Counter(dict(counter.most_common(vocab_size))), min_freq=min_freq)\n",
|
||||
" return train_dataset, test_dataset, classes, vocab, tokenizer"
|
||||
],
|
||||
"metadata": {
|
||||
"id": "wdZuygtgiuLG"
|
||||
},
|
||||
"execution_count": null,
|
||||
"outputs": []
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"source": [
|
||||
"train_dataset, test_dataset, _, vocab, tokenizer = load_dataset()"
|
||||
],
|
||||
"metadata": {
|
||||
"colab": {
|
||||
"base_uri": "https://localhost:8080/"
|
||||
},
|
||||
"id": "4d1nU1gsivGu",
|
||||
"outputId": "949fe272-ae0e-49f5-c373-6703458b3a74"
|
||||
},
|
||||
"execution_count": null,
|
||||
"outputs": [
|
||||
{
|
||||
"output_type": "stream",
|
||||
"name": "stdout",
|
||||
"text": [
|
||||
"Loading dataset...\n",
|
||||
"Building vocab...\n"
|
||||
]
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"source": [
|
||||
"def encode(x, vocabulary, tokenizer = tokenizer):\n",
|
||||
" return [vocabulary[s] for s in tokenizer(x)]"
|
||||
],
|
||||
"metadata": {
|
||||
"id": "1XDYNhG8ToFV"
|
||||
},
|
||||
"execution_count": null,
|
||||
"outputs": []
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {
|
||||
"id": "LIlQk6_PaHVY"
|
||||
},
|
||||
"source": [
|
||||
"## CBoW मॉडल\n",
|
||||
"\n",
|
||||
"CBoW $2N$ पड़ोसी शब्दों के आधार पर एक शब्द की भविष्यवाणी करना सीखता है। उदाहरण के लिए, जब $N=1$ हो, तो हमें वाक्य *I like to train networks* से निम्नलिखित जोड़े मिलेंगे: (like,I), (I, like), (to, like), (like,to), (train,to), (to, train), (networks, train), (train,networks)। यहाँ, पहला शब्द पड़ोसी शब्द है जिसे इनपुट के रूप में उपयोग किया गया है, और दूसरा शब्द वह है जिसकी हम भविष्यवाणी कर रहे हैं।\n",
|
||||
"\n",
|
||||
"अगले शब्द की भविष्यवाणी करने के लिए एक नेटवर्क बनाने के लिए, हमें पड़ोसी शब्द को इनपुट के रूप में देना होगा और शब्द संख्या को आउटपुट के रूप में प्राप्त करना होगा। CBoW नेटवर्क की संरचना निम्नलिखित है:\n",
|
||||
"\n",
|
||||
"* इनपुट शब्द को एम्बेडिंग लेयर से गुजारा जाता है। यही एम्बेडिंग लेयर हमारा Word2Vec एम्बेडिंग होगा, इसलिए हम इसे अलग से `embedder` वेरिएबल के रूप में परिभाषित करेंगे। इस उदाहरण में हम एम्बेडिंग साइज = 30 का उपयोग करेंगे, हालांकि आप उच्च डाइमेंशन के साथ प्रयोग करना चाह सकते हैं (वास्तविक Word2Vec में 300 होता है)।\n",
|
||||
"* एम्बेडिंग वेक्टर को फिर एक लीनियर लेयर में पास किया जाएगा, जो आउटपुट शब्द की भविष्यवाणी करेगा। इसलिए इसमें `vocab_size` न्यूरॉन्स होंगे।\n",
|
||||
"\n",
|
||||
"आउटपुट के लिए, यदि हम `CrossEntropyLoss` को लॉस फंक्शन के रूप में उपयोग करते हैं, तो हमें अपेक्षित परिणाम के रूप में केवल शब्द संख्या प्रदान करनी होगी, बिना वन-हॉट एनकोडिंग के।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"source": [
|
||||
"vocab_size = len(vocab)\n",
|
||||
"\n",
|
||||
"embedder = torch.nn.Embedding(num_embeddings = vocab_size, embedding_dim = 30)\n",
|
||||
"model = torch.nn.Sequential(\n",
|
||||
" embedder,\n",
|
||||
" torch.nn.Linear(in_features = 30, out_features = vocab_size),\n",
|
||||
")\n",
|
||||
"\n",
|
||||
"print(model)"
|
||||
],
|
||||
"metadata": {
|
||||
"colab": {
|
||||
"base_uri": "https://localhost:8080/"
|
||||
},
|
||||
"id": "akKTcKQKkfl2",
|
||||
"outputId": "da687e3e-a8ec-4c1a-e456-ab8cd6ac7dad"
|
||||
},
|
||||
"execution_count": null,
|
||||
"outputs": [
|
||||
{
|
||||
"output_type": "stream",
|
||||
"name": "stdout",
|
||||
"text": [
|
||||
"Sequential(\n",
|
||||
" (0): Embedding(5002, 30)\n",
|
||||
" (1): Linear(in_features=30, out_features=5002, bias=True)\n",
|
||||
")\n"
|
||||
]
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {
|
||||
"id": "Nud6jgGPaHVa"
|
||||
},
|
||||
"source": [
|
||||
"## प्रशिक्षण डेटा तैयार करना\n",
|
||||
"\n",
|
||||
"अब आइए मुख्य फ़ंक्शन को प्रोग्राम करते हैं, जो टेक्स्ट से CBoW शब्द जोड़े तैयार करेगा। यह फ़ंक्शन हमें विंडो साइज निर्दिष्ट करने की अनुमति देगा और इनपुट और आउटपुट शब्दों के जोड़े का एक सेट लौटाएगा। ध्यान दें कि इस फ़ंक्शन का उपयोग शब्दों पर किया जा सकता है, साथ ही वेक्टर/टेंसर पर भी - जो हमें टेक्स्ट को एन्कोड करने की अनुमति देगा, इससे पहले कि इसे `to_cbow` फ़ंक्शन में पास किया जाए।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {
|
||||
"colab": {
|
||||
"base_uri": "https://localhost:8080/"
|
||||
},
|
||||
"id": "x-dsXygOieXn",
|
||||
"outputId": "c2218280-e540-40ba-9546-efe48d0d714f"
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"output_type": "stream",
|
||||
"name": "stdout",
|
||||
"text": [
|
||||
"[['like', 'I'], ['to', 'I'], ['I', 'like'], ['to', 'like'], ['train', 'like'], ['I', 'to'], ['like', 'to'], ['train', 'to'], ['networks', 'to'], ['like', 'train'], ['to', 'train'], ['networks', 'train'], ['to', 'networks'], ['train', 'networks']]\n",
|
||||
"[[232, 172], [5, 172], [172, 232], [5, 232], [0, 232], [172, 5], [232, 5], [0, 5], [1202, 5], [232, 0], [5, 0], [1202, 0], [5, 1202], [0, 1202]]\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"def to_cbow(sent,window_size=2):\n",
|
||||
" res = []\n",
|
||||
" for i,x in enumerate(sent):\n",
|
||||
" for j in range(max(0,i-window_size),min(i+window_size+1,len(sent))):\n",
|
||||
" if i!=j:\n",
|
||||
" res.append([sent[j],x])\n",
|
||||
" return res\n",
|
||||
"\n",
|
||||
"print(to_cbow(['I','like','to','train','networks']))\n",
|
||||
"print(to_cbow(encode('I like to train networks', vocab)))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {
|
||||
"id": "XVaaDLjaaHVb"
|
||||
},
|
||||
"source": [
|
||||
"आइए प्रशिक्षण डेटासेट तैयार करें। हम सभी समाचारों को देखेंगे, `to_cbow` को कॉल करेंगे ताकि शब्द जोड़ों की सूची प्राप्त हो सके, और उन जोड़ों को `X` और `Y` में जोड़ देंगे। समय बचाने के लिए, हम केवल पहले 10k समाचार आइटम पर विचार करेंगे - यदि आपके पास अधिक समय है और बेहतर एम्बेडिंग प्राप्त करना चाहते हैं, तो आप आसानी से इस सीमा को हटा सकते हैं :)\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {
|
||||
"id": "54b-Gd9TieXo"
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"X = []\n",
|
||||
"Y = []\n",
|
||||
"for i, x in zip(range(10000), train_dataset):\n",
|
||||
" for w1, w2 in to_cbow(encode(x[1], vocab), window_size = 5):\n",
|
||||
" X.append(w1)\n",
|
||||
" Y.append(w2)\n",
|
||||
"\n",
|
||||
"X = torch.tensor(X)\n",
|
||||
"Y = torch.tensor(Y)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"source": [
|
||||
"हम उस डेटा को एक डेटासेट में भी बदलेंगे, और डाटालोडर बनाएंगे:\n"
|
||||
],
|
||||
"metadata": {
|
||||
"id": "cwWy0PzXWhN5"
|
||||
}
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"source": [
|
||||
"class SimpleIterableDataset(torch.utils.data.IterableDataset):\n",
|
||||
" def __init__(self, X, Y):\n",
|
||||
" super(SimpleIterableDataset).__init__()\n",
|
||||
" self.data = []\n",
|
||||
" for i in range(len(X)):\n",
|
||||
" self.data.append( (Y[i], X[i]) )\n",
|
||||
" random.shuffle(self.data)\n",
|
||||
"\n",
|
||||
" def __iter__(self):\n",
|
||||
" return iter(self.data)"
|
||||
],
|
||||
"metadata": {
|
||||
"id": "mfoAcGPFZU8p"
|
||||
},
|
||||
"execution_count": null,
|
||||
"outputs": []
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {
|
||||
"id": "e4NQ_-5waHVc"
|
||||
},
|
||||
"source": [
|
||||
"हम उस डेटा को एक डेटासेट में बदलेंगे, और डाटालोडर बनाएंगे:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {
|
||||
"id": "AbLUcojlieXo"
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"ds = SimpleIterableDataset(X, Y)\n",
|
||||
"dl = torch.utils.data.DataLoader(ds, batch_size = 256)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {
|
||||
"id": "pKQr7sXeaHVc"
|
||||
},
|
||||
"source": [
|
||||
"अब आइए वास्तविक प्रशिक्षण करें। हम `SGD` ऑप्टिमाइज़र का उपयोग करेंगे जिसमें काफी उच्च लर्निंग रेट होगा। आप अन्य ऑप्टिमाइज़र्स, जैसे `Adam`, के साथ भी प्रयोग कर सकते हैं। हम शुरुआत में 10 epochs के लिए प्रशिक्षण करेंगे - और यदि आप और भी कम हानि चाहते हैं तो आप इस सेल को पुनः चला सकते हैं।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"source": [
|
||||
"def train_epoch(net, dataloader, lr = 0.01, optimizer = None, loss_fn = torch.nn.CrossEntropyLoss(), epochs = None, report_freq = 1):\n",
|
||||
" optimizer = optimizer or torch.optim.Adam(net.parameters(), lr = lr)\n",
|
||||
" loss_fn = loss_fn.to(device)\n",
|
||||
" net.train()\n",
|
||||
"\n",
|
||||
" for i in range(epochs):\n",
|
||||
" total_loss, j = 0, 0, \n",
|
||||
" for labels, features in dataloader:\n",
|
||||
" optimizer.zero_grad()\n",
|
||||
" features, labels = features.to(device), labels.to(device)\n",
|
||||
" out = net(features)\n",
|
||||
" loss = loss_fn(out, labels)\n",
|
||||
" loss.backward()\n",
|
||||
" optimizer.step()\n",
|
||||
" total_loss += loss\n",
|
||||
" j += 1\n",
|
||||
" if i % report_freq == 0:\n",
|
||||
" print(f\"Epoch: {i+1}: loss={total_loss.item()/j}\")\n",
|
||||
"\n",
|
||||
" return total_loss.item()/j"
|
||||
],
|
||||
"metadata": {
|
||||
"id": "HeeCYKr_KF1w"
|
||||
},
|
||||
"execution_count": null,
|
||||
"outputs": []
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"source": [
|
||||
"train_epoch(net = model, dataloader = dl, optimizer = torch.optim.SGD(model.parameters(), lr = 0.1), loss_fn = torch.nn.CrossEntropyLoss(), epochs = 10)"
|
||||
],
|
||||
"metadata": {
|
||||
"colab": {
|
||||
"base_uri": "https://localhost:8080/"
|
||||
},
|
||||
"id": "KVgwGtDHgDlT",
|
||||
"outputId": "2447833f-f0e3-4566-c33d-addbfe2f451d"
|
||||
},
|
||||
"execution_count": null,
|
||||
"outputs": [
|
||||
{
|
||||
"output_type": "stream",
|
||||
"name": "stdout",
|
||||
"text": [
|
||||
"Epoch: 1: loss=5.664632366860172\n",
|
||||
"Epoch: 2: loss=5.632101973960962\n",
|
||||
"Epoch: 3: loss=5.610399051405015\n",
|
||||
"Epoch: 4: loss=5.594621561080262\n",
|
||||
"Epoch: 5: loss=5.582538017415446\n",
|
||||
"Epoch: 6: loss=5.572900234519603\n",
|
||||
"Epoch: 7: loss=5.564951676341915\n",
|
||||
"Epoch: 8: loss=5.558288112064614\n",
|
||||
"Epoch: 9: loss=5.552576955031129\n",
|
||||
"Epoch: 10: loss=5.547634165194347\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"output_type": "execute_result",
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"5.547634165194347"
|
||||
]
|
||||
},
|
||||
"metadata": {},
|
||||
"execution_count": 16
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {
|
||||
"id": "W8u2qXZmaHVd"
|
||||
},
|
||||
"source": [
|
||||
"## Word2Vec आज़माना\n",
|
||||
"\n",
|
||||
"Word2Vec का उपयोग करने के लिए, चलिए हमारे शब्दकोश में मौजूद सभी शब्दों के लिए वेक्टर निकालते हैं:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {
|
||||
"id": "r8TatcXjkU_t"
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"vectors = torch.stack([embedder(torch.tensor(vocab[s])) for s in vocab.itos], 0)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {
|
||||
"id": "3OcX21UOaHVd"
|
||||
},
|
||||
"source": []
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {
|
||||
"colab": {
|
||||
"base_uri": "https://localhost:8080/"
|
||||
},
|
||||
"id": "bz6tAeLzieXp",
|
||||
"outputId": "5b20850e-4342-45e9-f840-cfac2b4d61d8"
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"output_type": "stream",
|
||||
"name": "stdout",
|
||||
"text": [
|
||||
"tensor([-0.0915, 2.1224, -0.0281, -0.6819, 1.1219, 0.6458, -1.3704, -1.3314,\n",
|
||||
" -1.1437, 0.4496, 0.2301, -0.3515, -0.8485, 1.0481, 0.4386, -0.8949,\n",
|
||||
" 0.5644, 1.0939, -2.5096, 3.2949, -0.2601, -0.8640, 0.1421, -0.0804,\n",
|
||||
" -0.5083, -1.0560, 0.9753, -0.5949, -1.6046, 0.5774],\n",
|
||||
" grad_fn=<EmbeddingBackward>)\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"paris_vec = embedder(torch.tensor(vocab['paris']))\n",
|
||||
"print(paris_vec)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {
|
||||
"id": "pHTJlaeYaHVd"
|
||||
},
|
||||
"source": [
|
||||
"यह वर्ड2वेक का उपयोग करके पर्यायवाची शब्दों को खोजने में दिलचस्प है। निम्नलिखित फ़ंक्शन दिए गए इनपुट के लिए `n` सबसे निकटतम शब्दों को लौटाएगा। उन्हें खोजने के लिए, हम $|w_i - v|$ का मान निकालते हैं, जहाँ $v$ हमारे इनपुट शब्द के लिए संबंधित वेक्टर है, और $w_i$ शब्दावली में $i$-वें शब्द का एनकोडिंग है। इसके बाद हम ऐरे को सॉर्ट करते हैं और `argsort` का उपयोग करके संबंधित इंडेक्स लौटाते हैं, और सूची के पहले `n` तत्व लेते हैं, जो शब्दावली में निकटतम शब्दों की स्थिति को एनकोड करते हैं।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {
|
||||
"colab": {
|
||||
"base_uri": "https://localhost:8080/"
|
||||
},
|
||||
"id": "NlZyi-_olFar",
|
||||
"outputId": "b5dbb163-88c4-4d5a-eaf2-6751f700e98c"
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"output_type": "execute_result",
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"['microsoft', 'quoted', 'lp', 'rate', 'top']"
|
||||
]
|
||||
},
|
||||
"metadata": {},
|
||||
"execution_count": 56
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"def close_words(x, n = 5):\n",
|
||||
" vec = embedder(torch.tensor(vocab[x]))\n",
|
||||
" top5 = np.linalg.norm(vectors.detach().numpy() - vec.detach().numpy(), axis = 1).argsort()[:n]\n",
|
||||
" return [ vocab.itos[x] for x in top5 ]\n",
|
||||
"\n",
|
||||
"close_words('microsoft')"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {
|
||||
"colab": {
|
||||
"base_uri": "https://localhost:8080/"
|
||||
},
|
||||
"id": "-dQq7xeAln0U",
|
||||
"outputId": "66f768c3-c248-4bfd-ce4f-c8ffc6d0dd0d"
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"output_type": "execute_result",
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"['basketball', 'lot', 'sinai', 'states', 'healthdaynews']"
|
||||
]
|
||||
},
|
||||
"metadata": {},
|
||||
"execution_count": 51
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"close_words('basketball')"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {
|
||||
"colab": {
|
||||
"base_uri": "https://localhost:8080/"
|
||||
},
|
||||
"id": "fJXqK26b29sa",
|
||||
"outputId": "78f0baba-ffd0-485a-dd87-0a12bedfd7fa"
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"output_type": "execute_result",
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"['funds', 'travel', 'sydney', 'japan', 'business']"
|
||||
]
|
||||
},
|
||||
"metadata": {},
|
||||
"execution_count": 77
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"close_words('funds')"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {
|
||||
"id": "My0VeTDd3Ji8"
|
||||
},
|
||||
"source": [
|
||||
"## मुख्य बातें\n",
|
||||
"\n",
|
||||
"स्मार्ट तकनीकों जैसे CBoW का उपयोग करके, हम Word2Vec मॉडल को प्रशिक्षित कर सकते हैं। आप skip-gram मॉडल को प्रशिक्षित करने की कोशिश भी कर सकते हैं, जिसे केंद्रीय शब्द के आधार पर पड़ोसी शब्द की भविष्यवाणी करने के लिए प्रशिक्षित किया जाता है, और देख सकते हैं कि यह कितना अच्छा प्रदर्शन करता है।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"\n---\n\n**अस्वीकरण**: \nयह दस्तावेज़ AI अनुवाद सेवा [Co-op Translator](https://github.com/Azure/co-op-translator) का उपयोग करके अनुवादित किया गया है। जबकि हम सटीकता सुनिश्चित करने का प्रयास करते हैं, कृपया ध्यान दें कि स्वचालित अनुवाद में त्रुटियां या अशुद्धियां हो सकती हैं। मूल भाषा में उपलब्ध मूल दस्तावेज़ को प्रामाणिक स्रोत माना जाना चाहिए। महत्वपूर्ण जानकारी के लिए, पेशेवर मानव अनुवाद की सिफारिश की जाती है। इस अनुवाद के उपयोग से उत्पन्न किसी भी गलतफहमी या गलत व्याख्या के लिए हम उत्तरदायी नहीं हैं।\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
"colab": {
|
||||
"collapsed_sections": [],
|
||||
"name": "CBoW-PyTorch.ipynb",
|
||||
"provenance": []
|
||||
},
|
||||
"interpreter": {
|
||||
"hash": "16af2a8bbb083ea23e5e41c7f5787656b2ce26968575d8763f2c4b17f9cd711f"
|
||||
},
|
||||
"kernelspec": {
|
||||
"display_name": "Python 3.8.12 ('py38')",
|
||||
"language": "python",
|
||||
"name": "python3"
|
||||
},
|
||||
"language_info": {
|
||||
"codemirror_mode": {
|
||||
"name": "ipython",
|
||||
"version": 3
|
||||
},
|
||||
"file_extension": ".py",
|
||||
"mimetype": "text/x-python",
|
||||
"name": "python",
|
||||
"nbconvert_exporter": "python",
|
||||
"pygments_lexer": "ipython3",
|
||||
"version": "3.8.12"
|
||||
},
|
||||
"orig_nbformat": 4,
|
||||
"gpuClass": "standard",
|
||||
"coopTranslator": {
|
||||
"original_hash": "36df28efe3fe40b6fb0a7fa48fe3ea82",
|
||||
"translation_date": "2025-08-31T15:14:46+00:00",
|
||||
"source_file": "lessons/5-NLP/15-LanguageModeling/CBoW-PyTorch.ipynb",
|
||||
"language_code": "hi"
|
||||
}
|
||||
},
|
||||
"nbformat": 4,
|
||||
"nbformat_minor": 0
|
||||
}
|
||||
File diff suppressed because it is too large
Load Diff
|
|
@ -0,0 +1,479 @@
|
|||
{
|
||||
"cells": [
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"# पुनरावर्ती न्यूरल नेटवर्क\n",
|
||||
"\n",
|
||||
"पिछले मॉड्यूल में, हमने टेक्स्ट के समृद्ध सेमांटिक प्रतिनिधित्व और एम्बेडिंग के ऊपर एक साधारण लीनियर क्लासिफायर का उपयोग किया। यह आर्किटेक्चर वाक्य में शब्दों के समग्र अर्थ को कैप्चर करता है, लेकिन यह शब्दों के **क्रम** को ध्यान में नहीं रखता, क्योंकि एम्बेडिंग पर किया गया समग्र ऑपरेशन मूल टेक्स्ट से इस जानकारी को हटा देता है। चूंकि ये मॉडल शब्दों के क्रम को मॉडल नहीं कर सकते, वे टेक्स्ट जनरेशन या प्रश्नोत्तर जैसे अधिक जटिल या अस्पष्ट कार्यों को हल नहीं कर सकते।\n",
|
||||
"\n",
|
||||
"टेक्स्ट अनुक्रम के अर्थ को कैप्चर करने के लिए, हमें एक अन्य न्यूरल नेटवर्क आर्किटेक्चर का उपयोग करना होगा, जिसे **पुनरावर्ती न्यूरल नेटवर्क** या RNN कहा जाता है। RNN में, हम अपने वाक्य को नेटवर्क के माध्यम से एक बार में एक प्रतीक पास करते हैं, और नेटवर्क कुछ **स्थिति** (state) उत्पन्न करता है, जिसे हम अगले प्रतीक के साथ नेटवर्क में फिर से पास करते हैं।\n",
|
||||
"\n",
|
||||
"दिए गए इनपुट अनुक्रम $X_0,\\dots,X_n$ के लिए, RNN एक न्यूरल नेटवर्क ब्लॉकों का अनुक्रम बनाता है और इस अनुक्रम को बैक प्रोपेगेशन का उपयोग करके एंड-टू-एंड प्रशिक्षित करता है। प्रत्येक नेटवर्क ब्लॉक $(X_i,S_i)$ की एक जोड़ी को इनपुट के रूप में लेता है और परिणामस्वरूप $S_{i+1}$ उत्पन्न करता है। अंतिम स्थिति $S_n$ या आउटपुट $X_n$ को परिणाम उत्पन्न करने के लिए एक लीनियर क्लासिफायर में भेजा जाता है। सभी नेटवर्क ब्लॉक समान वेट्स साझा करते हैं और एक बैक प्रोपेगेशन पास का उपयोग करके एंड-टू-एंड प्रशिक्षित होते हैं।\n",
|
||||
"\n",
|
||||
"चूंकि स्थिति वेक्टर $S_0,\\dots,S_n$ नेटवर्क के माध्यम से पास किए जाते हैं, यह शब्दों के बीच अनुक्रमिक निर्भरताओं को सीखने में सक्षम होता है। उदाहरण के लिए, जब अनुक्रम में कहीं *not* शब्द आता है, तो यह स्थिति वेक्टर के कुछ तत्वों को नकारने (negate) के लिए सीख सकता है, जिससे नकारात्मकता (negation) उत्पन्न होती है।\n",
|
||||
"\n",
|
||||
"> चूंकि चित्र में सभी RNN ब्लॉकों के वेट्स साझा किए जाते हैं, इसलिए उसी चित्र को एक ब्लॉक (दाईं ओर) के रूप में दर्शाया जा सकता है, जिसमें एक पुनरावर्ती फीडबैक लूप होता है, जो नेटवर्क की आउटपुट स्थिति को इनपुट में वापस पास करता है।\n",
|
||||
"\n",
|
||||
"आइए देखें कि पुनरावर्ती न्यूरल नेटवर्क हमारे समाचार डेटासेट को वर्गीकृत करने में कैसे मदद कर सकते हैं।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 1,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Loading dataset...\n",
|
||||
"Building vocab...\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"import torch\n",
|
||||
"import torchtext\n",
|
||||
"from torchnlp import *\n",
|
||||
"train_dataset, test_dataset, classes, vocab = load_dataset()\n",
|
||||
"vocab_size = len(vocab)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## सरल RNN वर्गीकरणकर्ता\n",
|
||||
"\n",
|
||||
"सरल RNN के मामले में, प्रत्येक पुनरावर्ती यूनिट एक साधारण रैखिक नेटवर्क होता है, जो संयोजित इनपुट वेक्टर और स्टेट वेक्टर लेता है, और एक नया स्टेट वेक्टर उत्पन्न करता है। PyTorch इस यूनिट को `RNNCell` क्लास के साथ दर्शाता है, और ऐसे सेल्स के नेटवर्क को `RNN` लेयर के रूप में।\n",
|
||||
"\n",
|
||||
"एक RNN वर्गीकरणकर्ता को परिभाषित करने के लिए, हम पहले एक एम्बेडिंग लेयर लागू करेंगे ताकि इनपुट शब्दावली की आयामीयता को कम किया जा सके, और फिर इसके ऊपर RNN लेयर रखेंगे:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 2,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"class RNNClassifier(torch.nn.Module):\n",
|
||||
" def __init__(self, vocab_size, embed_dim, hidden_dim, num_class):\n",
|
||||
" super().__init__()\n",
|
||||
" self.hidden_dim = hidden_dim\n",
|
||||
" self.embedding = torch.nn.Embedding(vocab_size, embed_dim)\n",
|
||||
" self.rnn = torch.nn.RNN(embed_dim,hidden_dim,batch_first=True)\n",
|
||||
" self.fc = torch.nn.Linear(hidden_dim, num_class)\n",
|
||||
"\n",
|
||||
" def forward(self, x):\n",
|
||||
" batch_size = x.size(0)\n",
|
||||
" x = self.embedding(x)\n",
|
||||
" x,h = self.rnn(x)\n",
|
||||
" return self.fc(x.mean(dim=1))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"> **नोट:** यहां हम सरलता के लिए बिना प्रशिक्षित embedding layer का उपयोग कर रहे हैं, लेकिन और बेहतर परिणामों के लिए हम Word2Vec या GloVe embeddings के साथ प्री-ट्रेंड embedding layer का उपयोग कर सकते हैं, जैसा कि पिछले यूनिट में बताया गया है। बेहतर समझ के लिए, आप इस कोड को प्री-ट्रेंड embeddings के साथ काम करने के लिए अनुकूलित कर सकते हैं।\n",
|
||||
"\n",
|
||||
"हमारे मामले में, हम padded data loader का उपयोग करेंगे, जिससे प्रत्येक बैच में समान लंबाई वाले padded sequences होंगे। RNN layer embedding tensors के sequence को लेगी और दो आउटपुट उत्पन्न करेगी:\n",
|
||||
"* $x$ प्रत्येक चरण पर RNN सेल आउटपुट का sequence है\n",
|
||||
"* $h$ sequence के अंतिम तत्व के लिए अंतिम hidden state है\n",
|
||||
"\n",
|
||||
"इसके बाद हम एक fully-connected linear classifier लागू करेंगे ताकि वर्गों (classes) की संख्या प्राप्त की जा सके।\n",
|
||||
"\n",
|
||||
"> **नोट:** RNNs को प्रशिक्षित करना काफी कठिन होता है, क्योंकि जब RNN सेल्स को sequence की लंबाई के साथ unroll किया जाता है, तो back propagation में शामिल लेयर्स की संख्या काफी बड़ी हो जाती है। इसलिए हमें छोटा learning rate चुनने की आवश्यकता होती है और अच्छे परिणाम प्राप्त करने के लिए नेटवर्क को बड़े dataset पर प्रशिक्षित करना पड़ता है। इसमें काफी समय लग सकता है, इसलिए GPU का उपयोग करना बेहतर होता है।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 3,
|
||||
"metadata": {
|
||||
"scrolled": true
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"3200: acc=0.3090625\n",
|
||||
"6400: acc=0.38921875\n",
|
||||
"9600: acc=0.4590625\n",
|
||||
"12800: acc=0.511953125\n",
|
||||
"16000: acc=0.5506875\n",
|
||||
"19200: acc=0.57921875\n",
|
||||
"22400: acc=0.6070089285714285\n",
|
||||
"25600: acc=0.6304296875\n",
|
||||
"28800: acc=0.6484027777777778\n",
|
||||
"32000: acc=0.66509375\n",
|
||||
"35200: acc=0.6790056818181818\n",
|
||||
"38400: acc=0.6929166666666666\n",
|
||||
"41600: acc=0.7035817307692308\n",
|
||||
"44800: acc=0.7137276785714286\n",
|
||||
"48000: acc=0.72225\n",
|
||||
"51200: acc=0.73001953125\n",
|
||||
"54400: acc=0.7372794117647059\n",
|
||||
"57600: acc=0.7436631944444444\n",
|
||||
"60800: acc=0.7503947368421052\n",
|
||||
"64000: acc=0.75634375\n",
|
||||
"67200: acc=0.7615773809523809\n",
|
||||
"70400: acc=0.7662642045454545\n",
|
||||
"73600: acc=0.7708423913043478\n",
|
||||
"76800: acc=0.7751822916666666\n",
|
||||
"80000: acc=0.7790625\n",
|
||||
"83200: acc=0.7825\n",
|
||||
"86400: acc=0.7858564814814815\n",
|
||||
"89600: acc=0.7890513392857142\n",
|
||||
"92800: acc=0.7920474137931034\n",
|
||||
"96000: acc=0.7952708333333334\n",
|
||||
"99200: acc=0.7982258064516129\n",
|
||||
"102400: acc=0.80099609375\n",
|
||||
"105600: acc=0.8037594696969697\n",
|
||||
"108800: acc=0.8060569852941176\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"train_loader = torch.utils.data.DataLoader(train_dataset, batch_size=16, collate_fn=padify, shuffle=True)\n",
|
||||
"net = RNNClassifier(vocab_size,64,32,len(classes)).to(device)\n",
|
||||
"train_epoch(net,train_loader, lr=0.001)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## लॉन्ग शॉर्ट टर्म मेमोरी (LSTM)\n",
|
||||
"\n",
|
||||
"क्लासिकल RNNs की मुख्य समस्याओं में से एक है **vanishing gradients** समस्या। चूंकि RNNs को एक ही बैक-प्रोपेगेशन पास में एंड-टू-एंड ट्रेन किया जाता है, यह नेटवर्क की शुरुआती लेयर्स तक एरर को प्रोपेगेट करने में कठिनाई महसूस करता है, और इस कारण नेटवर्क दूरस्थ टोकन के बीच संबंधों को सीख नहीं पाता। इस समस्या से बचने के तरीकों में से एक है **explicit state management** को लागू करना, जिसे **gates** के उपयोग से किया जाता है। इस प्रकार की दो सबसे प्रसिद्ध आर्किटेक्चर हैं: **लॉन्ग शॉर्ट टर्म मेमोरी** (LSTM) और **गेटेड रिले यूनिट** (GRU)।\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"LSTM नेटवर्क RNN के समान तरीके से संगठित होता है, लेकिन इसमें दो स्टेट्स होते हैं जो लेयर से लेयर तक पास किए जाते हैं: वास्तविक स्टेट $c$, और हिडन वेक्टर $h$। प्रत्येक यूनिट पर, हिडन वेक्टर $h_i$ को इनपुट $x_i$ के साथ जोड़ दिया जाता है, और वे **gates** के माध्यम से स्टेट $c$ पर क्या प्रभाव पड़ेगा, इसे नियंत्रित करते हैं। प्रत्येक गेट एक न्यूरल नेटवर्क होता है जिसमें सिग्मॉइड एक्टिवेशन होता है (आउटपुट $[0,1]$ की रेंज में), जिसे स्टेट वेक्टर के साथ गुणा करने पर बिटवाइज मास्क के रूप में सोचा जा सकता है। निम्नलिखित गेट्स होते हैं (ऊपर दी गई तस्वीर में बाएं से दाएं):\n",
|
||||
"* **forget gate** हिडन वेक्टर लेता है और तय करता है कि वेक्टर $c$ के कौन से घटकों को भूलना है और कौन से पास करना है।\n",
|
||||
"* **input gate** इनपुट और हिडन वेक्टर से कुछ जानकारी लेता है और इसे स्टेट में डालता है।\n",
|
||||
"* **output gate** स्टेट को $\\tanh$ एक्टिवेशन के साथ किसी लीनियर लेयर के माध्यम से ट्रांसफॉर्म करता है, फिर हिडन वेक्टर $h_i$ का उपयोग करके इसके कुछ घटकों को चुनता है ताकि नया स्टेट $c_{i+1}$ उत्पन्न हो सके।\n",
|
||||
"\n",
|
||||
"स्टेट $c$ के घटकों को कुछ फ्लैग्स के रूप में सोचा जा सकता है जिन्हें ऑन और ऑफ किया जा सकता है। उदाहरण के लिए, जब हम सीक्वेंस में *Alice* नाम देखते हैं, तो हम मान सकते हैं कि यह एक महिला पात्र को संदर्भित करता है, और स्टेट में फ्लैग उठाते हैं कि हमारे पास वाक्य में महिला संज्ञा है। जब हम आगे *and Tom* वाक्यांश देखते हैं, तो हम फ्लैग उठाते हैं कि हमारे पास बहुवचन संज्ञा है। इस प्रकार स्टेट को मैनिपुलेट करके हम वाक्य के भागों के व्याकरणिक गुणों को ट्रैक कर सकते हैं।\n",
|
||||
"\n",
|
||||
"> **Note**: LSTM की आंतरिक संरचना को समझने के लिए एक बेहतरीन संसाधन है क्रिस्टोफर ओलाह का यह शानदार लेख [Understanding LSTM Networks](https://colah.github.io/posts/2015-08-Understanding-LSTMs/)।\n",
|
||||
"\n",
|
||||
"हालांकि LSTM सेल की आंतरिक संरचना जटिल लग सकती है, PyTorch इस इम्प्लीमेंटेशन को `LSTMCell` क्लास के अंदर छुपा देता है, और पूरे LSTM लेयर को दर्शाने के लिए `LSTM` ऑब्जेक्ट प्रदान करता है। इस प्रकार, LSTM क्लासिफायर का इम्प्लीमेंटेशन ऊपर देखे गए सिंपल RNN के समान ही होगा:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"class LSTMClassifier(torch.nn.Module):\n",
|
||||
" def __init__(self, vocab_size, embed_dim, hidden_dim, num_class):\n",
|
||||
" super().__init__()\n",
|
||||
" self.hidden_dim = hidden_dim\n",
|
||||
" self.embedding = torch.nn.Embedding(vocab_size, embed_dim)\n",
|
||||
" self.embedding.weight.data = torch.randn_like(self.embedding.weight.data)-0.5\n",
|
||||
" self.rnn = torch.nn.LSTM(embed_dim,hidden_dim,batch_first=True)\n",
|
||||
" self.fc = torch.nn.Linear(hidden_dim, num_class)\n",
|
||||
"\n",
|
||||
" def forward(self, x):\n",
|
||||
" batch_size = x.size(0)\n",
|
||||
" x = self.embedding(x)\n",
|
||||
" x,(h,c) = self.rnn(x)\n",
|
||||
" return self.fc(h[-1])"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": []
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 5,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"3200: acc=0.259375\n",
|
||||
"6400: acc=0.25859375\n",
|
||||
"9600: acc=0.26177083333333334\n",
|
||||
"12800: acc=0.2784375\n",
|
||||
"16000: acc=0.313\n",
|
||||
"19200: acc=0.3528645833333333\n",
|
||||
"22400: acc=0.3965625\n",
|
||||
"25600: acc=0.4385546875\n",
|
||||
"28800: acc=0.4752777777777778\n",
|
||||
"32000: acc=0.505375\n",
|
||||
"35200: acc=0.5326704545454546\n",
|
||||
"38400: acc=0.5557552083333334\n",
|
||||
"41600: acc=0.5760817307692307\n",
|
||||
"44800: acc=0.5954910714285714\n",
|
||||
"48000: acc=0.6118333333333333\n",
|
||||
"51200: acc=0.62681640625\n",
|
||||
"54400: acc=0.6404779411764706\n",
|
||||
"57600: acc=0.6520138888888889\n",
|
||||
"60800: acc=0.662828947368421\n",
|
||||
"64000: acc=0.673546875\n",
|
||||
"67200: acc=0.6831547619047619\n",
|
||||
"70400: acc=0.6917897727272727\n",
|
||||
"73600: acc=0.6997146739130434\n",
|
||||
"76800: acc=0.707109375\n",
|
||||
"80000: acc=0.714075\n",
|
||||
"83200: acc=0.7209134615384616\n",
|
||||
"86400: acc=0.727037037037037\n",
|
||||
"89600: acc=0.7326674107142858\n",
|
||||
"92800: acc=0.7379633620689655\n",
|
||||
"96000: acc=0.7433645833333333\n",
|
||||
"99200: acc=0.7479032258064516\n",
|
||||
"102400: acc=0.752119140625\n",
|
||||
"105600: acc=0.7562405303030303\n",
|
||||
"108800: acc=0.76015625\n",
|
||||
"112000: acc=0.7641339285714286\n",
|
||||
"115200: acc=0.7677777777777778\n",
|
||||
"118400: acc=0.7711233108108108\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"(0.03487814127604167, 0.7728)"
|
||||
]
|
||||
},
|
||||
"execution_count": 5,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"net = LSTMClassifier(vocab_size,64,32,len(classes)).to(device)\n",
|
||||
"train_epoch(net,train_loader, lr=0.001)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## पैक्ड सीक्वेंसेज़\n",
|
||||
"\n",
|
||||
"हमारे उदाहरण में, हमें मिनीबैच की सभी सीक्वेंसेज़ को शून्य वेक्टर से पैड करना पड़ा। हालांकि इससे कुछ मेमोरी की बर्बादी होती है, लेकिन RNNs के साथ यह अधिक महत्वपूर्ण है कि पैड किए गए इनपुट आइटम्स के लिए अतिरिक्त RNN सेल्स बनाए जाते हैं, जो प्रशिक्षण में भाग लेते हैं, लेकिन कोई महत्वपूर्ण इनपुट जानकारी नहीं ले जाते। यह बेहतर होगा कि RNN को केवल वास्तविक सीक्वेंस साइज के अनुसार प्रशिक्षित किया जाए।\n",
|
||||
"\n",
|
||||
"इसे करने के लिए, PyTorch में पैडेड सीक्वेंस स्टोरेज का एक विशेष फॉर्मेट पेश किया गया है। मान लीजिए हमारे पास ऐसा इनपुट पैडेड मिनीबैच है:\n",
|
||||
"```\n",
|
||||
"[[1,2,3,4,5],\n",
|
||||
" [6,7,8,0,0],\n",
|
||||
" [9,0,0,0,0]]\n",
|
||||
"```\n",
|
||||
"यहां 0 पैडेड वैल्यूज़ को दर्शाता है, और इनपुट सीक्वेंसेज़ की वास्तविक लंबाई का वेक्टर `[5,3,1]` है।\n",
|
||||
"\n",
|
||||
"पैडेड सीक्वेंस के साथ RNN को प्रभावी ढंग से प्रशिक्षित करने के लिए, हम चाहते हैं कि RNN सेल्स के पहले ग्रुप का प्रशिक्षण बड़े मिनीबैच (`[1,6,9]`) के साथ शुरू हो, लेकिन फिर तीसरी सीक्वेंस की प्रोसेसिंग समाप्त हो जाए, और छोटे मिनीबैचेज़ (`[2,7]`, `[3,8]`) के साथ प्रशिक्षण जारी रहे, और इसी तरह। इस प्रकार, पैक्ड सीक्वेंस को एक वेक्टर के रूप में दर्शाया जाता है - हमारे मामले में `[1,6,9,2,7,3,8,4,5]`, और लंबाई का वेक्टर (`[5,3,1]`), जिससे हम मूल पैडेड मिनीबैच को आसानी से पुनर्निर्मित कर सकते हैं।\n",
|
||||
"\n",
|
||||
"पैक्ड सीक्वेंस बनाने के लिए, हम `torch.nn.utils.rnn.pack_padded_sequence` फंक्शन का उपयोग कर सकते हैं। सभी रिकरेंट लेयर्स, जैसे RNN, LSTM और GRU, इनपुट के रूप में पैक्ड सीक्वेंसेज़ को सपोर्ट करते हैं, और पैक्ड आउटपुट उत्पन्न करते हैं, जिसे `torch.nn.utils.rnn.pad_packed_sequence` का उपयोग करके डिकोड किया जा सकता है।\n",
|
||||
"\n",
|
||||
"पैक्ड सीक्वेंस बनाने में सक्षम होने के लिए, हमें नेटवर्क को लंबाई का वेक्टर पास करना होगा, और इस प्रकार हमें मिनीबैच तैयार करने के लिए एक अलग फंक्शन की आवश्यकता होगी:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 6,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"def pad_length(b):\n",
|
||||
" # build vectorized sequence\n",
|
||||
" v = [encode(x[1]) for x in b]\n",
|
||||
" # compute max length of a sequence in this minibatch and length sequence itself\n",
|
||||
" len_seq = list(map(len,v))\n",
|
||||
" l = max(len_seq)\n",
|
||||
" return ( # tuple of three tensors - labels, padded features, length sequence\n",
|
||||
" torch.LongTensor([t[0]-1 for t in b]),\n",
|
||||
" torch.stack([torch.nn.functional.pad(torch.tensor(t),(0,l-len(t)),mode='constant',value=0) for t in v]),\n",
|
||||
" torch.tensor(len_seq)\n",
|
||||
" )\n",
|
||||
"\n",
|
||||
"train_loader_len = torch.utils.data.DataLoader(train_dataset, batch_size=16, collate_fn=pad_length, shuffle=True)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"वास्तविक नेटवर्क `LSTMClassifier` के समान होगा, लेकिन `forward` पास में दोनों, पैडेड मिनीबैच और अनुक्रम की लंबाई का वेक्टर प्राप्त होगा। एम्बेडिंग की गणना करने के बाद, हम पैक्ड अनुक्रम की गणना करते हैं, इसे LSTM लेयर में पास करते हैं, और फिर परिणाम को वापस अनपैक करते हैं।\n",
|
||||
"\n",
|
||||
"> **Note**: हम वास्तव में अनपैक किए गए परिणाम `x` का उपयोग नहीं करते हैं, क्योंकि हम अगले गणनाओं में छिपी हुई लेयर से आउटपुट का उपयोग करते हैं। इसलिए, हम इस कोड से अनपैकिंग को पूरी तरह से हटा सकते हैं। इसे यहां रखने का कारण यह है कि आप इस कोड को आसानी से संशोधित कर सकें, यदि आपको नेटवर्क आउटपुट का उपयोग आगे की गणनाओं में करना पड़े।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 7,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"class LSTMPackClassifier(torch.nn.Module):\n",
|
||||
" def __init__(self, vocab_size, embed_dim, hidden_dim, num_class):\n",
|
||||
" super().__init__()\n",
|
||||
" self.hidden_dim = hidden_dim\n",
|
||||
" self.embedding = torch.nn.Embedding(vocab_size, embed_dim)\n",
|
||||
" self.embedding.weight.data = torch.randn_like(self.embedding.weight.data)-0.5\n",
|
||||
" self.rnn = torch.nn.LSTM(embed_dim,hidden_dim,batch_first=True)\n",
|
||||
" self.fc = torch.nn.Linear(hidden_dim, num_class)\n",
|
||||
"\n",
|
||||
" def forward(self, x, lengths):\n",
|
||||
" batch_size = x.size(0)\n",
|
||||
" x = self.embedding(x)\n",
|
||||
" pad_x = torch.nn.utils.rnn.pack_padded_sequence(x,lengths,batch_first=True,enforce_sorted=False)\n",
|
||||
" pad_x,(h,c) = self.rnn(pad_x)\n",
|
||||
" x, _ = torch.nn.utils.rnn.pad_packed_sequence(pad_x,batch_first=True)\n",
|
||||
" return self.fc(h[-1])"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": []
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 8,
|
||||
"metadata": {
|
||||
"scrolled": true
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"3200: acc=0.285625\n",
|
||||
"6400: acc=0.33359375\n",
|
||||
"9600: acc=0.3876041666666667\n",
|
||||
"12800: acc=0.44078125\n",
|
||||
"16000: acc=0.4825\n",
|
||||
"19200: acc=0.5235416666666667\n",
|
||||
"22400: acc=0.5559821428571429\n",
|
||||
"25600: acc=0.58609375\n",
|
||||
"28800: acc=0.6116666666666667\n",
|
||||
"32000: acc=0.63340625\n",
|
||||
"35200: acc=0.6525284090909091\n",
|
||||
"38400: acc=0.668515625\n",
|
||||
"41600: acc=0.6822596153846154\n",
|
||||
"44800: acc=0.6948214285714286\n",
|
||||
"48000: acc=0.7052708333333333\n",
|
||||
"51200: acc=0.71521484375\n",
|
||||
"54400: acc=0.7239889705882353\n",
|
||||
"57600: acc=0.7315277777777778\n",
|
||||
"60800: acc=0.7388486842105263\n",
|
||||
"64000: acc=0.74571875\n",
|
||||
"67200: acc=0.7518303571428572\n",
|
||||
"70400: acc=0.7576988636363636\n",
|
||||
"73600: acc=0.7628940217391305\n",
|
||||
"76800: acc=0.7681510416666667\n",
|
||||
"80000: acc=0.7728125\n",
|
||||
"83200: acc=0.7772235576923077\n",
|
||||
"86400: acc=0.7815393518518519\n",
|
||||
"89600: acc=0.7857700892857142\n",
|
||||
"92800: acc=0.7895043103448276\n",
|
||||
"96000: acc=0.7930520833333333\n",
|
||||
"99200: acc=0.7959072580645161\n",
|
||||
"102400: acc=0.798994140625\n",
|
||||
"105600: acc=0.802064393939394\n",
|
||||
"108800: acc=0.8051378676470589\n",
|
||||
"112000: acc=0.8077857142857143\n",
|
||||
"115200: acc=0.8104600694444445\n",
|
||||
"118400: acc=0.8128293918918919\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"(0.029785829671223958, 0.8138166666666666)"
|
||||
]
|
||||
},
|
||||
"execution_count": 8,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"net = LSTMPackClassifier(vocab_size,64,32,len(classes)).to(device)\n",
|
||||
"train_epoch_emb(net,train_loader_len, lr=0.001,use_pack_sequence=True)\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"> **नोट:** आपने देखा होगा कि हम प्रशिक्षण फ़ंक्शन को `use_pack_sequence` पैरामीटर पास करते हैं। वर्तमान में, `pack_padded_sequence` फ़ंक्शन को लंबाई अनुक्रम टेंसर को CPU डिवाइस पर होना आवश्यक है, और इसलिए प्रशिक्षण फ़ंक्शन को प्रशिक्षण के समय लंबाई अनुक्रम डेटा को GPU पर स्थानांतरित करने से बचना पड़ता है। आप [`torchnlp.py`](../../../../../lessons/5-NLP/16-RNN/torchnlp.py) फ़ाइल में `train_emb` फ़ंक्शन के कार्यान्वयन को देख सकते हैं।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## द्विदिश और बहुस्तरीय RNNs\n",
|
||||
"\n",
|
||||
"हमारे उदाहरणों में, सभी पुनरावर्ती नेटवर्क एक दिशा में काम करते थे, अनुक्रम की शुरुआत से अंत तक। यह स्वाभाविक लगता है, क्योंकि यह उस तरीके जैसा है जैसे हम पढ़ते हैं और भाषण सुनते हैं। हालांकि, कई व्यावहारिक मामलों में हमारे पास इनपुट अनुक्रम तक रैंडम एक्सेस होता है, इसलिए दोनों दिशाओं में पुनरावर्ती गणना चलाना समझदारी हो सकती है। ऐसे नेटवर्क को **द्विदिश** RNNs कहा जाता है, और इन्हें RNN/LSTM/GRU कंस्ट्रक्टर में `bidirectional=True` पैरामीटर पास करके बनाया जा सकता है।\n",
|
||||
"\n",
|
||||
"द्विदिश नेटवर्क के साथ काम करते समय, हमें दो छिपे हुए स्टेट वेक्टर की आवश्यकता होगी, प्रत्येक दिशा के लिए एक। PyTorch इन वेक्टरों को एक बड़े आकार के वेक्टर के रूप में एन्कोड करता है, जो काफी सुविधाजनक है, क्योंकि आमतौर पर आप परिणामी छिपे हुए स्टेट को पूरी तरह से कनेक्टेड लीनियर लेयर में पास करते हैं, और आपको केवल लेयर बनाते समय इस आकार में वृद्धि को ध्यान में रखना होगा।\n",
|
||||
"\n",
|
||||
"पुनरावर्ती नेटवर्क, चाहे वह एक-दिशात्मक हो या द्विदिश, अनुक्रम के भीतर कुछ पैटर्न को कैप्चर करता है और उन्हें स्टेट वेक्टर में स्टोर कर सकता है या आउटपुट में पास कर सकता है। जैसे कि कन्वोल्यूशनल नेटवर्क्स के साथ होता है, हम पहले लेयर द्वारा निकाले गए निम्न-स्तरीय पैटर्न से उच्च-स्तरीय पैटर्न कैप्चर करने के लिए पहले लेयर के ऊपर एक और पुनरावर्ती लेयर बना सकते हैं। यह हमें **बहुस्तरीय RNN** की अवधारणा तक ले जाता है, जिसमें दो या अधिक पुनरावर्ती नेटवर्क होते हैं, जहां पिछले लेयर का आउटपुट अगले लेयर को इनपुट के रूप में पास किया जाता है।\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"*यह चित्र [इस शानदार पोस्ट](https://towardsdatascience.com/from-a-lstm-cell-to-a-multilayer-lstm-network-with-pytorch-2899eb5696f3) से लिया गया है, जिसे फर्नांडो लोपेज़ ने लिखा है।*\n",
|
||||
"\n",
|
||||
"PyTorch ऐसे नेटवर्क बनाना आसान बनाता है, क्योंकि आपको केवल RNN/LSTM/GRU कंस्ट्रक्टर में `num_layers` पैरामीटर पास करना होता है, जिससे पुनरावृत्ति की कई लेयर स्वचालित रूप से बन जाती हैं। इसका मतलब यह भी होगा कि छिपे हुए/स्टेट वेक्टर का आकार आनुपातिक रूप से बढ़ेगा, और आपको पुनरावर्ती लेयर के आउटपुट को संभालते समय इसे ध्यान में रखना होगा।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## अन्य कार्यों के लिए RNNs\n",
|
||||
"\n",
|
||||
"इस यूनिट में, हमने देखा कि RNNs का उपयोग अनुक्रम वर्गीकरण के लिए किया जा सकता है, लेकिन वास्तव में, वे कई और कार्यों को संभाल सकते हैं, जैसे कि टेक्स्ट जनरेशन, मशीन अनुवाद, और अधिक। हम इन कार्यों पर अगले यूनिट में विचार करेंगे।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"\n---\n\n**अस्वीकरण**: \nयह दस्तावेज़ AI अनुवाद सेवा [Co-op Translator](https://github.com/Azure/co-op-translator) का उपयोग करके अनुवादित किया गया है। जबकि हम सटीकता के लिए प्रयासरत हैं, कृपया ध्यान दें कि स्वचालित अनुवाद में त्रुटियां या अशुद्धियां हो सकती हैं। मूल भाषा में उपलब्ध मूल दस्तावेज़ को आधिकारिक स्रोत माना जाना चाहिए। महत्वपूर्ण जानकारी के लिए, पेशेवर मानव अनुवाद की सिफारिश की जाती है। इस अनुवाद के उपयोग से उत्पन्न किसी भी गलतफहमी या गलत व्याख्या के लिए हम उत्तरदायी नहीं हैं।\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
"interpreter": {
|
||||
"hash": "16af2a8bbb083ea23e5e41c7f5787656b2ce26968575d8763f2c4b17f9cd711f"
|
||||
},
|
||||
"kernelspec": {
|
||||
"display_name": "Python 3.8.12 ('py38')",
|
||||
"language": "python",
|
||||
"name": "python3"
|
||||
},
|
||||
"language_info": {
|
||||
"codemirror_mode": {
|
||||
"name": "ipython",
|
||||
"version": 3
|
||||
},
|
||||
"file_extension": ".py",
|
||||
"mimetype": "text/x-python",
|
||||
"name": "python",
|
||||
"nbconvert_exporter": "python",
|
||||
"pygments_lexer": "ipython3",
|
||||
"version": "3.8.12"
|
||||
},
|
||||
"coopTranslator": {
|
||||
"original_hash": "522ee52ae3d5ae933e283286254e9a55",
|
||||
"translation_date": "2025-08-31T15:24:12+00:00",
|
||||
"source_file": "lessons/5-NLP/16-RNN/RNNPyTorch.ipynb",
|
||||
"language_code": "hi"
|
||||
}
|
||||
},
|
||||
"nbformat": 4,
|
||||
"nbformat_minor": 2
|
||||
}
|
||||
|
|
@ -0,0 +1,460 @@
|
|||
{
|
||||
"cells": [
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"# पुनरावर्ती न्यूरल नेटवर्क\n",
|
||||
"\n",
|
||||
"पिछले मॉड्यूल में, हमने टेक्स्ट के समृद्ध अर्थपूर्ण प्रतिनिधित्व को कवर किया। जिस आर्किटेक्चर का हम उपयोग कर रहे हैं, वह वाक्य में शब्दों के समग्र अर्थ को कैप्चर करता है, लेकिन यह शब्दों के **क्रम** को ध्यान में नहीं रखता है, क्योंकि एम्बेडिंग के बाद का समेकन ऑपरेशन मूल टेक्स्ट से इस जानकारी को हटा देता है। चूंकि ये मॉडल शब्दों के क्रम को प्रदर्शित करने में असमर्थ हैं, वे टेक्स्ट जनरेशन या प्रश्न उत्तर जैसे अधिक जटिल या अस्पष्ट कार्यों को हल नहीं कर सकते।\n",
|
||||
"\n",
|
||||
"टेक्स्ट अनुक्रम के अर्थ को कैप्चर करने के लिए, हम एक न्यूरल नेटवर्क आर्किटेक्चर का उपयोग करेंगे जिसे **पुनरावर्ती न्यूरल नेटवर्क** या RNN कहा जाता है। RNN का उपयोग करते समय, हम अपने वाक्य को नेटवर्क के माध्यम से एक-एक टोकन करके पास करते हैं, और नेटवर्क कुछ **स्थिति** उत्पन्न करता है, जिसे हम अगले टोकन के साथ नेटवर्क में फिर से पास करते हैं।\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"दिए गए टोकन अनुक्रम $X_0,\\dots,X_n$ के लिए, RNN एक न्यूरल नेटवर्क ब्लॉकों का अनुक्रम बनाता है, और इस अनुक्रम को बैकप्रोपेगेशन का उपयोग करके एंड-टू-एंड ट्रेन करता है। प्रत्येक नेटवर्क ब्लॉक $(X_i,S_i)$ की एक जोड़ी को इनपुट के रूप में लेता है, और $S_{i+1}$ को परिणाम के रूप में उत्पन्न करता है। अंतिम स्थिति $S_n$ या आउटपुट $Y_n$ एक रैखिक वर्गीकरणकर्ता में जाती है ताकि परिणाम उत्पन्न हो सके। सभी नेटवर्क ब्लॉक समान वेट्स साझा करते हैं, और एक बैकप्रोपेगेशन पास का उपयोग करके एंड-टू-एंड ट्रेन किए जाते हैं।\n",
|
||||
"\n",
|
||||
"> ऊपर दी गई आकृति पुनरावर्ती न्यूरल नेटवर्क को अनरोल्ड रूप (बाईं ओर) और अधिक कॉम्पैक्ट पुनरावर्ती प्रतिनिधित्व (दाईं ओर) में दिखाती है। यह समझना महत्वपूर्ण है कि सभी RNN सेल्स के समान **शेयर करने योग्य वेट्स** होते हैं।\n",
|
||||
"\n",
|
||||
"चूंकि स्थिति वेक्टर $S_0,\\dots,S_n$ नेटवर्क के माध्यम से पास किए जाते हैं, RNN शब्दों के बीच क्रमिक निर्भरता सीखने में सक्षम होता है। उदाहरण के लिए, जब अनुक्रम में कहीं *not* शब्द आता है, तो यह स्थिति वेक्टर के भीतर कुछ तत्वों को नकारना सीख सकता है।\n",
|
||||
"\n",
|
||||
"अंदर, प्रत्येक RNN सेल में दो वेट मैट्रिक्स होते हैं: $W_H$ और $W_I$, और बायस $b$। प्रत्येक RNN चरण में, दिए गए इनपुट $X_i$ और इनपुट स्थिति $S_i$, आउटपुट स्थिति की गणना इस प्रकार की जाती है: $S_{i+1} = f(W_H\\times S_i + W_I\\times X_i+b)$, जहां $f$ एक सक्रियण फ़ंक्शन है (अक्सर $\\tanh$)।\n",
|
||||
"\n",
|
||||
"> टेक्स्ट जनरेशन (जिसे हम अगले यूनिट में कवर करेंगे) या मशीन अनुवाद जैसी समस्याओं के लिए, हम प्रत्येक RNN चरण में कुछ आउटपुट मान भी प्राप्त करना चाहते हैं। इस मामले में, एक और मैट्रिक्स $W_O$ होता है, और आउटपुट की गणना इस प्रकार की जाती है: $Y_i=f(W_O\\times S_i+b_O)$।\n",
|
||||
"\n",
|
||||
"आइए देखें कि पुनरावर्ती न्यूरल नेटवर्क हमारे समाचार डेटासेट को वर्गीकृत करने में कैसे मदद कर सकते हैं।\n",
|
||||
"\n",
|
||||
"> सैंडबॉक्स वातावरण के लिए, हमें यह सुनिश्चित करने के लिए निम्नलिखित सेल चलाना होगा कि आवश्यक लाइब्रेरी इंस्टॉल हो गई है, और डेटा प्रीफेच हो गया है। यदि आप लोकल रूप से चला रहे हैं, तो आप निम्नलिखित सेल को छोड़ सकते हैं।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 1,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"import sys\n",
|
||||
"!{sys.executable} -m pip install --quiet tensorflow_datasets==4.4.0\n",
|
||||
"!cd ~ && wget -q -O - https://mslearntensorflowlp.blob.core.windows.net/data/tfds-ag-news.tgz | tar xz"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 2,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"import tensorflow as tf\n",
|
||||
"from tensorflow import keras\n",
|
||||
"import tensorflow_datasets as tfds\n",
|
||||
"import numpy as np\n",
|
||||
"\n",
|
||||
"# We are going to be training pretty large models. In order not to face errors, we need\n",
|
||||
"# to set tensorflow option to grow GPU memory allocation when required\n",
|
||||
"physical_devices = tf.config.list_physical_devices('GPU') \n",
|
||||
"if len(physical_devices)>0:\n",
|
||||
" tf.config.experimental.set_memory_growth(physical_devices[0], True)\n",
|
||||
"\n",
|
||||
"ds_train, ds_test = tfds.load('ag_news_subset').values()"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {
|
||||
"nteract": {
|
||||
"transient": {
|
||||
"deleting": false
|
||||
}
|
||||
}
|
||||
},
|
||||
"source": [
|
||||
"जब बड़े मॉडल का प्रशिक्षण किया जाता है, तो GPU मेमोरी आवंटन एक समस्या बन सकता है। हमें विभिन्न मिनीबैच आकारों के साथ प्रयोग करने की आवश्यकता हो सकती है, ताकि डेटा हमारे GPU मेमोरी में फिट हो जाए और प्रशिक्षण पर्याप्त तेज़ हो। यदि आप इस कोड को अपने GPU मशीन पर चला रहे हैं, तो आप प्रशिक्षण को तेज़ करने के लिए मिनीबैच आकार को समायोजित करने के साथ प्रयोग कर सकते हैं।\n",
|
||||
"\n",
|
||||
"> **Note**: NVidia ड्राइवरों के कुछ संस्करणों के बारे में ज्ञात है कि वे मॉडल का प्रशिक्षण समाप्त होने के बाद मेमोरी को रिलीज़ नहीं करते। हम इस नोटबुक में कई उदाहरण चला रहे हैं, और यह कुछ सेटअप में मेमोरी समाप्त होने का कारण बन सकता है, खासकर यदि आप इसी नोटबुक में अपने स्वयं के प्रयोग कर रहे हैं। यदि आप मॉडल का प्रशिक्षण शुरू करते समय कुछ अजीब त्रुटियों का सामना करते हैं, तो आप नोटबुक कर्नेल को पुनः आरंभ करना चाह सकते हैं।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 3,
|
||||
"metadata": {
|
||||
"collapsed": true,
|
||||
"jupyter": {
|
||||
"outputs_hidden": false,
|
||||
"source_hidden": false
|
||||
},
|
||||
"nteract": {
|
||||
"transient": {
|
||||
"deleting": false
|
||||
}
|
||||
}
|
||||
},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"batch_size = 16\n",
|
||||
"embed_size = 64"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## सरल RNN क्लासिफायर\n",
|
||||
"\n",
|
||||
"सरल RNN के मामले में, प्रत्येक पुनरावर्ती यूनिट एक साधारण रैखिक नेटवर्क होता है, जो एक इनपुट वेक्टर और एक स्टेट वेक्टर लेता है, और एक नया स्टेट वेक्टर उत्पन्न करता है। Keras में, इसे `SimpleRNN` लेयर द्वारा दर्शाया जा सकता है।\n",
|
||||
"\n",
|
||||
"हालांकि हम RNN लेयर को सीधे वन-हॉट एन्कोडेड टोकन पास कर सकते हैं, यह एक अच्छा विचार नहीं है क्योंकि उनकी उच्च आयामीयता होती है। इसलिए, हम शब्द वेक्टर की आयामीयता को कम करने के लिए एक एम्बेडिंग लेयर का उपयोग करेंगे, इसके बाद एक RNN लेयर और अंत में एक `Dense` क्लासिफायर।\n",
|
||||
"\n",
|
||||
"> **Note**: उन मामलों में जहां आयामीयता इतनी अधिक नहीं होती, जैसे कि जब कैरेक्टर-लेवल टोकनाइज़ेशन का उपयोग किया जाता है, तो वन-हॉट एन्कोडेड टोकन को सीधे RNN सेल में पास करना समझदारी हो सकती है।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Model: \"sequential\"\n",
|
||||
"_________________________________________________________________\n",
|
||||
"Layer (type) Output Shape Param # \n",
|
||||
"=================================================================\n",
|
||||
"text_vectorization (TextVect (None, None) 0 \n",
|
||||
"_________________________________________________________________\n",
|
||||
"embedding (Embedding) (None, None, 64) 1280000 \n",
|
||||
"_________________________________________________________________\n",
|
||||
"simple_rnn (SimpleRNN) (None, 16) 1296 \n",
|
||||
"_________________________________________________________________\n",
|
||||
"dense (Dense) (None, 4) 68 \n",
|
||||
"=================================================================\n",
|
||||
"Total params: 1,281,364\n",
|
||||
"Trainable params: 1,281,364\n",
|
||||
"Non-trainable params: 0\n",
|
||||
"_________________________________________________________________\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"vocab_size = 20000\n",
|
||||
"\n",
|
||||
"vectorizer = keras.layers.experimental.preprocessing.TextVectorization(\n",
|
||||
" max_tokens=vocab_size,\n",
|
||||
" input_shape=(1,))\n",
|
||||
"\n",
|
||||
"model = keras.models.Sequential([\n",
|
||||
" vectorizer,\n",
|
||||
" keras.layers.Embedding(vocab_size, embed_size),\n",
|
||||
" keras.layers.SimpleRNN(16),\n",
|
||||
" keras.layers.Dense(4,activation='softmax')\n",
|
||||
"])\n",
|
||||
"\n",
|
||||
"model.summary()"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"> **ध्यान दें:** यहाँ हम सरलता के लिए एक अप्रशिक्षित एम्बेडिंग लेयर का उपयोग कर रहे हैं, लेकिन बेहतर परिणामों के लिए हम Word2Vec का उपयोग करके एक पूर्व-प्रशिक्षित एम्बेडिंग लेयर का उपयोग कर सकते हैं, जैसा कि पिछले यूनिट में बताया गया है। यह आपके लिए एक अच्छा अभ्यास होगा कि आप इस कोड को पूर्व-प्रशिक्षित एम्बेडिंग के साथ काम करने के लिए अनुकूलित करें।\n",
|
||||
"\n",
|
||||
"अब चलिए अपने RNN को प्रशिक्षित करते हैं। सामान्यतः RNN को प्रशिक्षित करना काफी कठिन होता है, क्योंकि जब RNN सेल्स को अनुक्रम की लंबाई के साथ अनरोल किया जाता है, तो बैकप्रोपेगेशन में शामिल लेयर्स की संख्या काफी अधिक हो जाती है। इसलिए हमें एक छोटा लर्निंग रेट चुनने की आवश्यकता होती है, और अच्छे परिणाम प्राप्त करने के लिए नेटवर्क को एक बड़े डेटासेट पर प्रशिक्षित करना होता है। इसमें काफी समय लग सकता है, इसलिए GPU का उपयोग करना बेहतर होता है।\n",
|
||||
"\n",
|
||||
"प्रक्रिया को तेज़ करने के लिए, हम केवल समाचार शीर्षकों पर RNN मॉडल को प्रशिक्षित करेंगे और विवरण को छोड़ देंगे। आप विवरण के साथ प्रशिक्षण का प्रयास कर सकते हैं और देख सकते हैं कि क्या आप मॉडल को प्रशिक्षित कर सकते हैं।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 5,
|
||||
"metadata": {
|
||||
"scrolled": true
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Training vectorizer\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"def extract_title(x):\n",
|
||||
" return x['title']\n",
|
||||
"\n",
|
||||
"def tupelize_title(x):\n",
|
||||
" return (extract_title(x),x['label'])\n",
|
||||
"\n",
|
||||
"print('Training vectorizer')\n",
|
||||
"vectorizer.adapt(ds_train.take(2000).map(extract_title))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 6,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"7500/7500 [==============================] - 82s 11ms/step - loss: 0.6629 - acc: 0.7623 - val_loss: 0.5559 - val_acc: 0.7995\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"<tensorflow.python.keras.callbacks.History at 0x7f3e0030d350>"
|
||||
]
|
||||
},
|
||||
"execution_count": 6,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"model.compile(loss='sparse_categorical_crossentropy',metrics=['acc'], optimizer='adam')\n",
|
||||
"model.fit(ds_train.map(tupelize_title).batch(batch_size),validation_data=ds_test.map(tupelize_title).batch(batch_size))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {
|
||||
"nteract": {
|
||||
"transient": {
|
||||
"deleting": false
|
||||
}
|
||||
}
|
||||
},
|
||||
"source": [
|
||||
"> **नोट** कि सटीकता यहां कम हो सकती है, क्योंकि हम केवल समाचार शीर्षकों पर प्रशिक्षण कर रहे हैं।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## वेरिएबल अनुक्रमों पर पुनर्विचार \n",
|
||||
"\n",
|
||||
"याद रखें कि `TextVectorization` लेयर स्वचालित रूप से एक मिनीबैच में वेरिएबल लंबाई वाले अनुक्रमों को पैड टोकन के साथ पैड कर देती है। यह देखा गया है कि ये टोकन भी प्रशिक्षण में भाग लेते हैं, और वे मॉडल के कन्वर्जेंस को जटिल बना सकते हैं।\n",
|
||||
"\n",
|
||||
"पैडिंग की मात्रा को कम करने के लिए हम कई तरीकों का उपयोग कर सकते हैं। उनमें से एक है डेटा सेट को अनुक्रम की लंबाई के अनुसार पुनः व्यवस्थित करना और सभी अनुक्रमों को उनके आकार के अनुसार समूहित करना। इसे `tf.data.experimental.bucket_by_sequence_length` फ़ंक्शन का उपयोग करके किया जा सकता है (देखें [डॉक्यूमेंटेशन](https://www.tensorflow.org/api_docs/python/tf/data/experimental/bucket_by_sequence_length))।\n",
|
||||
"\n",
|
||||
"एक अन्य तरीका **मास्किंग** का उपयोग करना है। Keras में, कुछ लेयर अतिरिक्त इनपुट का समर्थन करती हैं जो दिखाती हैं कि प्रशिक्षण के दौरान किन टोकनों को ध्यान में रखा जाना चाहिए। मास्किंग को अपने मॉडल में शामिल करने के लिए, हम या तो एक अलग `Masking` लेयर ([डॉक्स](https://keras.io/api/layers/core_layers/masking/)) जोड़ सकते हैं, या हम अपने `Embedding` लेयर में `mask_zero=True` पैरामीटर निर्दिष्ट कर सकते हैं।\n",
|
||||
"\n",
|
||||
"> **Note**: इस प्रशिक्षण में पूरे डेटा सेट पर एक एपोक पूरा करने में लगभग 5 मिनट लगेंगे। यदि आप धैर्य खो दें तो प्रशिक्षण को किसी भी समय रोकने के लिए स्वतंत्र महसूस करें। आप यह भी कर सकते हैं कि प्रशिक्षण के लिए उपयोग किए जाने वाले डेटा की मात्रा को सीमित करें, `ds_train` और `ds_test` डेटा सेट के बाद `.take(...)` क्लॉज जोड़कर।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 7,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"7500/7500 [==============================] - 371s 49ms/step - loss: 0.5401 - acc: 0.8079 - val_loss: 0.3780 - val_acc: 0.8822\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"<tensorflow.python.keras.callbacks.History at 0x7f3dec118850>"
|
||||
]
|
||||
},
|
||||
"execution_count": 7,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"def extract_text(x):\n",
|
||||
" return x['title']+' '+x['description']\n",
|
||||
"\n",
|
||||
"def tupelize(x):\n",
|
||||
" return (extract_text(x),x['label'])\n",
|
||||
"\n",
|
||||
"model = keras.models.Sequential([\n",
|
||||
" vectorizer,\n",
|
||||
" keras.layers.Embedding(vocab_size,embed_size,mask_zero=True),\n",
|
||||
" keras.layers.SimpleRNN(16),\n",
|
||||
" keras.layers.Dense(4,activation='softmax')\n",
|
||||
"])\n",
|
||||
"\n",
|
||||
"model.compile(loss='sparse_categorical_crossentropy',metrics=['acc'], optimizer='adam')\n",
|
||||
"model.fit(ds_train.map(tupelize).batch(batch_size),validation_data=ds_test.map(tupelize).batch(batch_size))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"अब जब हम मास्किंग का उपयोग कर रहे हैं, तो हम शीर्षकों और विवरणों के पूरे डेटासेट पर मॉडल को प्रशिक्षित कर सकते हैं।\n",
|
||||
"\n",
|
||||
"> **नोट**: क्या आपने देखा है कि हम उस वेक्टराइज़र का उपयोग कर रहे हैं जिसे समाचार शीर्षकों पर प्रशिक्षित किया गया था, न कि लेख के पूरे मुख्य भाग पर? संभवतः, इससे कुछ टोकन अनदेखे हो सकते हैं, इसलिए वेक्टराइज़र को फिर से प्रशिक्षित करना बेहतर होगा। हालांकि, इसका प्रभाव बहुत छोटा हो सकता है, इसलिए हम सरलता के लिए पहले से प्रशिक्षित वेक्टराइज़र का उपयोग जारी रखेंगे।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## LSTM: लंबी अवधि की स्मृति\n",
|
||||
"\n",
|
||||
"RNNs की मुख्य समस्याओं में से एक है **vanishing gradients**। RNNs काफी लंबे हो सकते हैं, और बैकप्रोपेगेशन के दौरान नेटवर्क की पहली परत तक ग्रेडिएंट्स को पूरी तरह से वापस ले जाना मुश्किल हो सकता है। जब ऐसा होता है, तो नेटवर्क दूरस्थ टोकन के बीच संबंधों को सीखने में असमर्थ हो जाता है। इस समस्या से बचने का एक तरीका है **स्पष्ट स्थिति प्रबंधन** को **गेट्स** का उपयोग करके लागू करना। गेट्स को पेश करने वाली दो सबसे सामान्य आर्किटेक्चर हैं **लंबी अवधि की स्मृति** (LSTM) और **गेटेड रिले यूनिट** (GRU)। यहां हम LSTMs को कवर करेंगे।\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"एक LSTM नेटवर्क को RNN के समान तरीके से व्यवस्थित किया जाता है, लेकिन इसमें दो अवस्थाएं होती हैं जो परत से परत तक पास की जाती हैं: वास्तविक स्थिति $c$, और छिपा हुआ वेक्टर $h$। प्रत्येक यूनिट पर, छिपा हुआ वेक्टर $h_{t-1}$ को इनपुट $x_t$ के साथ जोड़ा जाता है, और ये दोनों मिलकर **गेट्स** के माध्यम से स्थिति $c_t$ और आउटपुट $h_{t}$ पर नियंत्रण करते हैं। प्रत्येक गेट में सिग्मॉइड सक्रियता होती है (आउटपुट $[0,1]$ की सीमा में), जिसे स्थिति वेक्टर के साथ गुणा करने पर बिटवाइज मास्क के रूप में सोचा जा सकता है। LSTMs में निम्नलिखित गेट्स होते हैं (ऊपर दी गई तस्वीर में बाएं से दाएं):\n",
|
||||
"* **भूल गेट** जो यह निर्धारित करता है कि वेक्टर $c_{t-1}$ के कौन से घटकों को हमें भूलना है, और कौन से पास करने हैं।\n",
|
||||
"* **इनपुट गेट** जो यह तय करता है कि इनपुट वेक्टर और पिछले छिपे हुए वेक्टर से कितनी जानकारी को स्थिति वेक्टर में शामिल करना चाहिए।\n",
|
||||
"* **आउटपुट गेट** जो नई स्थिति वेक्टर लेता है और तय करता है कि इसके कौन से घटकों का उपयोग नए छिपे हुए वेक्टर $h_t$ को उत्पन्न करने के लिए किया जाएगा।\n",
|
||||
"\n",
|
||||
"स्थिति $c$ के घटकों को ऐसे फ्लैग्स के रूप में सोचा जा सकता है जिन्हें चालू और बंद किया जा सकता है। उदाहरण के लिए, जब हम अनुक्रम में नाम *Alice* का सामना करते हैं, तो हम अनुमान लगाते हैं कि यह एक महिला को संदर्भित करता है, और स्थिति में वह फ्लैग उठाते हैं जो कहता है कि हमारे पास वाक्य में एक स्त्रीलिंग संज्ञा है। जब हम आगे *and Tom* शब्दों का सामना करते हैं, तो हम वह फ्लैग उठाते हैं जो कहता है कि हमारे पास बहुवचन संज्ञा है। इस प्रकार स्थिति में हेरफेर करके हम वाक्य के व्याकरणिक गुणों का ट्रैक रख सकते हैं।\n",
|
||||
"\n",
|
||||
"> **Note**: LSTMs की आंतरिक संरचना को समझने के लिए यहां एक शानदार संसाधन है: [Understanding LSTM Networks](https://colah.github.io/posts/2015-08-Understanding-LSTMs/) क्रिस्टोफर ओलाह द्वारा।\n",
|
||||
"\n",
|
||||
"हालांकि LSTM सेल की आंतरिक संरचना जटिल लग सकती है, Keras इस कार्यान्वयन को `LSTM` लेयर के अंदर छुपा देता है, इसलिए ऊपर दिए गए उदाहरण में हमें केवल पुनरावर्ती लेयर को बदलने की आवश्यकता है:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 8,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"15000/15000 [==============================] - 188s 13ms/step - loss: 0.5692 - acc: 0.7916 - val_loss: 0.3441 - val_acc: 0.8870\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"<tensorflow.python.keras.callbacks.History at 0x7f3d6af5c350>"
|
||||
]
|
||||
},
|
||||
"execution_count": 8,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"model = keras.models.Sequential([\n",
|
||||
" vectorizer,\n",
|
||||
" keras.layers.Embedding(vocab_size, embed_size),\n",
|
||||
" keras.layers.LSTM(8),\n",
|
||||
" keras.layers.Dense(4,activation='softmax')\n",
|
||||
"])\n",
|
||||
"\n",
|
||||
"model.compile(loss='sparse_categorical_crossentropy',metrics=['acc'], optimizer='adam')\n",
|
||||
"model.fit(ds_train.map(tupelize).batch(8),validation_data=ds_test.map(tupelize).batch(8))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": []
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## द्विदिश और बहुस्तरीय RNNs\n",
|
||||
"\n",
|
||||
"अब तक के हमारे उदाहरणों में, पुनरावर्ती नेटवर्क एक अनुक्रम की शुरुआत से अंत तक काम करते हैं। यह हमें स्वाभाविक लगता है क्योंकि यह उसी दिशा का अनुसरण करता है जिसमें हम पढ़ते हैं या भाषण सुनते हैं। हालांकि, उन परिस्थितियों के लिए जहां इनपुट अनुक्रम का रैंडम एक्सेस आवश्यक है, दोनों दिशाओं में पुनरावर्ती गणना चलाना अधिक समझदारी भरा होता है। ऐसे RNNs जो दोनों दिशाओं में गणना की अनुमति देते हैं, उन्हें **द्विदिश** RNNs कहा जाता है, और इन्हें एक विशेष `Bidirectional` लेयर के साथ पुनरावर्ती लेयर को लपेटकर बनाया जा सकता है।\n",
|
||||
"\n",
|
||||
"> **Note**: `Bidirectional` लेयर अपनी भीतर की लेयर की दो प्रतियां बनाती है और उनमें से एक की `go_backwards` प्रॉपर्टी को `True` सेट करती है, जिससे वह अनुक्रम के साथ विपरीत दिशा में जाती है।\n",
|
||||
"\n",
|
||||
"पुनरावर्ती नेटवर्क, चाहे एक दिशा में हो या द्विदिश, अनुक्रम के भीतर पैटर्न को कैप्चर करते हैं और उन्हें स्टेट वेक्टर में स्टोर करते हैं या आउटपुट के रूप में लौटाते हैं। जैसे कि कन्वोल्यूशनल नेटवर्क्स में होता है, हम पहले लेयर द्वारा निकाले गए निम्न स्तर के पैटर्न से उच्च स्तर के पैटर्न कैप्चर करने के लिए पहले लेयर के बाद एक और पुनरावर्ती लेयर बना सकते हैं। यह हमें **बहुस्तरीय RNN** की अवधारणा तक ले जाता है, जिसमें दो या अधिक पुनरावर्ती नेटवर्क होते हैं, जहां पिछले लेयर का आउटपुट अगले लेयर को इनपुट के रूप में दिया जाता है।\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"*फर्नांडो लोपेज़ द्वारा [इस शानदार पोस्ट](https://towardsdatascience.com/from-a-lstm-cell-to-a-multilayer-lstm-network-with-pytorch-2899eb5696f3) से ली गई तस्वीर।*\n",
|
||||
"\n",
|
||||
"Keras इन नेटवर्क्स को बनाना आसान बनाता है, क्योंकि आपको बस मॉडल में अधिक पुनरावर्ती लेयर जोड़नी होती है। अंतिम लेयर को छोड़कर सभी लेयर के लिए, हमें `return_sequences=True` पैरामीटर निर्दिष्ट करना होता है, क्योंकि हमें लेयर से सभी मध्यवर्ती स्टेट्स चाहिए, न कि केवल पुनरावर्ती गणना की अंतिम स्टेट।\n",
|
||||
"\n",
|
||||
"आइए हमारे वर्गीकरण समस्या के लिए एक दो-लेयर द्विदिश LSTM बनाते हैं।\n",
|
||||
"\n",
|
||||
"> **Note** यह कोड फिर से पूरा होने में काफी समय लेता है, लेकिन यह हमें अब तक देखी गई सबसे अधिक सटीकता देता है। तो शायद इंतजार करना और परिणाम देखना उचित हो सकता है।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 9,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"5044/7500 [===================>..........] - ETA: 2:33 - loss: 0.3709 - acc: 0.8706\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\b\r5045/7500 [===================>..........] - ETA: 2:33 - loss: 0.3709 - acc: 0.8706"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"model = keras.models.Sequential([\n",
|
||||
" vectorizer,\n",
|
||||
" keras.layers.Embedding(vocab_size, 128, mask_zero=True),\n",
|
||||
" keras.layers.Bidirectional(keras.layers.LSTM(64,return_sequences=True)),\n",
|
||||
" keras.layers.Bidirectional(keras.layers.LSTM(64)), \n",
|
||||
" keras.layers.Dense(4,activation='softmax')\n",
|
||||
"])\n",
|
||||
"\n",
|
||||
"model.compile(loss='sparse_categorical_crossentropy',metrics=['acc'], optimizer='adam')\n",
|
||||
"model.fit(ds_train.map(tupelize).batch(batch_size),\n",
|
||||
" validation_data=ds_test.map(tupelize).batch(batch_size))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## अन्य कार्यों के लिए RNNs\n",
|
||||
"\n",
|
||||
"अब तक, हमने RNNs का उपयोग टेक्स्ट अनुक्रमों को वर्गीकृत करने के लिए किया है। लेकिन वे और भी कई कार्य संभाल सकते हैं, जैसे टेक्स्ट जनरेशन और मशीन ट्रांसलेशन — हम इन कार्यों पर अगले यूनिट में विचार करेंगे।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"\n---\n\n**अस्वीकरण**: \nयह दस्तावेज़ AI अनुवाद सेवा [Co-op Translator](https://github.com/Azure/co-op-translator) का उपयोग करके अनुवादित किया गया है। जबकि हम सटीकता सुनिश्चित करने का प्रयास करते हैं, कृपया ध्यान दें कि स्वचालित अनुवाद में त्रुटियां या अशुद्धियां हो सकती हैं। मूल भाषा में उपलब्ध मूल दस्तावेज़ को प्रामाणिक स्रोत माना जाना चाहिए। महत्वपूर्ण जानकारी के लिए, पेशेवर मानव अनुवाद की सिफारिश की जाती है। इस अनुवाद के उपयोग से उत्पन्न किसी भी गलतफहमी या गलत व्याख्या के लिए हम जिम्मेदार नहीं हैं।\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
"kernel_info": {
|
||||
"name": "conda-env-py37_tensorflow-py"
|
||||
},
|
||||
"kernelspec": {
|
||||
"display_name": "py37_tensorflow",
|
||||
"language": "python",
|
||||
"name": "conda-env-py37_tensorflow-py"
|
||||
},
|
||||
"language_info": {
|
||||
"codemirror_mode": {
|
||||
"name": "ipython",
|
||||
"version": 3
|
||||
},
|
||||
"file_extension": ".py",
|
||||
"mimetype": "text/x-python",
|
||||
"name": "python",
|
||||
"nbconvert_exporter": "python",
|
||||
"pygments_lexer": "ipython3",
|
||||
"version": "3.7.9"
|
||||
},
|
||||
"nteract": {
|
||||
"version": "nteract-front-end@1.0.0"
|
||||
},
|
||||
"coopTranslator": {
|
||||
"original_hash": "81351e61f619b432ff51010a4f993194",
|
||||
"translation_date": "2025-08-31T15:22:33+00:00",
|
||||
"source_file": "lessons/5-NLP/16-RNN/RNNTF.ipynb",
|
||||
"language_code": "hi"
|
||||
}
|
||||
},
|
||||
"nbformat": 4,
|
||||
"nbformat_minor": 4
|
||||
}
|
||||
|
|
@ -0,0 +1,414 @@
|
|||
{
|
||||
"cells": [
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"# जनरेटिव नेटवर्क्स\n",
|
||||
"\n",
|
||||
"रिकरेंट न्यूरल नेटवर्क्स (RNNs) और उनके गेटेड सेल वेरिएंट्स जैसे लॉन्ग शॉर्ट टर्म मेमोरी सेल्स (LSTMs) और गेटेड रिकारेंट यूनिट्स (GRUs) ने भाषा मॉडलिंग के लिए एक तंत्र प्रदान किया, यानी वे शब्दों के क्रम को सीख सकते हैं और अनुक्रम में अगले शब्द की भविष्यवाणी कर सकते हैं। यह हमें RNNs का उपयोग **जनरेटिव कार्यों** के लिए करने की अनुमति देता है, जैसे साधारण टेक्स्ट जनरेशन, मशीन ट्रांसलेशन, और यहां तक कि इमेज कैप्शनिंग।\n",
|
||||
"\n",
|
||||
"पिछली यूनिट में हमने जिस RNN आर्किटेक्चर पर चर्चा की थी, उसमें प्रत्येक RNN यूनिट ने अगले हिडन स्टेट को आउटपुट के रूप में उत्पन्न किया। हालांकि, हम प्रत्येक रिकारेंट यूनिट में एक और आउटपुट जोड़ सकते हैं, जो हमें एक **अनुक्रम** आउटपुट करने की अनुमति देगा (जो मूल अनुक्रम की लंबाई के बराबर होगा)। इसके अलावा, हम ऐसे RNN यूनिट्स का उपयोग कर सकते हैं जो प्रत्येक चरण में इनपुट स्वीकार नहीं करते, बल्कि केवल एक प्रारंभिक स्टेट वेक्टर लेते हैं और फिर आउटपुट का एक अनुक्रम उत्पन्न करते हैं।\n",
|
||||
"\n",
|
||||
"इस नोटबुक में, हम सरल जनरेटिव मॉडलों पर ध्यान केंद्रित करेंगे जो हमें टेक्स्ट जनरेट करने में मदद करते हैं। सरलता के लिए, चलिए **कैरेक्टर-लेवल नेटवर्क** बनाते हैं, जो अक्षर दर अक्षर टेक्स्ट जनरेट करता है। प्रशिक्षण के दौरान, हमें कुछ टेक्स्ट कॉर्पस लेना होगा और इसे अक्षर अनुक्रमों में विभाजित करना होगा।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 1,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Loading dataset...\n",
|
||||
"Building vocab...\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"import torch\n",
|
||||
"import torchtext\n",
|
||||
"import numpy as np\n",
|
||||
"from torchnlp import *\n",
|
||||
"train_dataset,test_dataset,classes,vocab = load_dataset()"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## वर्णमाला शब्दावली बनाना\n",
|
||||
"\n",
|
||||
"वर्ण-स्तरीय जनरेटिव नेटवर्क बनाने के लिए, हमें टेक्स्ट को शब्दों के बजाय व्यक्तिगत अक्षरों में विभाजित करना होगा। यह एक अलग टोकनाइज़र को परिभाषित करके किया जा सकता है:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 2,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Vocabulary size = 82\n",
|
||||
"Encoding of 'a' is 1\n",
|
||||
"Character with code 13 is c\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"def char_tokenizer(words):\n",
|
||||
" return list(words) #[word for word in words]\n",
|
||||
"\n",
|
||||
"counter = collections.Counter()\n",
|
||||
"for (label, line) in train_dataset:\n",
|
||||
" counter.update(char_tokenizer(line))\n",
|
||||
"vocab = torchtext.vocab.vocab(counter)\n",
|
||||
"\n",
|
||||
"vocab_size = len(vocab)\n",
|
||||
"print(f\"Vocabulary size = {vocab_size}\")\n",
|
||||
"print(f\"Encoding of 'a' is {vocab.get_stoi()['a']}\")\n",
|
||||
"print(f\"Character with code 13 is {vocab.get_itos()[13]}\")"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"आइए देखें कि हम अपने डेटासेट से टेक्स्ट को कैसे एन्कोड कर सकते हैं:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 3,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"tensor([ 0, 1, 2, 2, 3, 4, 5, 6, 3, 7, 8, 1, 9, 10, 3, 11, 2, 1,\n",
|
||||
" 12, 3, 7, 1, 13, 14, 3, 15, 16, 5, 17, 3, 5, 18, 8, 3, 7, 2,\n",
|
||||
" 1, 13, 14, 3, 19, 20, 8, 21, 5, 8, 9, 10, 22, 3, 20, 8, 21, 5,\n",
|
||||
" 8, 9, 10, 3, 23, 3, 4, 18, 17, 9, 5, 23, 10, 8, 2, 2, 8, 9,\n",
|
||||
" 10, 24, 3, 0, 1, 2, 2, 3, 4, 5, 9, 8, 8, 5, 25, 10, 3, 26,\n",
|
||||
" 12, 27, 16, 26, 2, 27, 16, 28, 29, 30, 1, 16, 26, 3, 17, 31, 3, 21,\n",
|
||||
" 2, 5, 9, 1, 23, 13, 32, 16, 27, 13, 10, 24, 3, 1, 9, 8, 3, 10,\n",
|
||||
" 8, 8, 27, 16, 28, 3, 28, 9, 8, 8, 16, 3, 1, 28, 1, 27, 16, 6])"
|
||||
]
|
||||
},
|
||||
"execution_count": 3,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"def enc(x):\n",
|
||||
" return torch.LongTensor(encode(x,voc=vocab,tokenizer=char_tokenizer))\n",
|
||||
"\n",
|
||||
"enc(train_dataset[0][1])"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## जनरेटिव RNN को प्रशिक्षित करना\n",
|
||||
"\n",
|
||||
"हम RNN को टेक्स्ट जनरेट करने के लिए इस प्रकार प्रशिक्षित करेंगे। हर चरण में, हम `nchars` लंबाई के अक्षरों की एक श्रृंखला लेंगे और नेटवर्क से प्रत्येक इनपुट अक्षर के लिए अगला आउटपुट अक्षर जनरेट करने के लिए कहेंगे:\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"वास्तविक परिदृश्य के आधार पर, हम कुछ विशेष अक्षरों को भी शामिल करना चाह सकते हैं, जैसे *end-of-sequence* `<eos>`। हमारे मामले में, हम नेटवर्क को अंतहीन टेक्स्ट जनरेशन के लिए प्रशिक्षित करना चाहते हैं, इसलिए हम प्रत्येक श्रृंखला का आकार `nchars` टोकन के बराबर तय करेंगे। परिणामस्वरूप, प्रत्येक प्रशिक्षण उदाहरण में `nchars` इनपुट और `nchars` आउटपुट होंगे (जो इनपुट श्रृंखला को एक प्रतीक बाईं ओर शिफ्ट करके प्राप्त किए जाते हैं)। मिनीबैच में ऐसी कई श्रृंखलाएं शामिल होंगी।\n",
|
||||
"\n",
|
||||
"हम मिनीबैच को इस प्रकार जनरेट करेंगे कि प्रत्येक समाचार टेक्स्ट जिसकी लंबाई `l` है, से सभी संभावित इनपुट-आउटपुट संयोजन बनाएंगे (ऐसे `l-nchars` संयोजन होंगे)। ये एक मिनीबैच बनाएंगे, और प्रत्येक प्रशिक्षण चरण में मिनीबैच का आकार अलग-अलग होगा।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"(tensor([[ 0, 1, 2, ..., 28, 29, 30],\n",
|
||||
" [ 1, 2, 2, ..., 29, 30, 1],\n",
|
||||
" [ 2, 2, 3, ..., 30, 1, 16],\n",
|
||||
" ...,\n",
|
||||
" [20, 8, 21, ..., 1, 28, 1],\n",
|
||||
" [ 8, 21, 5, ..., 28, 1, 27],\n",
|
||||
" [21, 5, 8, ..., 1, 27, 16]]),\n",
|
||||
" tensor([[ 1, 2, 2, ..., 29, 30, 1],\n",
|
||||
" [ 2, 2, 3, ..., 30, 1, 16],\n",
|
||||
" [ 2, 3, 4, ..., 1, 16, 26],\n",
|
||||
" ...,\n",
|
||||
" [ 8, 21, 5, ..., 28, 1, 27],\n",
|
||||
" [21, 5, 8, ..., 1, 27, 16],\n",
|
||||
" [ 5, 8, 9, ..., 27, 16, 6]]))"
|
||||
]
|
||||
},
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"nchars = 100\n",
|
||||
"\n",
|
||||
"def get_batch(s,nchars=nchars):\n",
|
||||
" ins = torch.zeros(len(s)-nchars,nchars,dtype=torch.long,device=device)\n",
|
||||
" outs = torch.zeros(len(s)-nchars,nchars,dtype=torch.long,device=device)\n",
|
||||
" for i in range(len(s)-nchars):\n",
|
||||
" ins[i] = enc(s[i:i+nchars])\n",
|
||||
" outs[i] = enc(s[i+1:i+nchars+1])\n",
|
||||
" return ins,outs\n",
|
||||
"\n",
|
||||
"get_batch(train_dataset[0][1])"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"अब हम जनरेटर नेटवर्क को परिभाषित करते हैं। यह किसी भी पुनरावर्ती सेल पर आधारित हो सकता है जिसे हमने पिछले यूनिट में चर्चा की थी (सिंपल, LSTM या GRU)। हमारे उदाहरण में, हम LSTM का उपयोग करेंगे।\n",
|
||||
"\n",
|
||||
"चूंकि नेटवर्क अक्षरों को इनपुट के रूप में लेता है और शब्दावली का आकार काफी छोटा है, हमें एम्बेडिंग लेयर की आवश्यकता नहीं है। वन-हॉट-एनकोडेड इनपुट सीधे LSTM सेल में जा सकता है। हालांकि, क्योंकि हम अक्षरों के नंबर को इनपुट के रूप में पास करते हैं, हमें उन्हें LSTM में पास करने से पहले वन-हॉट-एनकोड करना होगा। यह `forward` पास के दौरान `one_hot` फ़ंक्शन को कॉल करके किया जाता है। आउटपुट एनकोडर एक लीनियर लेयर होगा जो हिडन स्टेट को वन-हॉट-एनकोडेड आउटपुट में बदल देगा।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 5,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"class LSTMGenerator(torch.nn.Module):\n",
|
||||
" def __init__(self, vocab_size, hidden_dim):\n",
|
||||
" super().__init__()\n",
|
||||
" self.rnn = torch.nn.LSTM(vocab_size,hidden_dim,batch_first=True)\n",
|
||||
" self.fc = torch.nn.Linear(hidden_dim, vocab_size)\n",
|
||||
"\n",
|
||||
" def forward(self, x, s=None):\n",
|
||||
" x = torch.nn.functional.one_hot(x,vocab_size).to(torch.float32)\n",
|
||||
" x,s = self.rnn(x,s)\n",
|
||||
" return self.fc(x),s"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"प्रशिक्षण के दौरान, हम उत्पन्न किए गए टेक्स्ट का नमूना लेना चाहते हैं। इसे करने के लिए, हम `generate` फ़ंक्शन को परिभाषित करेंगे, जो प्रारंभिक स्ट्रिंग `start` से शुरू करते हुए, लंबाई `size` का आउटपुट स्ट्रिंग उत्पन्न करेगा।\n",
|
||||
"\n",
|
||||
"इसका काम करने का तरीका निम्नलिखित है। सबसे पहले, हम पूरी प्रारंभिक स्ट्रिंग को नेटवर्क के माध्यम से पास करेंगे, और आउटपुट स्थिति `s` और अगला अनुमानित अक्षर `out` प्राप्त करेंगे। चूंकि `out` वन-हॉट एन्कोडेड होता है, हम `argmax` का उपयोग करके शब्दावली में अक्षर `nc` का इंडेक्स प्राप्त करेंगे, और `itos` का उपयोग करके वास्तविक अक्षर का पता लगाएंगे और इसे अक्षरों की परिणामी सूची `chars` में जोड़ देंगे। इस प्रक्रिया को एक अक्षर उत्पन्न करने के लिए `size` बार दोहराया जाता है, ताकि आवश्यक संख्या में अक्षर उत्पन्न किए जा सकें।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 8,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"def generate(net,size=100,start='today '):\n",
|
||||
" chars = list(start)\n",
|
||||
" out, s = net(enc(chars).view(1,-1).to(device))\n",
|
||||
" for i in range(size):\n",
|
||||
" nc = torch.argmax(out[0][-1])\n",
|
||||
" chars.append(vocab.get_itos()[nc])\n",
|
||||
" out, s = net(nc.view(1,-1),s)\n",
|
||||
" return ''.join(chars)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"अब चलिए प्रशिक्षण शुरू करते हैं! प्रशिक्षण लूप लगभग हमारे सभी पिछले उदाहरणों जैसा ही है, लेकिन सटीकता (accuracy) के बजाय हम हर 1000 epochs पर उत्पन्न किया गया नमूना टेक्स्ट प्रिंट करते हैं।\n",
|
||||
"\n",
|
||||
"विशेष ध्यान उस तरीके पर देना होगा जिससे हम हानि (loss) की गणना करते हैं। हमें हानि की गणना एक-हॉट-एन्कोडेड आउटपुट `out` और अपेक्षित टेक्स्ट `text_out` (जो कि कैरेक्टर इंडेक्स की सूची है) के आधार पर करनी होगी। सौभाग्य से, `cross_entropy` फ़ंक्शन अननॉर्मलाइज़्ड नेटवर्क आउटपुट को पहले तर्क के रूप में और क्लास नंबर को दूसरे तर्क के रूप में अपेक्षित करता है, जो कि हमारे पास पहले से ही है। यह मिनीबैच साइज पर स्वचालित औसत भी करता है।\n",
|
||||
"\n",
|
||||
"हम प्रशिक्षण को `samples_to_train` सैंपल्स तक सीमित करते हैं, ताकि बहुत अधिक समय न लगे। हम आपको प्रोत्साहित करते हैं कि आप प्रयोग करें और लंबे समय तक प्रशिक्षण आज़माएं, संभवतः कई epochs के लिए (ऐसे मामले में आपको इस कोड के चारों ओर एक और लूप बनाना होगा)।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 9,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Current loss = 4.398899078369141\n",
|
||||
"today sr sr sr sr sr sr sr sr sr sr sr sr sr sr sr sr sr sr sr sr sr sr sr sr sr sr sr sr sr sr sr sr sr s\n",
|
||||
"Current loss = 2.161320447921753\n",
|
||||
"today and to the tor to to the tor to to the tor to to the tor to to the tor to to the tor to to the tor t\n",
|
||||
"Current loss = 1.6722588539123535\n",
|
||||
"today and the court to the could to the could to the could to the could to the could to the could to the c\n",
|
||||
"Current loss = 2.423795223236084\n",
|
||||
"today and a second to the conternation of the conternation of the conternation of the conternation of the \n",
|
||||
"Current loss = 1.702607274055481\n",
|
||||
"today and the company to the company to the company to the company to the company to the company to the co\n",
|
||||
"Current loss = 1.692358136177063\n",
|
||||
"today and the company to the company to the company to the company to the company to the company to the co\n",
|
||||
"Current loss = 1.9722288846969604\n",
|
||||
"today and the control the control the control the control the control the control the control the control \n",
|
||||
"Current loss = 1.8705692291259766\n",
|
||||
"today and the second to the second to the second to the second to the second to the second to the second t\n",
|
||||
"Current loss = 1.7626899480819702\n",
|
||||
"today and a security and a security and a security and a security and a security and a security and a secu\n",
|
||||
"Current loss = 1.5574463605880737\n",
|
||||
"today and the company and the company and the company and the company and the company and the company and \n",
|
||||
"Current loss = 1.5620026588439941\n",
|
||||
"today and the be that the be the be that the be the be that the be the be that the be the be that the be t\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"net = LSTMGenerator(vocab_size,64).to(device)\n",
|
||||
"\n",
|
||||
"samples_to_train = 10000\n",
|
||||
"optimizer = torch.optim.Adam(net.parameters(),0.01)\n",
|
||||
"loss_fn = torch.nn.CrossEntropyLoss()\n",
|
||||
"net.train()\n",
|
||||
"for i,x in enumerate(train_dataset):\n",
|
||||
" # x[0] is class label, x[1] is text\n",
|
||||
" if len(x[1])-nchars<10:\n",
|
||||
" continue\n",
|
||||
" samples_to_train-=1\n",
|
||||
" if not samples_to_train: break\n",
|
||||
" text_in, text_out = get_batch(x[1])\n",
|
||||
" optimizer.zero_grad()\n",
|
||||
" out,s = net(text_in)\n",
|
||||
" loss = torch.nn.functional.cross_entropy(out.view(-1,vocab_size),text_out.flatten()) #cross_entropy(out,labels)\n",
|
||||
" loss.backward()\n",
|
||||
" optimizer.step()\n",
|
||||
" if i%1000==0:\n",
|
||||
" print(f\"Current loss = {loss.item()}\")\n",
|
||||
" print(generate(net))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"यह उदाहरण पहले से ही काफी अच्छा टेक्स्ट उत्पन्न करता है, लेकिन इसे कई तरीकों से और बेहतर बनाया जा सकता है:\n",
|
||||
"\n",
|
||||
"* **बेहतर मिनीबैच जनरेशन**। जिस तरीके से हमने ट्रेनिंग के लिए डेटा तैयार किया, वह था एक सैंपल से एक मिनीबैच बनाना। यह आदर्श नहीं है, क्योंकि मिनीबैच अलग-अलग आकार के होते हैं, और कुछ तो बनाए भी नहीं जा सकते, क्योंकि टेक्स्ट `nchars` से छोटा होता है। इसके अलावा, छोटे मिनीबैच GPU को पर्याप्त रूप से लोड नहीं करते। यह अधिक समझदारी होगी कि सभी सैंपल से एक बड़ा टेक्स्ट का हिस्सा लिया जाए, फिर सभी इनपुट-आउटपुट जोड़े बनाए जाएं, उन्हें शफल किया जाए, और समान आकार के मिनीबैच बनाए जाएं।\n",
|
||||
"\n",
|
||||
"* **मल्टीलायर LSTM**। 2 या 3 लेयर के LSTM सेल्स को आज़माना समझदारी होगी। जैसा कि हमने पिछले यूनिट में बताया था, LSTM की प्रत्येक लेयर टेक्स्ट से कुछ विशेष पैटर्न निकालती है, और कैरेक्टर-लेवल जनरेटर के मामले में हम उम्मीद कर सकते हैं कि निचली LSTM लेयर अक्षरों के समूह (syllables) को निकालने के लिए जिम्मेदार होगी, और ऊपरी लेयर शब्दों और शब्द संयोजनों के लिए। इसे आसानी से LSTM कंस्ट्रक्टर में लेयर की संख्या का पैरामीटर पास करके लागू किया जा सकता है।\n",
|
||||
"\n",
|
||||
"* आप **GRU यूनिट्स** के साथ भी प्रयोग कर सकते हैं और देख सकते हैं कि कौन से बेहतर प्रदर्शन करते हैं, और **अलग-अलग हिडन लेयर साइज** के साथ भी। बहुत बड़ी हिडन लेयर ओवरफिटिंग का कारण बन सकती है (जैसे कि नेटवर्क सटीक टेक्स्ट सीख लेगा), और छोटी साइज अच्छे परिणाम नहीं दे सकती।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## सॉफ्ट टेक्स्ट जनरेशन और टेम्परेचर\n",
|
||||
"\n",
|
||||
"`generate` की पिछली परिभाषा में, हम हमेशा उस कैरेक्टर को अगला कैरेक्टर चुनते थे जिसकी संभावना सबसे अधिक होती थी। इसका परिणाम यह होता था कि टेक्स्ट अक्सर बार-बार एक ही कैरेक्टर सीक्वेंस में \"चक्रित\" हो जाता था, जैसे इस उदाहरण में:\n",
|
||||
"```\n",
|
||||
"today of the second the company and a second the company ...\n",
|
||||
"```\n",
|
||||
"\n",
|
||||
"हालांकि, अगर हम अगले कैरेक्टर के लिए संभावना वितरण को देखें, तो यह हो सकता है कि कुछ उच्चतम संभावनाओं के बीच का अंतर बहुत बड़ा न हो, जैसे कि एक कैरेक्टर की संभावना 0.2 हो, और दूसरे की 0.19। उदाहरण के लिए, जब हम '*play*' सीक्वेंस में अगले कैरेक्टर की तलाश कर रहे हों, तो अगला कैरेक्टर स्पेस या **e** (जैसे शब्द *player* में) दोनों ही हो सकते हैं।\n",
|
||||
"\n",
|
||||
"इससे यह निष्कर्ष निकलता है कि हमेशा उच्चतम संभावना वाले कैरेक्टर को चुनना \"न्यायसंगत\" नहीं है, क्योंकि दूसरे उच्चतम को चुनने से भी अर्थपूर्ण टेक्स्ट बन सकता है। यह अधिक समझदारी होगी कि नेटवर्क आउटपुट द्वारा दी गई संभावना वितरण से **सैंपल** करके कैरेक्टर चुने जाएं।\n",
|
||||
"\n",
|
||||
"यह सैंपलिंग `multinomial` फंक्शन का उपयोग करके की जा सकती है, जो तथाकथित **मल्टिनोमियल वितरण** को लागू करता है। एक फंक्शन जो इस **सॉफ्ट** टेक्स्ट जनरेशन को लागू करता है, नीचे परिभाषित है:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 10,
|
||||
"metadata": {
|
||||
"scrolled": true
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"--- Temperature = 0.3\n",
|
||||
"Today and a company and complete an all the land the restrational the as a security and has provers the pay to and a report and the computer in the stand has filities and working the law the stations for a company and with the company and the final the first company and refight of the state and and workin\n",
|
||||
"\n",
|
||||
"--- Temperature = 0.8\n",
|
||||
"Today he oniis its first to Aus bomblaties the marmation a to manan boogot that pirate assaid a relaid their that goverfin the the Cappets Ecrotional Assonia Cition targets it annight the w scyments Blamity #39;s TVeer Diercheg Reserals fran envyuil that of ster said access what succers of Dour-provelith\n",
|
||||
"\n",
|
||||
"--- Temperature = 1.0\n",
|
||||
"Today holy they a 11 will meda a toket subsuaties, engins for Chanos, they's has stainger past to opening orital his thempting new Nattona was al innerforder advan-than #36;s night year his religuled talitatian what the but with Wednesday to Justment will wemen of Mark CCC Camp as Timed Nae wome a leaders\n",
|
||||
"\n",
|
||||
"--- Temperature = 1.3\n",
|
||||
"Today gpone 2.5 fech atcusion poor cocles toparsdorM.cht Line Pamage put 43 his calt lowed to the book, that has authh-the silia rruch ailing to'ory andhes beutirsimi- Aefffive heading offil an auf eacklets is charged evis, Gunymy oy) Mony has it after-sloythyor loveId out filme, the Natabl -Najuntaxiggs \n",
|
||||
"\n",
|
||||
"--- Temperature = 1.8\n",
|
||||
"Today plary, P.slan chly\\401 mardregationly #39;t 8.1Mide) closes ,filtcon alfly playin roven!\\grea.-QFBEP: Iss onfarchQ/itilia CCf Zivesigntwasta orce.-Peul-aw.uicrin of fuglinfsut aftaningwo, MIEX awayew Aice Woiduar Corvagiugge oppo esig ThusBratourid canthly-RyI.co lagitems\\eexciaishes.conBabntusmor I\n",
|
||||
"\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"def generate_soft(net,size=100,start='today ',temperature=1.0):\n",
|
||||
" chars = list(start)\n",
|
||||
" out, s = net(enc(chars).view(1,-1).to(device))\n",
|
||||
" for i in range(size):\n",
|
||||
" #nc = torch.argmax(out[0][-1])\n",
|
||||
" out_dist = out[0][-1].div(temperature).exp()\n",
|
||||
" nc = torch.multinomial(out_dist,1)[0]\n",
|
||||
" chars.append(vocab.get_itos()[nc])\n",
|
||||
" out, s = net(nc.view(1,-1),s)\n",
|
||||
" return ''.join(chars)\n",
|
||||
" \n",
|
||||
"for i in [0.3,0.8,1.0,1.3,1.8]:\n",
|
||||
" print(f\"--- Temperature = {i}\\n{generate_soft(net,size=300,start='Today ',temperature=i)}\\n\")"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"हमने एक और पैरामीटर **तापमान** पेश किया है, जिसका उपयोग यह संकेत देने के लिए किया जाता है कि हमें उच्चतम संभावना से कितनी दृढ़ता से चिपकना चाहिए। यदि तापमान 1.0 है, तो हम निष्पक्ष बहुपद नमूना लेते हैं, और जब तापमान अनंत तक जाता है - सभी संभावनाएँ समान हो जाती हैं, और हम अगला वर्ण यादृच्छिक रूप से चुनते हैं। नीचे दिए गए उदाहरण में हम देख सकते हैं कि जब हम तापमान को बहुत अधिक बढ़ाते हैं तो पाठ अर्थहीन हो जाता है, और जब यह 0 के करीब होता है तो यह \"चक्रित\" कठोर-जनित पाठ जैसा दिखता है।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"\n---\n\n**अस्वीकरण**: \nयह दस्तावेज़ AI अनुवाद सेवा [Co-op Translator](https://github.com/Azure/co-op-translator) का उपयोग करके अनुवादित किया गया है। जबकि हम सटीकता सुनिश्चित करने का प्रयास करते हैं, कृपया ध्यान दें कि स्वचालित अनुवाद में त्रुटियां या अशुद्धियां हो सकती हैं। मूल भाषा में उपलब्ध मूल दस्तावेज़ को आधिकारिक स्रोत माना जाना चाहिए। महत्वपूर्ण जानकारी के लिए, पेशेवर मानव अनुवाद की सिफारिश की जाती है। इस अनुवाद के उपयोग से उत्पन्न किसी भी गलतफहमी या गलत व्याख्या के लिए हम जिम्मेदार नहीं हैं।\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
"interpreter": {
|
||||
"hash": "16af2a8bbb083ea23e5e41c7f5787656b2ce26968575d8763f2c4b17f9cd711f"
|
||||
},
|
||||
"kernelspec": {
|
||||
"display_name": "Python 3.8.12 ('py38')",
|
||||
"language": "python",
|
||||
"name": "python3"
|
||||
},
|
||||
"language_info": {
|
||||
"codemirror_mode": {
|
||||
"name": "ipython",
|
||||
"version": 3
|
||||
},
|
||||
"file_extension": ".py",
|
||||
"mimetype": "text/x-python",
|
||||
"name": "python",
|
||||
"nbconvert_exporter": "python",
|
||||
"pygments_lexer": "ipython3",
|
||||
"version": "3.8.12"
|
||||
},
|
||||
"coopTranslator": {
|
||||
"original_hash": "7673cd150d96c74c6d6011460094efb4",
|
||||
"translation_date": "2025-08-31T15:13:47+00:00",
|
||||
"source_file": "lessons/5-NLP/17-GenerativeNetworks/GenerativePyTorch.ipynb",
|
||||
"language_code": "hi"
|
||||
}
|
||||
},
|
||||
"nbformat": 4,
|
||||
"nbformat_minor": 4
|
||||
}
|
||||
|
|
@ -0,0 +1,495 @@
|
|||
{
|
||||
"cells": [
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"# जनरेटिव नेटवर्क्स\n",
|
||||
"\n",
|
||||
"रिकरंट न्यूरल नेटवर्क्स (RNNs) और उनके गेटेड सेल वेरिएंट्स जैसे लॉन्ग शॉर्ट टर्म मेमोरी सेल्स (LSTMs) और गेटेड रिकारंट यूनिट्स (GRUs) ने भाषा मॉडलिंग के लिए एक तंत्र प्रदान किया है, यानी वे शब्दों के क्रम को सीख सकते हैं और अनुक्रम में अगले शब्द की भविष्यवाणी कर सकते हैं। यह हमें RNNs का उपयोग **जनरेटिव कार्यों** के लिए करने की अनुमति देता है, जैसे साधारण टेक्स्ट जनरेशन, मशीन ट्रांसलेशन, और यहां तक कि इमेज कैप्शनिंग।\n",
|
||||
"\n",
|
||||
"पिछली यूनिट में हमने जिस RNN आर्किटेक्चर पर चर्चा की थी, उसमें प्रत्येक RNN यूनिट ने अगले हिडन स्टेट को आउटपुट के रूप में उत्पन्न किया। हालांकि, हम प्रत्येक रिकारंट यूनिट में एक और आउटपुट जोड़ सकते हैं, जो हमें एक **अनुक्रम** आउटपुट करने की अनुमति देगा (जो मूल अनुक्रम की लंबाई के बराबर होगा)। इसके अलावा, हम ऐसे RNN यूनिट्स का उपयोग कर सकते हैं जो प्रत्येक चरण में इनपुट स्वीकार नहीं करते, बल्कि केवल एक प्रारंभिक स्टेट वेक्टर लेते हैं और फिर आउटपुट का एक अनुक्रम उत्पन्न करते हैं।\n",
|
||||
"\n",
|
||||
"इस नोटबुक में, हम सरल जनरेटिव मॉडलों पर ध्यान केंद्रित करेंगे जो हमें टेक्स्ट जनरेट करने में मदद करते हैं। सरलता के लिए, चलिए एक **कैरेक्टर-लेवल नेटवर्क** बनाते हैं, जो अक्षर दर अक्षर टेक्स्ट जनरेट करता है। प्रशिक्षण के दौरान, हमें कुछ टेक्स्ट कॉर्पस लेना होगा और उसे अक्षर अनुक्रमों में विभाजित करना होगा।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 1,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"import tensorflow as tf\n",
|
||||
"from tensorflow import keras\n",
|
||||
"import tensorflow_datasets as tfds\n",
|
||||
"import numpy as np\n",
|
||||
"\n",
|
||||
"ds_train, ds_test = tfds.load('ag_news_subset').values()"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## वर्णमाला शब्दावली बनाना\n",
|
||||
"\n",
|
||||
"वर्ण-स्तरीय जनरेटिव नेटवर्क बनाने के लिए, हमें टेक्स्ट को शब्दों के बजाय व्यक्तिगत अक्षरों में विभाजित करना होगा। `TextVectorization` लेयर, जिसे हमने पहले उपयोग किया था, ऐसा नहीं कर सकती, इसलिए हमारे पास दो विकल्प हैं:\n",
|
||||
"\n",
|
||||
"* टेक्स्ट को मैन्युअली लोड करें और 'हाथ से' टोकनाइज़ेशन करें, जैसा कि [इस आधिकारिक Keras उदाहरण](https://keras.io/examples/generative/lstm_character_level_text_generation/) में दिखाया गया है।\n",
|
||||
"* वर्ण-स्तरीय टोकनाइज़ेशन के लिए `Tokenizer` क्लास का उपयोग करें।\n",
|
||||
"\n",
|
||||
"हम दूसरे विकल्प को चुनेंगे। `Tokenizer` का उपयोग शब्दों में टोकनाइज़ करने के लिए भी किया जा सकता है, इसलिए कोई आसानी से वर्ण-स्तरीय से शब्द-स्तरीय टोकनाइज़ेशन में स्विच कर सकता है।\n",
|
||||
"\n",
|
||||
"वर्ण-स्तरीय टोकनाइज़ेशन करने के लिए, हमें `char_level=True` पैरामीटर पास करना होगा:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 2,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"def extract_text(x):\n",
|
||||
" return x['title']+' '+x['description']\n",
|
||||
"\n",
|
||||
"def tupelize(x):\n",
|
||||
" return (extract_text(x),x['label'])\n",
|
||||
"\n",
|
||||
"tokenizer = keras.preprocessing.text.Tokenizer(char_level=True,lower=False)\n",
|
||||
"tokenizer.fit_on_texts([x['title'].numpy().decode('utf-8') for x in ds_train])"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"हम एक विशेष टोकन का उपयोग करना चाहते हैं ताकि **अनुक्रम के अंत** को दर्शाया जा सके, जिसे हम `<eos>` कहेंगे। आइए इसे मैन्युअल रूप से शब्दावली में जोड़ें:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 3,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"eos_token = len(tokenizer.word_index)+1\n",
|
||||
"tokenizer.word_index['<eos>'] = eos_token\n",
|
||||
"\n",
|
||||
"vocab_size = eos_token + 1"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": []
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"[[48, 2, 10, 10, 5, 44, 1, 25, 5, 8, 10, 13, 78]]"
|
||||
]
|
||||
},
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"tokenizer.texts_to_sequences(['Hello, world!'])"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## शीर्षक उत्पन्न करने के लिए एक जनरेटिव RNN को प्रशिक्षित करना\n",
|
||||
"\n",
|
||||
"हम RNN को समाचार शीर्षक उत्पन्न करने के लिए इस प्रकार प्रशिक्षित करेंगे। प्रत्येक चरण में, हम एक शीर्षक लेंगे, जिसे RNN में दिया जाएगा, और प्रत्येक इनपुट अक्षर के लिए हम नेटवर्क से अगला आउटपुट अक्षर उत्पन्न करने के लिए कहेंगे:\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"हमारे अनुक्रम के अंतिम अक्षर के लिए, हम नेटवर्क से `<eos>` टोकन उत्पन्न करने के लिए कहेंगे।\n",
|
||||
"\n",
|
||||
"यहां उपयोग किए जा रहे जनरेटिव RNN का मुख्य अंतर यह है कि हम RNN के प्रत्येक चरण से आउटपुट लेंगे, न कि केवल अंतिम सेल से। इसे RNN सेल में `return_sequences` पैरामीटर निर्दिष्ट करके प्राप्त किया जा सकता है।\n",
|
||||
"\n",
|
||||
"इस प्रकार, प्रशिक्षण के दौरान, नेटवर्क में इनपुट कुछ लंबाई के एन्कोडेड अक्षरों का अनुक्रम होगा, और आउटपुट उसी लंबाई का अनुक्रम होगा, लेकिन एक तत्व द्वारा शिफ्ट किया गया और `<eos>` से समाप्त किया गया। मिनीबैच में कई ऐसे अनुक्रम शामिल होंगे, और हमें सभी अनुक्रमों को संरेखित करने के लिए **पैडिंग** का उपयोग करना होगा।\n",
|
||||
"\n",
|
||||
"आइए ऐसी फ़ंक्शन बनाएं जो हमारे लिए डेटासेट को परिवर्तित करें। क्योंकि हम मिनीबैच स्तर पर अनुक्रमों को पैड करना चाहते हैं, हम पहले `.batch()` कॉल करके डेटासेट को बैच करेंगे, और फिर इसे `map` करेंगे ताकि परिवर्तन किया जा सके। इसलिए, परिवर्तन फ़ंक्शन पूरे मिनीबैच को एक पैरामीटर के रूप में लेगा:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 5,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"def title_batch(x):\n",
|
||||
" x = [t.numpy().decode('utf-8') for t in x]\n",
|
||||
" z = tokenizer.texts_to_sequences(x)\n",
|
||||
" z = tf.keras.preprocessing.sequence.pad_sequences(z)\n",
|
||||
" return tf.one_hot(z,vocab_size), tf.one_hot(tf.concat([z[:,1:],tf.constant(eos_token,shape=(len(z),1))],axis=1),vocab_size)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"कुछ महत्वपूर्ण बातें जो हम यहाँ करते हैं:\n",
|
||||
"* सबसे पहले हम स्ट्रिंग टेंसर से वास्तविक टेक्स्ट को निकालते हैं\n",
|
||||
"* `text_to_sequences` स्ट्रिंग्स की सूची को पूर्णांक टेंसर की सूची में बदल देता है\n",
|
||||
"* `pad_sequences` उन टेंसर को उनकी अधिकतम लंबाई तक पैड करता है\n",
|
||||
"* अंत में हम सभी अक्षरों को वन-हॉट एन्कोड करते हैं, और साथ ही शिफ्टिंग और `<eos>` जोड़ने का काम भी करते हैं। हम जल्द ही देखेंगे कि हमें वन-हॉट-एन्कोडेड अक्षरों की आवश्यकता क्यों है\n",
|
||||
"\n",
|
||||
"हालांकि, यह फ़ंक्शन **Pythonic** है, यानी इसे Tensorflow के कम्प्यूटेशनल ग्राफ में स्वचालित रूप से अनुवादित नहीं किया जा सकता। अगर हम इस फ़ंक्शन को सीधे `Dataset.map` फ़ंक्शन में उपयोग करने की कोशिश करेंगे, तो हमें त्रुटियाँ मिलेंगी। हमें इस Pythonic कॉल को `py_function` रैपर का उपयोग करके संलग्न करना होगा:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 6,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"def title_batch_fn(x):\n",
|
||||
" x = x['title']\n",
|
||||
" a,b = tf.py_function(title_batch,inp=[x],Tout=(tf.float32,tf.float32))\n",
|
||||
" return a,b"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"> **Note**: पायथनिक और टेन्सरफ्लो ट्रांसफॉर्मेशन फंक्शन्स के बीच अंतर करना थोड़ा जटिल लग सकता है, और आप सोच सकते हैं कि हम डेटा सेट को `fit` में पास करने से पहले मानक पायथन फंक्शन्स का उपयोग करके क्यों नहीं बदलते। हालांकि यह निश्चित रूप से किया जा सकता है, `Dataset.map` का उपयोग करने का एक बड़ा लाभ है, क्योंकि डेटा ट्रांसफॉर्मेशन पाइपलाइन टेन्सरफ्लो कम्प्यूटेशनल ग्राफ का उपयोग करके निष्पादित होती है, जो GPU कम्प्यूटेशन का लाभ उठाती है और CPU/GPU के बीच डेटा पास करने की आवश्यकता को कम करती है।\n",
|
||||
"\n",
|
||||
"अब हम अपना जनरेटर नेटवर्क बना सकते हैं और प्रशिक्षण शुरू कर सकते हैं। इसे किसी भी पुनरावर्ती सेल पर आधारित किया जा सकता है जिसे हमने पिछले यूनिट में चर्चा की थी (सिंपल, LSTM या GRU)। हमारे उदाहरण में हम LSTM का उपयोग करेंगे।\n",
|
||||
"\n",
|
||||
"चूंकि नेटवर्क इनपुट के रूप में अक्षरों को लेता है, और शब्दावली का आकार काफी छोटा है, हमें एम्बेडिंग लेयर की आवश्यकता नहीं है। वन-हॉट-एनकोडेड इनपुट सीधे LSTM सेल में जा सकता है। आउटपुट लेयर एक `Dense` क्लासिफायर होगी जो LSTM आउटपुट को वन-हॉट-एनकोडेड टोकन नंबरों में बदल देगी।\n",
|
||||
"\n",
|
||||
"इसके अलावा, चूंकि हम वेरिएबल-लेंथ सीक्वेंस के साथ काम कर रहे हैं, हम `Masking` लेयर का उपयोग कर सकते हैं ताकि एक मास्क बनाया जा सके जो स्ट्रिंग के पैडेड हिस्से को अनदेखा कर दे। यह सख्ती से आवश्यक नहीं है, क्योंकि हम `<eos>` टोकन से आगे की चीजों में बहुत अधिक रुचि नहीं रखते हैं, लेकिन हम इस लेयर प्रकार के साथ कुछ अनुभव प्राप्त करने के लिए इसका उपयोग करेंगे। `input_shape` `(None, vocab_size)` होगा, जहां `None` वेरिएबल लंबाई की सीक्वेंस को इंगित करता है, और आउटपुट आकार भी `(None, vocab_size)` होगा, जैसा कि आप `summary` से देख सकते हैं:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 7,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Model: \"sequential\"\n",
|
||||
"_________________________________________________________________\n",
|
||||
"Layer (type) Output Shape Param # \n",
|
||||
"=================================================================\n",
|
||||
"masking (Masking) (None, None, 84) 0 \n",
|
||||
"_________________________________________________________________\n",
|
||||
"lstm (LSTM) (None, None, 128) 109056 \n",
|
||||
"_________________________________________________________________\n",
|
||||
"dense (Dense) (None, None, 84) 10836 \n",
|
||||
"=================================================================\n",
|
||||
"Total params: 119,892\n",
|
||||
"Trainable params: 119,892\n",
|
||||
"Non-trainable params: 0\n",
|
||||
"_________________________________________________________________\n",
|
||||
"15000/15000 [==============================] - 229s 15ms/step - loss: 1.5385\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"<tensorflow.python.keras.callbacks.History at 0x7fa40c1245e0>"
|
||||
]
|
||||
},
|
||||
"execution_count": 7,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"model = keras.models.Sequential([\n",
|
||||
" keras.layers.Masking(input_shape=(None,vocab_size)),\n",
|
||||
" keras.layers.LSTM(128,return_sequences=True),\n",
|
||||
" keras.layers.Dense(vocab_size,activation='softmax')\n",
|
||||
"])\n",
|
||||
"\n",
|
||||
"model.summary()\n",
|
||||
"model.compile(loss='categorical_crossentropy')\n",
|
||||
"\n",
|
||||
"model.fit(ds_train.batch(8).map(title_batch_fn))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## आउटपुट उत्पन्न करना\n",
|
||||
"\n",
|
||||
"अब जब हमने मॉडल को प्रशिक्षित कर लिया है, तो हम इसे कुछ आउटपुट उत्पन्न करने के लिए उपयोग करना चाहते हैं। सबसे पहले, हमें टोकन नंबरों की एक श्रृंखला द्वारा दर्शाए गए टेक्स्ट को डिकोड करने का एक तरीका चाहिए। इसके लिए, हम `tokenizer.sequences_to_texts` फ़ंक्शन का उपयोग कर सकते हैं; हालांकि, यह कैरेक्टर-लेवल टोकनाइजेशन के साथ अच्छी तरह से काम नहीं करता। इसलिए, हम टोकनाइज़र से टोकन की एक डिक्शनरी (जिसे `word_index` कहा जाता है) लेंगे, एक रिवर्स मैप बनाएंगे, और अपना खुद का डिकोडिंग फ़ंक्शन लिखेंगे:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 10,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"reverse_map = {val:key for key, val in tokenizer.word_index.items()}\n",
|
||||
"\n",
|
||||
"def decode(x):\n",
|
||||
" return ''.join([reverse_map[t] for t in x])"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"अब, चलिए जनरेशन करते हैं। हम किसी स्ट्रिंग `start` से शुरुआत करेंगे, इसे एक अनुक्रम `inp` में एन्कोड करेंगे, और फिर हर चरण में हम अपने नेटवर्क को कॉल करेंगे ताकि अगला कैरेक्टर अनुमानित किया जा सके।\n",
|
||||
"\n",
|
||||
"नेटवर्क का आउटपुट `out` एक वेक्टर होता है जिसमें `vocab_size` तत्व होते हैं, जो प्रत्येक टोकन की संभावनाओं का प्रतिनिधित्व करते हैं। हम `argmax` का उपयोग करके सबसे संभावित टोकन नंबर ढूंढ सकते हैं। इसके बाद हम इस कैरेक्टर को जनरेट किए गए टोकन्स की सूची में जोड़ते हैं और जनरेशन की प्रक्रिया जारी रखते हैं। इस प्रक्रिया में एक कैरेक्टर जनरेट करने की प्रक्रिया को `size` बार दोहराया जाता है ताकि आवश्यक संख्या में कैरेक्टर्स जनरेट किए जा सकें, और हम जल्दी समाप्त कर देते हैं जब `eos_token` मिल जाता है।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 12,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"'Today #39;s lead to strike for the strike for the strike for the strike (AFP)'"
|
||||
]
|
||||
},
|
||||
"execution_count": 12,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"def generate(model,size=100,start='Today '):\n",
|
||||
" inp = tokenizer.texts_to_sequences([start])[0]\n",
|
||||
" chars = inp\n",
|
||||
" for i in range(size):\n",
|
||||
" out = model(tf.expand_dims(tf.one_hot(inp,vocab_size),0))[0][-1]\n",
|
||||
" nc = tf.argmax(out)\n",
|
||||
" if nc==eos_token:\n",
|
||||
" break\n",
|
||||
" chars.append(nc.numpy())\n",
|
||||
" inp = inp+[nc]\n",
|
||||
" return decode(chars)\n",
|
||||
" \n",
|
||||
"generate(model)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## प्रशिक्षण के दौरान आउटपुट का नमूना लेना\n",
|
||||
"\n",
|
||||
"चूंकि हमारे पास *सटीकता* जैसे कोई उपयोगी मेट्रिक्स नहीं हैं, इसलिए यह देखने का एकमात्र तरीका कि हमारा मॉडल बेहतर हो रहा है, **प्रशिक्षण के दौरान उत्पन्न स्ट्रिंग का नमूना लेना** है। इसे करने के लिए, हम **कॉलबैक्स** का उपयोग करेंगे, यानी ऐसी फ़ंक्शन्स जिन्हें हम `fit` फ़ंक्शन में पास कर सकते हैं, और जो प्रशिक्षण के दौरान समय-समय पर कॉल की जाएंगी।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 13,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Epoch 1/3\n",
|
||||
"15000/15000 [==============================] - 226s 15ms/step - loss: 1.2703\n",
|
||||
"Today #39;s a lead in the company for the strike\n",
|
||||
"Epoch 2/3\n",
|
||||
"15000/15000 [==============================] - 227s 15ms/step - loss: 1.2057\n",
|
||||
"Today #39;s the Market Service on Security Start (AP)\n",
|
||||
"Epoch 3/3\n",
|
||||
"15000/15000 [==============================] - 226s 15ms/step - loss: 1.1752\n",
|
||||
"Today #39;s a line on the strike to start for the start\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"<tensorflow.python.keras.callbacks.History at 0x7fa40c74e3d0>"
|
||||
]
|
||||
},
|
||||
"execution_count": 13,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"sampling_callback = keras.callbacks.LambdaCallback(\n",
|
||||
" on_epoch_end = lambda batch, logs: print(generate(model))\n",
|
||||
")\n",
|
||||
"\n",
|
||||
"model.fit(ds_train.batch(8).map(title_batch_fn),callbacks=[sampling_callback],epochs=3)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"यह उदाहरण पहले से ही काफी अच्छा पाठ उत्पन्न करता है, लेकिन इसे कई तरीकों से और बेहतर बनाया जा सकता है:\n",
|
||||
"\n",
|
||||
"* **अधिक पाठ**। हमने अपने कार्य के लिए केवल शीर्षकों का उपयोग किया है, लेकिन आप पूरे पाठ के साथ प्रयोग करना चाह सकते हैं। याद रखें कि RNNs लंबे अनुक्रमों को संभालने में बहुत अच्छे नहीं होते हैं, इसलिए उन्हें छोटे वाक्यों में विभाजित करना या हमेशा किसी पूर्वनिर्धारित मान `num_chars` (जैसे, 256) की निश्चित अनुक्रम लंबाई पर प्रशिक्षण देना समझदारी हो सकती है। आप ऊपर दिए गए उदाहरण को ऐसी संरचना में बदलने की कोशिश कर सकते हैं, [आधिकारिक Keras ट्यूटोरियल](https://keras.io/examples/generative/lstm_character_level_text_generation/) को प्रेरणा के रूप में उपयोग करते हुए।\n",
|
||||
"\n",
|
||||
"* **मल्टीलेयर LSTM**। LSTM कोशिकाओं की 2 या 3 परतों को आज़माना समझदारी हो सकता है। जैसा कि हमने पिछले यूनिट में उल्लेख किया था, LSTM की प्रत्येक परत पाठ से कुछ पैटर्न निकालती है, और कैरेक्टर-लेवल जनरेटर के मामले में हम उम्मीद कर सकते हैं कि निचली LSTM परत अक्षरों को निकालने के लिए जिम्मेदार होगी, और ऊपरी परतें - शब्द और शब्द संयोजन के लिए। इसे LSTM कंस्ट्रक्टर में परतों की संख्या का पैरामीटर पास करके आसानी से लागू किया जा सकता है।\n",
|
||||
"\n",
|
||||
"* आप **GRU यूनिट्स** के साथ भी प्रयोग करना चाह सकते हैं और देख सकते हैं कि कौन सा बेहतर प्रदर्शन करता है, और **विभिन्न छिपी परत के आकार** के साथ भी। बहुत बड़ी छिपी परत ओवरफिटिंग का कारण बन सकती है (जैसे, नेटवर्क सटीक पाठ सीख लेगा), और छोटा आकार अच्छा परिणाम उत्पन्न नहीं कर सकता।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## सॉफ्ट टेक्स्ट जनरेशन और टेम्परेचर\n",
|
||||
"\n",
|
||||
"`generate` की पिछली परिभाषा में, हम हमेशा उस अक्षर को अगला अक्षर चुनते थे जिसकी संभावना सबसे अधिक होती थी। इसका परिणाम यह होता था कि टेक्स्ट अक्सर बार-बार एक ही अक्षर अनुक्रम में \"चक्रित\" हो जाता था, जैसे इस उदाहरण में:\n",
|
||||
"```\n",
|
||||
"today of the second the company and a second the company ...\n",
|
||||
"```\n",
|
||||
"\n",
|
||||
"हालांकि, अगर हम अगले अक्षर के लिए संभावना वितरण को देखें, तो यह हो सकता है कि कुछ उच्चतम संभावनाओं के बीच का अंतर बहुत बड़ा न हो, जैसे कि एक अक्षर की संभावना 0.2 हो, और दूसरे की 0.19। उदाहरण के लिए, जब अनुक्रम '*play*' में अगले अक्षर की तलाश की जाती है, तो अगला अक्षर समान रूप से स्पेस या **e** (जैसे शब्द *player* में) हो सकता है।\n",
|
||||
"\n",
|
||||
"इससे यह निष्कर्ष निकलता है कि हमेशा उच्चतम संभावना वाले अक्षर को चुनना \"न्यायसंगत\" नहीं है, क्योंकि दूसरे उच्चतम को चुनना भी हमें सार्थक टेक्स्ट की ओर ले जा सकता है। यह अधिक समझदारी होगी कि नेटवर्क आउटपुट द्वारा दी गई संभावना वितरण से अक्षरों को **सैंपल** किया जाए।\n",
|
||||
"\n",
|
||||
"यह सैंपलिंग `np.multinomial` फंक्शन का उपयोग करके की जा सकती है, जो तथाकथित **मल्टिनोमियल वितरण** को लागू करता है। एक फंक्शन जो इस **सॉफ्ट** टेक्स्ट जनरेशन को लागू करता है, नीचे परिभाषित है:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 33,
|
||||
"metadata": {
|
||||
"scrolled": true
|
||||
},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"\n",
|
||||
"--- Temperature = 0.3\n",
|
||||
"Today #39;s strike #39; to start at the store return\n",
|
||||
"On Sunday PO to Be Data Profit Up (Reuters)\n",
|
||||
"Moscow, SP wins straight to the Microsoft #39;s control of the space start\n",
|
||||
"President olding of the blast start for the strike to pay <b>...</b>\n",
|
||||
"Little red riding hood ficed to the spam countered in European <b>...</b>\n",
|
||||
"\n",
|
||||
"--- Temperature = 0.8\n",
|
||||
"Today countie strikes ryder missile faces food market blut\n",
|
||||
"On Sunday collores lose-toppy of sale of Bullment in <b>...</b>\n",
|
||||
"Moscow, IBM Diffeiting in Afghan Software Hotels (Reuters)\n",
|
||||
"President Ol Luster for Profit Peaced Raised (AP)\n",
|
||||
"Little red riding hood dace on depart talks #39; bank up\n",
|
||||
"\n",
|
||||
"--- Temperature = 1.0\n",
|
||||
"Today wits House buiting debate fixes #39; supervice stake again\n",
|
||||
"On Sunday arling digital poaching In for level\n",
|
||||
"Moscow, DS Up 7, Top Proble Protest Caprey Mamarian Strike\n",
|
||||
"President teps help of roubler stepted lessabul-Dhalitics (AFP)\n",
|
||||
"Little red riding hood signs on cash in Carter-youb\n",
|
||||
"\n",
|
||||
"--- Temperature = 1.3\n",
|
||||
"Today wits flawer ro, pSIA figat's co DroftwavesIs Talo up\n",
|
||||
"On Sunday hround elitwing wint EU Powerburlinetien\n",
|
||||
"Moscow, Bazz #39;s sentries olymen winnelds' next for Olympite Huc?\n",
|
||||
"President lost securitys from power Elections in Smiltrials\n",
|
||||
"Little red riding hood vides profit, exponituity, profitmainalist-at said listers\n",
|
||||
"\n",
|
||||
"--- Temperature = 1.8\n",
|
||||
"Today #39;It: He deat: N.KA Asside\n",
|
||||
"On Sunday i arry Par aldeup patient Wo stele1\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"ename": "KeyError",
|
||||
"evalue": "0",
|
||||
"output_type": "error",
|
||||
"traceback": [
|
||||
"\u001b[0;31m---------------------------------------------------------------------------\u001b[0m",
|
||||
"\u001b[0;31mKeyError\u001b[0m Traceback (most recent call last)",
|
||||
"\u001b[0;32m<ipython-input-33-db32367a0feb>\u001b[0m in \u001b[0;36m<module>\u001b[0;34m\u001b[0m\n\u001b[1;32m 18\u001b[0m \u001b[0mprint\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0;34mf\"\\n--- Temperature = {i}\"\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[1;32m 19\u001b[0m \u001b[0;32mfor\u001b[0m \u001b[0mj\u001b[0m \u001b[0;32min\u001b[0m \u001b[0mrange\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0;36m5\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m:\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0;32m---> 20\u001b[0;31m \u001b[0mprint\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0mgenerate_soft\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0mmodel\u001b[0m\u001b[0;34m,\u001b[0m\u001b[0msize\u001b[0m\u001b[0;34m=\u001b[0m\u001b[0;36m300\u001b[0m\u001b[0;34m,\u001b[0m\u001b[0mstart\u001b[0m\u001b[0;34m=\u001b[0m\u001b[0mwords\u001b[0m\u001b[0;34m[\u001b[0m\u001b[0mj\u001b[0m\u001b[0;34m]\u001b[0m\u001b[0;34m,\u001b[0m\u001b[0mtemperature\u001b[0m\u001b[0;34m=\u001b[0m\u001b[0mi\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0m",
|
||||
"\u001b[0;32m<ipython-input-33-db32367a0feb>\u001b[0m in \u001b[0;36mgenerate_soft\u001b[0;34m(model, size, start, temperature)\u001b[0m\n\u001b[1;32m 11\u001b[0m \u001b[0mchars\u001b[0m\u001b[0;34m.\u001b[0m\u001b[0mappend\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0mnc\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[1;32m 12\u001b[0m \u001b[0minp\u001b[0m \u001b[0;34m=\u001b[0m \u001b[0minp\u001b[0m\u001b[0;34m+\u001b[0m\u001b[0;34m[\u001b[0m\u001b[0mnc\u001b[0m\u001b[0;34m]\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0;32m---> 13\u001b[0;31m \u001b[0;32mreturn\u001b[0m \u001b[0mdecode\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0mchars\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0m\u001b[1;32m 14\u001b[0m \u001b[0;34m\u001b[0m\u001b[0m\n\u001b[1;32m 15\u001b[0m \u001b[0mwords\u001b[0m \u001b[0;34m=\u001b[0m \u001b[0;34m[\u001b[0m\u001b[0;34m'Today '\u001b[0m\u001b[0;34m,\u001b[0m\u001b[0;34m'On Sunday '\u001b[0m\u001b[0;34m,\u001b[0m\u001b[0;34m'Moscow, '\u001b[0m\u001b[0;34m,\u001b[0m\u001b[0;34m'President '\u001b[0m\u001b[0;34m,\u001b[0m\u001b[0;34m'Little red riding hood '\u001b[0m\u001b[0;34m]\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n",
|
||||
"\u001b[0;32m<ipython-input-10-3f5fa6130b1d>\u001b[0m in \u001b[0;36mdecode\u001b[0;34m(x)\u001b[0m\n\u001b[1;32m 2\u001b[0m \u001b[0;34m\u001b[0m\u001b[0m\n\u001b[1;32m 3\u001b[0m \u001b[0;32mdef\u001b[0m \u001b[0mdecode\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0mx\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m:\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0;32m----> 4\u001b[0;31m \u001b[0;32mreturn\u001b[0m \u001b[0;34m''\u001b[0m\u001b[0;34m.\u001b[0m\u001b[0mjoin\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0;34m[\u001b[0m\u001b[0mreverse_map\u001b[0m\u001b[0;34m[\u001b[0m\u001b[0mt\u001b[0m\u001b[0;34m]\u001b[0m \u001b[0;32mfor\u001b[0m \u001b[0mt\u001b[0m \u001b[0;32min\u001b[0m \u001b[0mx\u001b[0m\u001b[0;34m]\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0m",
|
||||
"\u001b[0;32m<ipython-input-10-3f5fa6130b1d>\u001b[0m in \u001b[0;36m<listcomp>\u001b[0;34m(.0)\u001b[0m\n\u001b[1;32m 2\u001b[0m \u001b[0;34m\u001b[0m\u001b[0m\n\u001b[1;32m 3\u001b[0m \u001b[0;32mdef\u001b[0m \u001b[0mdecode\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0mx\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m:\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0;32m----> 4\u001b[0;31m \u001b[0;32mreturn\u001b[0m \u001b[0;34m''\u001b[0m\u001b[0;34m.\u001b[0m\u001b[0mjoin\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0;34m[\u001b[0m\u001b[0mreverse_map\u001b[0m\u001b[0;34m[\u001b[0m\u001b[0mt\u001b[0m\u001b[0;34m]\u001b[0m \u001b[0;32mfor\u001b[0m \u001b[0mt\u001b[0m \u001b[0;32min\u001b[0m \u001b[0mx\u001b[0m\u001b[0;34m]\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0m",
|
||||
"\u001b[0;31mKeyError\u001b[0m: 0"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"def generate_soft(model,size=100,start='Today ',temperature=1.0):\n",
|
||||
" inp = tokenizer.texts_to_sequences([start])[0]\n",
|
||||
" chars = inp\n",
|
||||
" for i in range(size):\n",
|
||||
" out = model(tf.expand_dims(tf.one_hot(inp,vocab_size),0))[0][-1]\n",
|
||||
" probs = tf.exp(tf.math.log(out)/temperature).numpy().astype(np.float64)\n",
|
||||
" probs = probs/np.sum(probs)\n",
|
||||
" nc = np.argmax(np.random.multinomial(1,probs,1))\n",
|
||||
" if nc==eos_token:\n",
|
||||
" break\n",
|
||||
" chars.append(nc)\n",
|
||||
" inp = inp+[nc]\n",
|
||||
" return decode(chars)\n",
|
||||
"\n",
|
||||
"words = ['Today ','On Sunday ','Moscow, ','President ','Little red riding hood ']\n",
|
||||
" \n",
|
||||
"for i in [0.3,0.8,1.0,1.3,1.8]:\n",
|
||||
" print(f\"\\n--- Temperature = {i}\")\n",
|
||||
" for j in range(5):\n",
|
||||
" print(generate_soft(model,size=300,start=words[j],temperature=i))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"हमने एक और पैरामीटर **तापमान** पेश किया है, जिसका उपयोग यह संकेत देने के लिए किया जाता है कि हमें उच्चतम संभावना से कितनी दृढ़ता से चिपकना चाहिए। यदि तापमान 1.0 है, तो हम निष्पक्ष बहुपद नमूना लेते हैं, और जब तापमान अनंत तक जाता है - सभी संभावनाएँ समान हो जाती हैं, और हम अगला वर्ण यादृच्छिक रूप से चुनते हैं। नीचे दिए गए उदाहरण में हम देख सकते हैं कि जब हम तापमान को बहुत अधिक बढ़ाते हैं तो पाठ अर्थहीन हो जाता है, और जब यह 0 के करीब होता है तो यह \"चक्रित\" कठोर-जनित पाठ जैसा दिखता है।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"\n---\n\n**अस्वीकरण**: \nयह दस्तावेज़ AI अनुवाद सेवा [Co-op Translator](https://github.com/Azure/co-op-translator) का उपयोग करके अनुवादित किया गया है। जबकि हम सटीकता सुनिश्चित करने का प्रयास करते हैं, कृपया ध्यान दें कि स्वचालित अनुवाद में त्रुटियां या अशुद्धियां हो सकती हैं। मूल भाषा में उपलब्ध मूल दस्तावेज़ को प्रामाणिक स्रोत माना जाना चाहिए। महत्वपूर्ण जानकारी के लिए, पेशेवर मानव अनुवाद की सिफारिश की जाती है। इस अनुवाद के उपयोग से उत्पन्न किसी भी गलतफहमी या गलत व्याख्या के लिए हम जिम्मेदार नहीं हैं।\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
"interpreter": {
|
||||
"hash": "16af2a8bbb083ea23e5e41c7f5787656b2ce26968575d8763f2c4b17f9cd711f"
|
||||
},
|
||||
"kernelspec": {
|
||||
"display_name": "Python 3.8.12 ('py38')",
|
||||
"language": "python",
|
||||
"name": "python3"
|
||||
},
|
||||
"language_info": {
|
||||
"codemirror_mode": {
|
||||
"name": "ipython",
|
||||
"version": 3
|
||||
},
|
||||
"file_extension": ".py",
|
||||
"mimetype": "text/x-python",
|
||||
"name": "python",
|
||||
"nbconvert_exporter": "python",
|
||||
"pygments_lexer": "ipython3",
|
||||
"version": "3.8.12"
|
||||
},
|
||||
"coopTranslator": {
|
||||
"original_hash": "9fbb7d5fda708537649f71f5f646fcde",
|
||||
"translation_date": "2025-08-31T15:12:26+00:00",
|
||||
"source_file": "lessons/5-NLP/17-GenerativeNetworks/GenerativeTF.ipynb",
|
||||
"language_code": "hi"
|
||||
}
|
||||
},
|
||||
"nbformat": 4,
|
||||
"nbformat_minor": 4
|
||||
}
|
||||
|
|
@ -0,0 +1,353 @@
|
|||
{
|
||||
"cells": [
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"# ध्यान तंत्र और ट्रांसफॉर्मर\n",
|
||||
"\n",
|
||||
"पुनरावर्ती नेटवर्क्स (Recurrent Networks) की एक बड़ी कमी यह है कि किसी अनुक्रम (sequence) के सभी शब्दों का परिणाम पर समान प्रभाव होता है। यह समस्या नामित इकाई पहचान (Named Entity Recognition) और मशीन अनुवाद (Machine Translation) जैसे अनुक्रम-से-अनुक्रम कार्यों के लिए मानक LSTM एन्कोडर-डिकोडर मॉडल्स के प्रदर्शन को कम कर देती है। वास्तविकता में, इनपुट अनुक्रम के कुछ विशेष शब्दों का अनुक्रमिक आउटपुट पर अन्य शब्दों की तुलना में अधिक प्रभाव होता है।\n",
|
||||
"\n",
|
||||
"मान लीजिए कि हमारे पास एक अनुक्रम-से-अनुक्रम मॉडल है, जैसे मशीन अनुवाद। इसे दो पुनरावर्ती नेटवर्क्स द्वारा लागू किया जाता है, जहां एक नेटवर्क (**एन्कोडर**) इनपुट अनुक्रम को छिपी हुई अवस्था (hidden state) में संकुचित करता है, और दूसरा नेटवर्क (**डिकोडर**) इस छिपी हुई अवस्था को अनुवादित परिणाम में विस्तारित करता है। इस दृष्टिकोण की समस्या यह है कि नेटवर्क की अंतिम अवस्था वाक्य की शुरुआत को याद रखने में कठिनाई महसूस करती है, जिससे लंबे वाक्यों पर मॉडल की गुणवत्ता खराब हो जाती है।\n",
|
||||
"\n",
|
||||
"**ध्यान तंत्र (Attention Mechanisms)** प्रत्येक इनपुट वेक्टर के संदर्भीय प्रभाव को RNN के प्रत्येक आउटपुट भविष्यवाणी पर भारित करने का एक साधन प्रदान करते हैं। इसे लागू करने का तरीका यह है कि इनपुट RNN की मध्यवर्ती अवस्थाओं और आउटपुट RNN के बीच शॉर्टकट्स बनाए जाते हैं। इस प्रकार, जब आउटपुट प्रतीक $y_t$ उत्पन्न किया जाता है, तो हम सभी इनपुट छिपी अवस्थाओं $h_i$ को विभिन्न भार गुणांक $\\alpha_{t,i}$ के साथ ध्यान में रखेंगे।\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"*एन्कोडर-डिकोडर मॉडल एडिटिव ध्यान तंत्र के साथ [Bahdanau et al., 2015](https://arxiv.org/pdf/1409.0473.pdf) से, [इस ब्लॉग पोस्ट](https://lilianweng.github.io/lil-log/2018/06/24/attention-attention.html) से उद्धृत*\n",
|
||||
"\n",
|
||||
"ध्यान मैट्रिक्स $\\{\\alpha_{i,j}\\}$ यह दर्शाता है कि आउटपुट अनुक्रम में किसी दिए गए शब्द को उत्पन्न करने में कौन से इनपुट शब्द कितनी भूमिका निभाते हैं। नीचे इस प्रकार की एक मैट्रिक्स का उदाहरण दिया गया है:\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"*[Bahdanau et al., 2015](https://arxiv.org/pdf/1409.0473.pdf) (Fig.3) से लिया गया चित्र]*\n",
|
||||
"\n",
|
||||
"ध्यान तंत्र वर्तमान या लगभग वर्तमान प्राकृतिक भाषा प्रसंस्करण (Natural Language Processing) में अत्याधुनिक तकनीकों के लिए जिम्मेदार हैं। हालांकि, ध्यान जोड़ने से मॉडल के पैरामीटरों की संख्या में काफी वृद्धि होती है, जिससे RNNs के साथ स्केलिंग समस्याएं उत्पन्न होती हैं। RNNs को स्केल करने की एक प्रमुख बाधा यह है कि मॉडल की पुनरावर्ती प्रकृति प्रशिक्षण को बैच और समानांतर बनाने में चुनौतीपूर्ण बनाती है। RNN में अनुक्रम के प्रत्येक तत्व को क्रमिक रूप से संसाधित करना पड़ता है, जिससे इसे आसानी से समानांतर नहीं किया जा सकता।\n",
|
||||
"\n",
|
||||
"ध्यान तंत्रों को अपनाने और इस बाधा ने आज के अत्याधुनिक ट्रांसफॉर्मर मॉडल्स के निर्माण का मार्ग प्रशस्त किया, जिन्हें हम BERT से OpenGPT3 तक जानते और उपयोग करते हैं।\n",
|
||||
"\n",
|
||||
"## ट्रांसफॉर्मर मॉडल्स\n",
|
||||
"\n",
|
||||
"प्रत्येक पूर्वानुमान के संदर्भ को अगले मूल्यांकन चरण में अग्रेषित करने के बजाय, **ट्रांसफॉर्मर मॉडल्स** **स्थिति एन्कोडिंग्स (positional encodings)** और ध्यान का उपयोग करके दिए गए इनपुट के संदर्भ को एक निर्दिष्ट पाठ विंडो के भीतर कैप्चर करते हैं। नीचे दी गई छवि दिखाती है कि स्थिति एन्कोडिंग्स और ध्यान का उपयोग करके किसी विंडो के भीतर संदर्भ को कैसे कैप्चर किया जा सकता है।\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"चूंकि प्रत्येक इनपुट स्थिति को स्वतंत्र रूप से प्रत्येक आउटपुट स्थिति पर मैप किया जाता है, ट्रांसफॉर्मर RNNs की तुलना में बेहतर समानांतरता प्रदान करते हैं, जिससे बहुत बड़े और अधिक अभिव्यक्तिपूर्ण भाषा मॉडल्स सक्षम होते हैं। प्रत्येक ध्यान हेड का उपयोग शब्दों के बीच विभिन्न संबंधों को सीखने के लिए किया जा सकता है, जो डाउनस्ट्रीम प्राकृतिक भाषा प्रसंस्करण कार्यों को बेहतर बनाता है।\n",
|
||||
"\n",
|
||||
"**BERT** (Bidirectional Encoder Representations from Transformers) एक बहुत बड़ा बहु-स्तरीय ट्रांसफॉर्मर नेटवर्क है, जिसमें *BERT-base* के लिए 12 परतें और *BERT-large* के लिए 24 परतें होती हैं। इस मॉडल को पहले बड़े पाठ डेटा (विकिपीडिया + किताबें) पर असुपरवाइज्ड प्रशिक्षण (वाक्य में छिपे हुए शब्दों की भविष्यवाणी) का उपयोग करके प्रशिक्षित किया जाता है। पूर्व-प्रशिक्षण के दौरान, मॉडल भाषा की महत्वपूर्ण समझ को आत्मसात करता है, जिसे फिर अन्य डेटासेट्स के साथ फाइन-ट्यूनिंग के माध्यम से उपयोग किया जा सकता है। इस प्रक्रिया को **ट्रांसफर लर्निंग** कहा जाता है।\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"ट्रांसफॉर्मर आर्किटेक्चर के कई प्रकार हैं, जैसे BERT, DistilBERT, BigBird, OpenGPT3 और अन्य, जिन्हें फाइन-ट्यून किया जा सकता है। [HuggingFace पैकेज](https://github.com/huggingface/) PyTorch के साथ इनमें से कई आर्किटेक्चर को प्रशिक्षित करने के लिए एक रिपॉजिटरी प्रदान करता है।\n",
|
||||
"\n",
|
||||
"## टेक्स्ट वर्गीकरण के लिए BERT का उपयोग\n",
|
||||
"\n",
|
||||
"आइए देखें कि हम अपने पारंपरिक कार्य को हल करने के लिए पूर्व-प्रशिक्षित BERT मॉडल का उपयोग कैसे कर सकते हैं: अनुक्रम वर्गीकरण। हम अपने मूल AG News डेटासेट को वर्गीकृत करेंगे।\n",
|
||||
"\n",
|
||||
"सबसे पहले, HuggingFace लाइब्रेरी और हमारा डेटासेट लोड करें:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 10,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Loading dataset...\n",
|
||||
"Building vocab...\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"import torch\n",
|
||||
"import torchtext\n",
|
||||
"from torchnlp import *\n",
|
||||
"import transformers\n",
|
||||
"train_dataset, test_dataset, classes, vocab = load_dataset()\n",
|
||||
"vocab_len = len(vocab)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"क्योंकि हम प्री-ट्रेंड BERT मॉडल का उपयोग करेंगे, हमें एक विशेष टोकनाइज़र का उपयोग करना होगा। सबसे पहले, हम प्री-ट्रेंड BERT मॉडल से जुड़े टोकनाइज़र को लोड करेंगे।\n",
|
||||
"\n",
|
||||
"HuggingFace लाइब्रेरी में प्री-ट्रेंड मॉडल्स का एक रिपॉजिटरी है, जिसे आप केवल उनके नामों को `from_pretrained` फंक्शन के आर्ग्युमेंट्स के रूप में देकर उपयोग कर सकते हैं। मॉडल के लिए आवश्यक सभी बाइनरी फाइलें स्वचालित रूप से डाउनलोड हो जाएंगी।\n",
|
||||
"\n",
|
||||
"हालांकि, कुछ स्थितियों में आपको अपने खुद के मॉडल्स लोड करने की आवश्यकता हो सकती है। ऐसे मामलों में, आप उस डायरेक्टरी को निर्दिष्ट कर सकते हैं जिसमें सभी संबंधित फाइलें हों, जैसे टोकनाइज़र के लिए पैरामीटर्स, मॉडल पैरामीटर्स के साथ `config.json` फाइल, बाइनरी वेट्स आदि।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 11,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"# To load the model from Internet repository using model name. \n",
|
||||
"# Use this if you are running from your own copy of the notebooks\n",
|
||||
"bert_model = 'bert-base-uncased' \n",
|
||||
"\n",
|
||||
"# To load the model from the directory on disk. Use this for Microsoft Learn module, because we have\n",
|
||||
"# prepared all required files for you.\n",
|
||||
"bert_model = './bert'\n",
|
||||
"\n",
|
||||
"tokenizer = transformers.BertTokenizer.from_pretrained(bert_model)\n",
|
||||
"\n",
|
||||
"MAX_SEQ_LEN = 128\n",
|
||||
"PAD_INDEX = tokenizer.convert_tokens_to_ids(tokenizer.pad_token)\n",
|
||||
"UNK_INDEX = tokenizer.convert_tokens_to_ids(tokenizer.unk_token)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"`tokenizer` ऑब्जेक्ट में `encode` फ़ंक्शन होता है जिसे सीधे टेक्स्ट को एन्कोड करने के लिए उपयोग किया जा सकता है:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 15,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"[101, 1052, 22123, 2953, 2818, 2003, 1037, 2307, 7705, 2005, 17953, 2361, 102]"
|
||||
]
|
||||
},
|
||||
"execution_count": 15,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"tokenizer.encode('PyTorch is a great framework for NLP')"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"फिर, आइए ऐसे इटरेटर बनाते हैं जिन्हें हम प्रशिक्षण के दौरान डेटा तक पहुंचने के लिए उपयोग करेंगे। क्योंकि BERT अपनी स्वयं की एनकोडिंग फ़ंक्शन का उपयोग करता है, हमें एक पैडिंग फ़ंक्शन को परिभाषित करने की आवश्यकता होगी जो पहले परिभाषित `padify` के समान हो:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"def pad_bert(b):\n",
|
||||
" # b is the list of tuples of length batch_size\n",
|
||||
" # - first element of a tuple = label, \n",
|
||||
" # - second = feature (text sequence)\n",
|
||||
" # build vectorized sequence\n",
|
||||
" v = [tokenizer.encode(x[1]) for x in b]\n",
|
||||
" # compute max length of a sequence in this minibatch\n",
|
||||
" l = max(map(len,v))\n",
|
||||
" return ( # tuple of two tensors - labels and features\n",
|
||||
" torch.LongTensor([t[0] for t in b]),\n",
|
||||
" torch.stack([torch.nn.functional.pad(torch.tensor(t),(0,l-len(t)),mode='constant',value=0) for t in v])\n",
|
||||
" )\n",
|
||||
"\n",
|
||||
"train_loader = torch.utils.data.DataLoader(train_dataset, batch_size=8, collate_fn=pad_bert, shuffle=True)\n",
|
||||
"test_loader = torch.utils.data.DataLoader(test_dataset, batch_size=8, collate_fn=pad_bert)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"हमारे मामले में, हम पूर्व-प्रशिक्षित BERT मॉडल का उपयोग करेंगे जिसे `bert-base-uncased` कहा जाता है। आइए मॉडल को `BertForSequenceClassfication` पैकेज का उपयोग करके लोड करें। यह सुनिश्चित करता है कि हमारे मॉडल में पहले से ही वर्गीकरण के लिए आवश्यक संरचना है, जिसमें अंतिम वर्गीकर्ता भी शामिल है। आपको एक चेतावनी संदेश दिखाई देगा जिसमें कहा जाएगा कि अंतिम वर्गीकर्ता के वज़न प्रारंभ नहीं किए गए हैं, और मॉडल को पूर्व-प्रशिक्षण की आवश्यकता होगी - यह पूरी तरह से ठीक है, क्योंकि यही हम करने जा रहे हैं!\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 9,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stderr",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Some weights of the model checkpoint at ./bert were not used when initializing BertForSequenceClassification: ['cls.predictions.bias', 'cls.predictions.transform.dense.weight', 'cls.predictions.transform.dense.bias', 'cls.predictions.decoder.weight', 'cls.seq_relationship.weight', 'cls.seq_relationship.bias', 'cls.predictions.transform.LayerNorm.weight', 'cls.predictions.transform.LayerNorm.bias']\n",
|
||||
"- This IS expected if you are initializing BertForSequenceClassification from the checkpoint of a model trained on another task or with another architecture (e.g. initializing a BertForSequenceClassification model from a BertForPreTraining model).\n",
|
||||
"- This IS NOT expected if you are initializing BertForSequenceClassification from the checkpoint of a model that you expect to be exactly identical (initializing a BertForSequenceClassification model from a BertForSequenceClassification model).\n",
|
||||
"Some weights of BertForSequenceClassification were not initialized from the model checkpoint at ./bert and are newly initialized: ['classifier.weight', 'classifier.bias']\n",
|
||||
"You should probably TRAIN this model on a down-stream task to be able to use it for predictions and inference.\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"model = transformers.BertForSequenceClassification.from_pretrained(bert_model,num_labels=4).to(device)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"अब हम प्रशिक्षण शुरू करने के लिए तैयार हैं! क्योंकि BERT पहले से ही प्री-ट्रेंड है, हम एक बहुत ही छोटे लर्निंग रेट से शुरुआत करना चाहते हैं ताकि प्रारंभिक वज़न खराब न हो जाएं।\n",
|
||||
"\n",
|
||||
"सारा कठिन काम `BertForSequenceClassification` मॉडल द्वारा किया जाता है। जब हम प्रशिक्षण डेटा पर मॉडल को कॉल करते हैं, तो यह इनपुट मिनीबैच के लिए लॉस और नेटवर्क आउटपुट दोनों लौटाता है। हम पैरामीटर ऑप्टिमाइज़ेशन के लिए लॉस का उपयोग करते हैं (`loss.backward()` बैकवर्ड पास करता है), और `out` का उपयोग प्रशिक्षण सटीकता की गणना के लिए करते हैं, जो प्राप्त लेबल `labs` (जो `argmax` का उपयोग करके गणना किए जाते हैं) को अपेक्षित `labels` के साथ तुलना करता है।\n",
|
||||
"\n",
|
||||
"प्रक्रिया को नियंत्रित करने के लिए, हम कई पुनरावृत्तियों में लॉस और सटीकता को संचित करते हैं, और हर `report_freq` प्रशिक्षण चक्रों के बाद उन्हें प्रिंट करते हैं।\n",
|
||||
"\n",
|
||||
"यह प्रशिक्षण संभवतः काफी समय लेगा, इसलिए हम पुनरावृत्तियों की संख्या को सीमित करते हैं।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 6,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Loss = 1.1254194641113282, Accuracy = 0.585\n",
|
||||
"Loss = 0.6194715118408203, Accuracy = 0.83\n",
|
||||
"Loss = 0.46665248870849607, Accuracy = 0.8475\n",
|
||||
"Loss = 0.4309701919555664, Accuracy = 0.8575\n",
|
||||
"Loss = 0.35427074432373046, Accuracy = 0.8825\n",
|
||||
"Loss = 0.3306886291503906, Accuracy = 0.8975\n",
|
||||
"Loss = 0.30340143203735354, Accuracy = 0.8975\n",
|
||||
"Loss = 0.26139299392700194, Accuracy = 0.915\n",
|
||||
"Loss = 0.26708646774291994, Accuracy = 0.9225\n",
|
||||
"Loss = 0.3667240524291992, Accuracy = 0.8675\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"optimizer = torch.optim.Adam(model.parameters(), lr=2e-5)\n",
|
||||
"\n",
|
||||
"report_freq = 50\n",
|
||||
"iterations = 500 # make this larger to train for longer time!\n",
|
||||
"\n",
|
||||
"model.train()\n",
|
||||
"\n",
|
||||
"i,c = 0,0\n",
|
||||
"acc_loss = 0\n",
|
||||
"acc_acc = 0\n",
|
||||
"\n",
|
||||
"for labels,texts in train_loader:\n",
|
||||
" labels = labels.to(device)-1 # get labels in the range 0-3 \n",
|
||||
" texts = texts.to(device)\n",
|
||||
" loss, out = model(texts, labels=labels)[:2]\n",
|
||||
" labs = out.argmax(dim=1)\n",
|
||||
" acc = torch.mean((labs==labels).type(torch.float32))\n",
|
||||
" optimizer.zero_grad()\n",
|
||||
" loss.backward()\n",
|
||||
" optimizer.step()\n",
|
||||
" acc_loss += loss\n",
|
||||
" acc_acc += acc\n",
|
||||
" i+=1\n",
|
||||
" c+=1\n",
|
||||
" if i%report_freq==0:\n",
|
||||
" print(f\"Loss = {acc_loss.item()/c}, Accuracy = {acc_acc.item()/c}\")\n",
|
||||
" c = 0\n",
|
||||
" acc_loss = 0\n",
|
||||
" acc_acc = 0\n",
|
||||
" iterations-=1\n",
|
||||
" if not iterations:\n",
|
||||
" break"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"आप देख सकते हैं (खासकर अगर आप iterations की संख्या बढ़ाते हैं और थोड़ा अधिक इंतजार करते हैं) कि BERT classification हमें काफी अच्छी सटीकता देता है! इसका कारण यह है कि BERT पहले से ही भाषा की संरचना को काफी अच्छी तरह समझता है, और हमें केवल अंतिम classifier को fine-tune करना होता है। हालांकि, क्योंकि BERT एक बड़ा मॉडल है, पूरा training प्रक्रिया काफी समय लेती है और इसके लिए गंभीर computational power की आवश्यकता होती है! (GPU, और बेहतर होगा कि एक से अधिक GPU हों)।\n",
|
||||
"\n",
|
||||
"> **Note:** हमारे उदाहरण में, हमने सबसे छोटे pre-trained BERT मॉडल्स में से एक का उपयोग किया है। बड़े मॉडल्स उपलब्ध हैं, जो संभवतः बेहतर परिणाम दे सकते हैं।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## मॉडल प्रदर्शन का मूल्यांकन\n",
|
||||
"\n",
|
||||
"अब हम अपने मॉडल के प्रदर्शन का परीक्षण डेटा सेट पर मूल्यांकन कर सकते हैं। मूल्यांकन लूप प्रशिक्षण लूप के समान ही है, लेकिन हमें यह नहीं भूलना चाहिए कि मॉडल को मूल्यांकन मोड में स्विच करने के लिए `model.eval()` कॉल करना आवश्यक है।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 10,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Final accuracy: 0.9047029702970297\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"model.eval()\n",
|
||||
"iterations = 100\n",
|
||||
"acc = 0\n",
|
||||
"i = 0\n",
|
||||
"for labels,texts in test_loader:\n",
|
||||
" labels = labels.to(device)-1 \n",
|
||||
" texts = texts.to(device)\n",
|
||||
" _, out = model(texts, labels=labels)[:2]\n",
|
||||
" labs = out.argmax(dim=1)\n",
|
||||
" acc += torch.mean((labs==labels).type(torch.float32))\n",
|
||||
" i+=1\n",
|
||||
" if i>iterations: break\n",
|
||||
" \n",
|
||||
"print(f\"Final accuracy: {acc.item()/i}\")"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## मुख्य बातें\n",
|
||||
"\n",
|
||||
"इस यूनिट में, हमने देखा कि **transformers** लाइब्रेरी से प्री-ट्रेंड भाषा मॉडल लेना और उसे हमारे टेक्स्ट क्लासिफिकेशन टास्क के लिए अनुकूलित करना कितना आसान है। इसी तरह, BERT मॉडल का उपयोग एंटिटी एक्सट्रैक्शन, प्रश्न उत्तर देने और अन्य NLP टास्क के लिए किया जा सकता है।\n",
|
||||
"\n",
|
||||
"ट्रांसफॉर्मर मॉडल NLP में वर्तमान में सबसे उन्नत तकनीक का प्रतिनिधित्व करते हैं, और अधिकांश मामलों में, जब आप कस्टम NLP समाधान लागू करना शुरू करते हैं, तो यह पहला विकल्प होना चाहिए जिसके साथ आप प्रयोग करें। हालांकि, यदि आप उन्नत न्यूरल मॉडल बनाना चाहते हैं, तो इस मॉड्यूल में चर्चा किए गए पुनरावर्ती न्यूरल नेटवर्क के मूलभूत सिद्धांतों को समझना अत्यंत महत्वपूर्ण है।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"\n---\n\n**अस्वीकरण**: \nयह दस्तावेज़ AI अनुवाद सेवा [Co-op Translator](https://github.com/Azure/co-op-translator) का उपयोग करके अनुवादित किया गया है। जबकि हम सटीकता सुनिश्चित करने का प्रयास करते हैं, कृपया ध्यान दें कि स्वचालित अनुवाद में त्रुटियां या अशुद्धियां हो सकती हैं। मूल भाषा में उपलब्ध मूल दस्तावेज़ को प्रामाणिक स्रोत माना जाना चाहिए। महत्वपूर्ण जानकारी के लिए, पेशेवर मानव अनुवाद की सिफारिश की जाती है। इस अनुवाद के उपयोग से उत्पन्न किसी भी गलतफहमी या गलत व्याख्या के लिए हम उत्तरदायी नहीं हैं।\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
"kernelspec": {
|
||||
"display_name": "py37_pytorch",
|
||||
"language": "python",
|
||||
"name": "conda-env-py37_pytorch-py"
|
||||
},
|
||||
"language_info": {
|
||||
"codemirror_mode": {
|
||||
"name": "ipython",
|
||||
"version": 3
|
||||
},
|
||||
"file_extension": ".py",
|
||||
"mimetype": "text/x-python",
|
||||
"name": "python",
|
||||
"nbconvert_exporter": "python",
|
||||
"pygments_lexer": "ipython3",
|
||||
"version": "3.7.7"
|
||||
},
|
||||
"coopTranslator": {
|
||||
"original_hash": "753865967678a92dbce7d7efbd36d980",
|
||||
"translation_date": "2025-08-31T15:17:09+00:00",
|
||||
"source_file": "lessons/5-NLP/18-Transformers/TransformersPyTorch.ipynb",
|
||||
"language_code": "hi"
|
||||
}
|
||||
},
|
||||
"nbformat": 4,
|
||||
"nbformat_minor": 4
|
||||
}
|
||||
|
|
@ -0,0 +1,819 @@
|
|||
{
|
||||
"cells": [
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"# ध्यान तंत्र और ट्रांसफॉर्मर्स\n",
|
||||
"\n",
|
||||
"पुनरावर्ती नेटवर्क्स (Recurrent Networks) की एक बड़ी कमी यह है कि अनुक्रम में सभी शब्दों का परिणाम पर समान प्रभाव होता है। यह नामित इकाई पहचान (Named Entity Recognition) और मशीन अनुवाद (Machine Translation) जैसे अनुक्रम-से-अनुक्रम कार्यों के लिए मानक LSTM एन्कोडर-डिकोडर मॉडल्स के साथ उप-इष्टतम प्रदर्शन का कारण बनता है। वास्तविकता में, इनपुट अनुक्रम के कुछ विशिष्ट शब्दों का अनुक्रमिक आउटपुट पर अन्य शब्दों की तुलना में अधिक प्रभाव होता है।\n",
|
||||
"\n",
|
||||
"मशीन अनुवाद जैसे अनुक्रम-से-अनुक्रम मॉडल पर विचार करें। इसे दो पुनरावर्ती नेटवर्क्स द्वारा लागू किया जाता है, जहां एक नेटवर्क (**एन्कोडर**) इनपुट अनुक्रम को छिपी हुई स्थिति (hidden state) में संक्षेपित करता है, और दूसरा नेटवर्क (**डिकोडर**) इस छिपी हुई स्थिति को अनुवादित परिणाम में बदलता है। इस दृष्टिकोण की समस्या यह है कि नेटवर्क की अंतिम स्थिति को वाक्य की शुरुआत को याद रखने में कठिनाई होती है, जिससे लंबे वाक्यों पर मॉडल की गुणवत्ता खराब हो जाती है।\n",
|
||||
"\n",
|
||||
"**ध्यान तंत्र (Attention Mechanisms)** प्रत्येक इनपुट वेक्टर के संदर्भ प्रभाव को RNN के प्रत्येक आउटपुट भविष्यवाणी पर भारित करने का एक साधन प्रदान करते हैं। इसे लागू करने का तरीका यह है कि इनपुट RNN की मध्यवर्ती अवस्थाओं और आउटपुट RNN के बीच शॉर्टकट्स बनाए जाते हैं। इस प्रकार, जब आउटपुट प्रतीक $y_t$ उत्पन्न किया जाता है, तो हम सभी इनपुट छिपी हुई अवस्थाओं $h_i$ को विभिन्न भार गुणांक $\\alpha_{t,i}$ के साथ ध्यान में रखेंगे।\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"*[Bahdanau et al., 2015](https://arxiv.org/pdf/1409.0473.pdf) में एडिटिव ध्यान तंत्र के साथ एन्कोडर-डिकोडर मॉडल, [इस ब्लॉग पोस्ट](https://lilianweng.github.io/lil-log/2018/06/24/attention-attention.html) से उद्धृत*\n",
|
||||
"\n",
|
||||
"ध्यान मैट्रिक्स $\\{\\alpha_{i,j}\\}$ यह दर्शाएगा कि आउटपुट अनुक्रम में किसी दिए गए शब्द को उत्पन्न करने में कौन से इनपुट शब्द कितनी भूमिका निभाते हैं। नीचे इस प्रकार की मैट्रिक्स का एक उदाहरण दिया गया है:\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"*[Bahdanau et al., 2015](https://arxiv.org/pdf/1409.0473.pdf) (Fig.3) से ली गई आकृति]*\n",
|
||||
"\n",
|
||||
"ध्यान तंत्र प्राकृतिक भाषा प्रसंस्करण (Natural Language Processing) में वर्तमान या निकट वर्तमान की अत्याधुनिक स्थिति के लिए जिम्मेदार हैं। हालांकि, ध्यान जोड़ने से मॉडल के पैरामीटरों की संख्या में काफी वृद्धि होती है, जिससे RNNs के साथ स्केलिंग समस्याएं उत्पन्न होती हैं। RNNs को स्केल करने की एक प्रमुख बाधा यह है कि मॉडल की पुनरावर्ती प्रकृति प्रशिक्षण को बैच और समानांतर बनाने में चुनौतीपूर्ण बनाती है। RNN में अनुक्रम के प्रत्येक तत्व को क्रमिक क्रम में संसाधित करना पड़ता है, जिसका अर्थ है कि इसे आसानी से समानांतर नहीं किया जा सकता।\n",
|
||||
"\n",
|
||||
"ध्यान तंत्रों को अपनाने और इस बाधा ने उन अत्याधुनिक ट्रांसफॉर्मर मॉडलों के निर्माण का मार्ग प्रशस्त किया, जिन्हें हम आज BERT से OpenGPT3 तक उपयोग करते हैं।\n",
|
||||
"\n",
|
||||
"## ट्रांसफॉर्मर मॉडल्स\n",
|
||||
"\n",
|
||||
"पिछली भविष्यवाणी के संदर्भ को अगले मूल्यांकन चरण में अग्रेषित करने के बजाय, **ट्रांसफॉर्मर मॉडल्स** **पोजिशनल एन्कोडिंग्स** और **ध्यान** का उपयोग करते हैं ताकि दिए गए इनपुट के संदर्भ को एक निर्दिष्ट पाठ विंडो के भीतर कैप्चर किया जा सके। नीचे दी गई छवि दिखाती है कि पोजिशनल एन्कोडिंग्स और ध्यान का उपयोग करके किसी विंडो के भीतर संदर्भ को कैसे कैप्चर किया जा सकता है।\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"चूंकि प्रत्येक इनपुट स्थिति को स्वतंत्र रूप से प्रत्येक आउटपुट स्थिति पर मैप किया जाता है, ट्रांसफॉर्मर्स RNNs की तुलना में बेहतर समानांतरता प्रदान कर सकते हैं, जिससे बड़े और अधिक अभिव्यक्तिपूर्ण भाषा मॉडल सक्षम होते हैं। प्रत्येक ध्यान हेड का उपयोग शब्दों के बीच विभिन्न संबंधों को सीखने के लिए किया जा सकता है, जो डाउनस्ट्रीम प्राकृतिक भाषा प्रसंस्करण कार्यों में सुधार करता है।\n",
|
||||
"\n",
|
||||
"## सरल ट्रांसफॉर्मर मॉडल बनाना\n",
|
||||
"\n",
|
||||
"Keras में बिल्ट-इन ट्रांसफॉर्मर लेयर नहीं है, लेकिन हम अपना खुद का बना सकते हैं। पहले की तरह, हम AG News डेटासेट के टेक्स्ट वर्गीकरण पर ध्यान केंद्रित करेंगे, लेकिन यह उल्लेख करना महत्वपूर्ण है कि ट्रांसफॉर्मर मॉडल अधिक कठिन NLP कार्यों में सर्वश्रेष्ठ परिणाम दिखाते हैं।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 1,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"import tensorflow as tf\n",
|
||||
"from tensorflow import keras\n",
|
||||
"import tensorflow_datasets as tfds\n",
|
||||
"import numpy as np\n",
|
||||
"\n",
|
||||
"ds_train, ds_test = tfds.load('ag_news_subset').values()\n",
|
||||
"\n",
|
||||
"def extract_text(x):\n",
|
||||
" return x['title']+' '+x['description']\n",
|
||||
"\n",
|
||||
"def tupelize(x):\n",
|
||||
" return (extract_text(x),x['label'])"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"केरस में नई लेयर्स को `Layer` क्लास को सबक्लास करना चाहिए और `call` मेथड को लागू करना चाहिए। चलिए **Positional Embedding** लेयर से शुरू करते हैं। हम [आधिकारिक केरस दस्तावेज़](https://keras.io/examples/nlp/text_classification_with_transformer/) से कुछ कोड का उपयोग करेंगे। हम मान लेंगे कि हम सभी इनपुट अनुक्रमों को `maxlen` लंबाई तक पैड करते हैं।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 2,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"class TokenAndPositionEmbedding(keras.layers.Layer):\n",
|
||||
" def __init__(self, maxlen, vocab_size, embed_dim):\n",
|
||||
" super(TokenAndPositionEmbedding, self).__init__()\n",
|
||||
" self.token_emb = keras.layers.Embedding(input_dim=vocab_size, output_dim=embed_dim)\n",
|
||||
" self.pos_emb = keras.layers.Embedding(input_dim=maxlen, output_dim=embed_dim)\n",
|
||||
" self.maxlen = maxlen\n",
|
||||
"\n",
|
||||
" def call(self, x):\n",
|
||||
" maxlen = self.maxlen\n",
|
||||
" positions = tf.range(start=0, limit=maxlen, delta=1)\n",
|
||||
" positions = self.pos_emb(positions)\n",
|
||||
" x = self.token_emb(x)\n",
|
||||
" return x+positions"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"यह लेयर दो `Embedding` लेयर्स से बनी होती है: एक टोकन को एम्बेड करने के लिए (जैसा कि हमने पहले चर्चा की है) और दूसरी टोकन की पोजीशन को एम्बेड करने के लिए। टोकन की पोजीशन को 0 से `maxlen` तक के प्राकृतिक संख्याओं के अनुक्रम के रूप में `tf.range` का उपयोग करके बनाया जाता है, और फिर इसे एम्बेडिंग लेयर में पास किया जाता है। इसके बाद, दो प्राप्त एम्बेडिंग वेक्टर को जोड़ा जाता है, जिससे इनपुट का पोजीशनली-एम्बेडेड प्रतिनिधित्व तैयार होता है, जिसका आकार `maxlen`$\\times$`embed_dim` होता है।\n",
|
||||
"\n",
|
||||
"अब, आइए ट्रांसफॉर्मर ब्लॉक को लागू करें। यह पहले परिभाषित एम्बेडिंग लेयर के आउटपुट को लेगा:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 3,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"class TransformerBlock(keras.layers.Layer):\n",
|
||||
" def __init__(self, embed_dim, num_heads, ff_dim, rate=0.1):\n",
|
||||
" super(TransformerBlock, self).__init__()\n",
|
||||
" self.att = keras.layers.MultiHeadAttention(num_heads=num_heads, key_dim=embed_dim, name='attn')\n",
|
||||
" self.ffn = keras.Sequential(\n",
|
||||
" [keras.layers.Dense(ff_dim, activation=\"relu\"), keras.layers.Dense(embed_dim),]\n",
|
||||
" )\n",
|
||||
" self.layernorm1 = keras.layers.LayerNormalization(epsilon=1e-6)\n",
|
||||
" self.layernorm2 = keras.layers.LayerNormalization(epsilon=1e-6)\n",
|
||||
" self.dropout1 = keras.layers.Dropout(rate)\n",
|
||||
" self.dropout2 = keras.layers.Dropout(rate)\n",
|
||||
"\n",
|
||||
" def call(self, inputs, training):\n",
|
||||
" attn_output = self.att(inputs, inputs)\n",
|
||||
" attn_output = self.dropout1(attn_output, training=training)\n",
|
||||
" out1 = self.layernorm1(inputs + attn_output)\n",
|
||||
" ffn_output = self.ffn(out1)\n",
|
||||
" ffn_output = self.dropout2(ffn_output, training=training)\n",
|
||||
" return self.layernorm2(out1 + ffn_output)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"अब, हम पूरा ट्रांसफॉर्मर मॉडल परिभाषित करने के लिए तैयार हैं:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Model: \"sequential_1\"\n",
|
||||
"_________________________________________________________________\n",
|
||||
"Layer (type) Output Shape Param # \n",
|
||||
"=================================================================\n",
|
||||
"text_vectorization (TextVect (None, 256) 0 \n",
|
||||
"_________________________________________________________________\n",
|
||||
"token_and_position_embedding (None, 256, 32) 648192 \n",
|
||||
"_________________________________________________________________\n",
|
||||
"transformer_block (Transform (None, 256, 32) 10656 \n",
|
||||
"_________________________________________________________________\n",
|
||||
"global_average_pooling1d (Gl (None, 32) 0 \n",
|
||||
"_________________________________________________________________\n",
|
||||
"dropout_2 (Dropout) (None, 32) 0 \n",
|
||||
"_________________________________________________________________\n",
|
||||
"dense_2 (Dense) (None, 20) 660 \n",
|
||||
"_________________________________________________________________\n",
|
||||
"dropout_3 (Dropout) (None, 20) 0 \n",
|
||||
"_________________________________________________________________\n",
|
||||
"dense_3 (Dense) (None, 4) 84 \n",
|
||||
"=================================================================\n",
|
||||
"Total params: 659,592\n",
|
||||
"Trainable params: 659,592\n",
|
||||
"Non-trainable params: 0\n",
|
||||
"_________________________________________________________________\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"embed_dim = 32 # Embedding size for each token\n",
|
||||
"num_heads = 2 # Number of attention heads\n",
|
||||
"ff_dim = 32 # Hidden layer size in feed forward network inside transformer\n",
|
||||
"maxlen = 256\n",
|
||||
"vocab_size = 20000\n",
|
||||
"\n",
|
||||
"model = keras.models.Sequential([\n",
|
||||
" keras.layers.experimental.preprocessing.TextVectorization(max_tokens=vocab_size,output_sequence_length=maxlen, input_shape=(1,)),\n",
|
||||
" TokenAndPositionEmbedding(maxlen, vocab_size, embed_dim),\n",
|
||||
" TransformerBlock(embed_dim, num_heads, ff_dim),\n",
|
||||
" keras.layers.GlobalAveragePooling1D(),\n",
|
||||
" keras.layers.Dropout(0.1),\n",
|
||||
" keras.layers.Dense(20, activation=\"relu\"),\n",
|
||||
" keras.layers.Dropout(0.1),\n",
|
||||
" keras.layers.Dense(4, activation=\"softmax\")\n",
|
||||
"])\n",
|
||||
"\n",
|
||||
"model.summary()"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 5,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Training tokenizer\n",
|
||||
"938/938 [==============================] - 45s 39ms/step - loss: 0.4978 - acc: 0.8068 - val_loss: 0.2808 - val_acc: 0.9124\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"<tensorflow.python.keras.callbacks.History at 0x7f9c2427a0d0>"
|
||||
]
|
||||
},
|
||||
"execution_count": 5,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"print('Training tokenizer')\n",
|
||||
"model.layers[0].adapt(ds_train.map(extract_text))\n",
|
||||
"model.compile(loss='sparse_categorical_crossentropy',metrics=['acc'], optimizer='adam')\n",
|
||||
"model.fit(ds_train.map(tupelize).batch(128),validation_data=ds_test.map(tupelize).batch(128))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## BERT ट्रांसफॉर्मर मॉडल्स\n",
|
||||
"\n",
|
||||
"**BERT** (Bidirectional Encoder Representations from Transformers) एक बहुत बड़ा मल्टी-लेयर ट्रांसफॉर्मर नेटवर्क है, जिसमें *BERT-base* के लिए 12 लेयर्स और *BERT-large* के लिए 24 लेयर्स होती हैं। इस मॉडल को पहले बड़े टेक्स्ट डेटा (WikiPedia + किताबें) के कॉर्पस पर अनसुपरवाइज्ड ट्रेनिंग (एक वाक्य में छुपे हुए शब्दों की भविष्यवाणी करना) का उपयोग करके प्री-ट्रेन किया जाता है। प्री-ट्रेनिंग के दौरान, मॉडल भाषा को समझने की एक महत्वपूर्ण क्षमता विकसित करता है, जिसे फिर अन्य डेटा सेट्स के साथ फाइन-ट्यूनिंग के जरिए उपयोग किया जा सकता है। इस प्रक्रिया को **ट्रांसफर लर्निंग** कहा जाता है। \n",
|
||||
"\n",
|
||||
"\n",
|
||||
"\n",
|
||||
"BERT, DistilBERT, BigBird, OpenGPT3 और अन्य जैसे ट्रांसफॉर्मर आर्किटेक्चर के कई प्रकार हैं, जिन्हें फाइन-ट्यून किया जा सकता है। \n",
|
||||
"\n",
|
||||
"आइए देखें कि हम प्री-ट्रेन किए गए BERT मॉडल का उपयोग करके अपनी पारंपरिक सीक्वेंस क्लासिफिकेशन समस्या को कैसे हल कर सकते हैं। हम [आधिकारिक डाक्यूमेंटेशन](https://www.tensorflow.org/text/tutorials/classify_text_with_bert) से विचार और कुछ कोड उधार लेंगे।\n",
|
||||
"\n",
|
||||
"प्री-ट्रेन किए गए मॉडल्स को लोड करने के लिए, हम **Tensorflow hub** का उपयोग करेंगे। सबसे पहले, आइए BERT-विशिष्ट वेक्टराइज़र लोड करें:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 1,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"ename": "ModuleNotFoundError",
|
||||
"evalue": "No module named 'tensorflow_text'",
|
||||
"output_type": "error",
|
||||
"traceback": [
|
||||
"\u001b[1;31m---------------------------------------------------------------------------\u001b[0m",
|
||||
"\u001b[1;31mModuleNotFoundError\u001b[0m Traceback (most recent call last)",
|
||||
"\u001b[1;32m~\\AppData\\Local\\Temp/ipykernel_41180/4216669875.py\u001b[0m in \u001b[0;36m<module>\u001b[1;34m\u001b[0m\n\u001b[1;32m----> 1\u001b[1;33m \u001b[1;32mimport\u001b[0m \u001b[0mtensorflow_text\u001b[0m\u001b[1;33m\u001b[0m\u001b[1;33m\u001b[0m\u001b[0m\n\u001b[0m\u001b[0;32m 2\u001b[0m \u001b[1;32mimport\u001b[0m \u001b[0mtensorflow_hub\u001b[0m \u001b[1;32mas\u001b[0m \u001b[0mhub\u001b[0m\u001b[1;33m\u001b[0m\u001b[1;33m\u001b[0m\u001b[0m\n\u001b[0;32m 3\u001b[0m \u001b[0mvectorizer\u001b[0m \u001b[1;33m=\u001b[0m \u001b[0mhub\u001b[0m\u001b[1;33m.\u001b[0m\u001b[0mKerasLayer\u001b[0m\u001b[1;33m(\u001b[0m\u001b[1;34m'https://tfhub.dev/tensorflow/bert_en_uncased_preprocess/3'\u001b[0m\u001b[1;33m)\u001b[0m\u001b[1;33m\u001b[0m\u001b[1;33m\u001b[0m\u001b[0m\n",
|
||||
"\u001b[1;31mModuleNotFoundError\u001b[0m: No module named 'tensorflow_text'"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"import tensorflow_text \n",
|
||||
"import tensorflow_hub as hub\n",
|
||||
"vectorizer = hub.KerasLayer('https://tfhub.dev/tensorflow/bert_en_uncased_preprocess/3')"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 7,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"{'input_type_ids': <tf.Tensor: shape=(1, 128), dtype=int32, numpy=\n",
|
||||
" array([[0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,\n",
|
||||
" 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,\n",
|
||||
" 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,\n",
|
||||
" 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,\n",
|
||||
" 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,\n",
|
||||
" 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]],\n",
|
||||
" dtype=int32)>,\n",
|
||||
" 'input_word_ids': <tf.Tensor: shape=(1, 128), dtype=int32, numpy=\n",
|
||||
" array([[ 101, 1045, 2293, 19081, 102, 0, 0, 0, 0,\n",
|
||||
" 0, 0, 0, 0, 0, 0, 0, 0, 0,\n",
|
||||
" 0, 0, 0, 0, 0, 0, 0, 0, 0,\n",
|
||||
" 0, 0, 0, 0, 0, 0, 0, 0, 0,\n",
|
||||
" 0, 0, 0, 0, 0, 0, 0, 0, 0,\n",
|
||||
" 0, 0, 0, 0, 0, 0, 0, 0, 0,\n",
|
||||
" 0, 0, 0, 0, 0, 0, 0, 0, 0,\n",
|
||||
" 0, 0, 0, 0, 0, 0, 0, 0, 0,\n",
|
||||
" 0, 0, 0, 0, 0, 0, 0, 0, 0,\n",
|
||||
" 0, 0, 0, 0, 0, 0, 0, 0, 0,\n",
|
||||
" 0, 0, 0, 0, 0, 0, 0, 0, 0,\n",
|
||||
" 0, 0, 0, 0, 0, 0, 0, 0, 0,\n",
|
||||
" 0, 0, 0, 0, 0, 0, 0, 0, 0,\n",
|
||||
" 0, 0, 0, 0, 0, 0, 0, 0, 0,\n",
|
||||
" 0, 0]], dtype=int32)>,\n",
|
||||
" 'input_mask': <tf.Tensor: shape=(1, 128), dtype=int32, numpy=\n",
|
||||
" array([[1, 1, 1, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,\n",
|
||||
" 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,\n",
|
||||
" 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,\n",
|
||||
" 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,\n",
|
||||
" 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,\n",
|
||||
" 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]],\n",
|
||||
" dtype=int32)>}"
|
||||
]
|
||||
},
|
||||
"execution_count": 7,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"vectorizer(['I love transformers'])"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"यह ज़रूरी है कि आप वही vectorizer इस्तेमाल करें जो मूल नेटवर्क पर ट्रेनिंग के दौरान उपयोग किया गया था। साथ ही, BERT vectorizer तीन घटक लौटाता है:\n",
|
||||
"* `input_word_ids`, जो इनपुट वाक्य के लिए टोकन नंबरों का अनुक्रम है\n",
|
||||
"* `input_mask`, जो दिखाता है कि अनुक्रम का कौन सा हिस्सा वास्तविक इनपुट है और कौन सा padding है। यह `Masking` लेयर द्वारा बनाए गए मास्क के समान है\n",
|
||||
"* `input_type_ids` भाषा मॉडलिंग कार्यों के लिए उपयोग किया जाता है, और एक अनुक्रम में दो इनपुट वाक्यों को निर्दिष्ट करने की अनुमति देता है।\n",
|
||||
"\n",
|
||||
"इसके बाद, हम BERT फीचर एक्सट्रैक्टर को instantiate कर सकते हैं:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 8,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"bert = hub.KerasLayer('https://tfhub.dev/tensorflow/small_bert/bert_en_uncased_L-4_H-128_A-2/1')"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 9,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"pooled_output -> (1, 128)\n",
|
||||
"encoder_outputs -> 4\n",
|
||||
"sequence_output -> (1, 128, 128)\n",
|
||||
"default -> (1, 128)\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"z = bert(vectorizer(['I love transformers']))\n",
|
||||
"for i,x in z.items():\n",
|
||||
" print(f\"{i} -> { len(x) if isinstance(x, list) else x.shape }\")"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"तो, BERT लेयर कई उपयोगी परिणाम लौटाती है:\n",
|
||||
"* `pooled_output` पूरे अनुक्रम के सभी टोकन का औसत निकालने का परिणाम है। इसे पूरे नेटवर्क का एक बुद्धिमान अर्थपूर्ण एम्बेडिंग माना जा सकता है। यह हमारे पिछले मॉडल में `GlobalAveragePooling1D` लेयर के आउटपुट के समकक्ष है।\n",
|
||||
"* `sequence_output` अंतिम ट्रांसफॉर्मर लेयर का आउटपुट है (जो हमारे ऊपर दिए गए मॉडल में `TransformerBlock` के आउटपुट के अनुरूप है)।\n",
|
||||
"* `encoder_outputs` सभी ट्रांसफॉर्मर लेयर्स के आउटपुट हैं। चूंकि हमने 4-लेयर BERT मॉडल लोड किया है (जैसा कि आप शायद नाम से अनुमान लगा सकते हैं, जिसमें `4_H` शामिल है), इसमें 4 टेन्सर हैं। अंतिम टेन्सर `sequence_output` के समान है।\n",
|
||||
"\n",
|
||||
"अब हम एंड-टू-एंड क्लासिफिकेशन मॉडल को परिभाषित करेंगे। हम *फंक्शनल मॉडल डिफिनिशन* का उपयोग करेंगे, जिसमें हम मॉडल का इनपुट परिभाषित करेंगे और फिर इसके आउटपुट की गणना के लिए एक श्रृंखला में अभिव्यक्तियाँ प्रदान करेंगे। हम BERT मॉडल के वेट्स को ट्रेन नहीं करेंगे और केवल अंतिम क्लासिफायर को ट्रेन करेंगे:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 10,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Model: \"model\"\n",
|
||||
"__________________________________________________________________________________________________\n",
|
||||
"Layer (type) Output Shape Param # Connected to \n",
|
||||
"==================================================================================================\n",
|
||||
"input_1 (InputLayer) [(None,)] 0 \n",
|
||||
"__________________________________________________________________________________________________\n",
|
||||
"keras_layer (KerasLayer) {'input_type_ids': ( 0 input_1[0][0] \n",
|
||||
"__________________________________________________________________________________________________\n",
|
||||
"keras_layer_1 (KerasLayer) {'pooled_output': (N 4782465 keras_layer[0][0] \n",
|
||||
" keras_layer[0][1] \n",
|
||||
" keras_layer[0][2] \n",
|
||||
"__________________________________________________________________________________________________\n",
|
||||
"dropout_4 (Dropout) (None, 128) 0 keras_layer_1[0][5] \n",
|
||||
"__________________________________________________________________________________________________\n",
|
||||
"dense_4 (Dense) (None, 4) 516 dropout_4[0][0] \n",
|
||||
"==================================================================================================\n",
|
||||
"Total params: 4,782,981\n",
|
||||
"Trainable params: 516\n",
|
||||
"Non-trainable params: 4,782,465\n",
|
||||
"__________________________________________________________________________________________________\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"inp = keras.Input(shape=(),dtype=tf.string)\n",
|
||||
"x = vectorizer(inp)\n",
|
||||
"x = bert(x)\n",
|
||||
"x = keras.layers.Dropout(0.1)(x['pooled_output'])\n",
|
||||
"out = keras.layers.Dense(4,activation='softmax')(x)\n",
|
||||
"model = keras.models.Model(inp,out)\n",
|
||||
"bert.trainable = False\n",
|
||||
"model.summary()"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 11,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"938/938 [==============================] - 528s 559ms/step - loss: 0.8056 - acc: 0.6983 - val_loss: 0.5953 - val_acc: 0.7888\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"<tensorflow.python.keras.callbacks.History at 0x7f9bb1e36d00>"
|
||||
]
|
||||
},
|
||||
"execution_count": 11,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"model.compile(loss='sparse_categorical_crossentropy',metrics=['acc'], optimizer='adam')\n",
|
||||
"model.fit(ds_train.map(tupelize).batch(128),validation_data=ds_test.map(tupelize).batch(128))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"हालांकि ट्रेन करने योग्य पैरामीटर बहुत कम हैं, प्रक्रिया काफी धीमी है क्योंकि BERT फीचर एक्सट्रैक्टर गणनात्मक रूप से भारी है। ऐसा लगता है कि हम उचित सटीकता प्राप्त करने में असमर्थ रहे, या तो प्रशिक्षण की कमी के कारण, या मॉडल पैरामीटर की कमी के कारण।\n",
|
||||
"\n",
|
||||
"आइए BERT वेट्स को अनफ्रीज़ करें और इसे भी ट्रेन करें। इसके लिए बहुत छोटे लर्निंग रेट की आवश्यकता होती है, और साथ ही **वार्मअप** के साथ अधिक सावधानीपूर्वक प्रशिक्षण रणनीति की आवश्यकता होती है, जिसमें **AdamW** ऑप्टिमाइज़र का उपयोग किया जाता है। हम `tf-models-official` पैकेज का उपयोग करके ऑप्टिमाइज़र बनाएंगे:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 12,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Model: \"model\"\n",
|
||||
"__________________________________________________________________________________________________\n",
|
||||
"Layer (type) Output Shape Param # Connected to \n",
|
||||
"==================================================================================================\n",
|
||||
"input_1 (InputLayer) [(None,)] 0 \n",
|
||||
"__________________________________________________________________________________________________\n",
|
||||
"keras_layer (KerasLayer) {'input_type_ids': ( 0 input_1[0][0] \n",
|
||||
"__________________________________________________________________________________________________\n",
|
||||
"keras_layer_1 (KerasLayer) {'pooled_output': (N 4782465 keras_layer[0][0] \n",
|
||||
" keras_layer[0][1] \n",
|
||||
" keras_layer[0][2] \n",
|
||||
"__________________________________________________________________________________________________\n",
|
||||
"dropout_4 (Dropout) (None, 128) 0 keras_layer_1[0][5] \n",
|
||||
"__________________________________________________________________________________________________\n",
|
||||
"dense_4 (Dense) (None, 4) 516 dropout_4[0][0] \n",
|
||||
"==================================================================================================\n",
|
||||
"Total params: 4,782,981\n",
|
||||
"Trainable params: 4,782,980\n",
|
||||
"Non-trainable params: 1\n",
|
||||
"__________________________________________________________________________________________________\n",
|
||||
"938/938 [==============================] - 629s 664ms/step - loss: 0.6344 - acc: 0.7658 - val_loss: 0.4876 - val_acc: 0.8247\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"<tensorflow.python.keras.callbacks.History at 0x7f9bb0bd0070>"
|
||||
]
|
||||
},
|
||||
"execution_count": 12,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"from official.nlp import optimization \n",
|
||||
"bert.trainable=True\n",
|
||||
"model.summary()\n",
|
||||
"epochs = 3\n",
|
||||
"opt = optimization.create_optimizer(\n",
|
||||
" init_lr=3e-5,\n",
|
||||
" num_train_steps=epochs*len(ds_train),\n",
|
||||
" num_warmup_steps=0.1*epochs*len(ds_train),\n",
|
||||
" optimizer_type='adamw')\n",
|
||||
"\n",
|
||||
"model.compile(loss='sparse_categorical_crossentropy',metrics=['acc'], optimizer=opt)\n",
|
||||
"model.fit(ds_train.map(tupelize).batch(128),validation_data=ds_test.map(tupelize).batch(128))"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"जैसा कि आप देख सकते हैं, प्रशिक्षण काफी धीमी गति से होता है - लेकिन आप कुछ epochs (5-10) के लिए मॉडल को प्रशिक्षित करने का प्रयास कर सकते हैं और देख सकते हैं कि क्या आप पहले उपयोग किए गए तरीकों की तुलना में सबसे अच्छा परिणाम प्राप्त कर सकते हैं।\n",
|
||||
"\n",
|
||||
"## Huggingface Transformers लाइब्रेरी\n",
|
||||
"\n",
|
||||
"Transformer मॉडल का उपयोग करने का एक और बहुत सामान्य (और थोड़ा सरल) तरीका [HuggingFace पैकेज](https://github.com/huggingface/) है, जो विभिन्न NLP कार्यों के लिए सरल बिल्डिंग ब्लॉक्स प्रदान करता है। यह Tensorflow और PyTorch, एक अन्य बहुत लोकप्रिय न्यूरल नेटवर्क फ्रेमवर्क, दोनों के लिए उपलब्ध है।\n",
|
||||
"\n",
|
||||
"> **Note**: यदि आप यह देखने में रुचि नहीं रखते कि Transformers लाइब्रेरी कैसे काम करती है - तो आप इस नोटबुक के अंत तक जा सकते हैं, क्योंकि आप ऊपर किए गए कार्यों से कुछ भी मौलिक रूप से अलग नहीं देखेंगे। हम BERT मॉडल को प्रशिक्षित करने के उन्हीं चरणों को दोहराएंगे, लेकिन एक अलग लाइब्रेरी और काफी बड़े मॉडल का उपयोग करेंगे। इसलिए, प्रक्रिया में कुछ लंबा प्रशिक्षण शामिल है, तो आप केवल कोड को देख सकते हैं।\n",
|
||||
"\n",
|
||||
"आइए देखें कि हमारा समस्या [Huggingface Transformers](http://huggingface.co) का उपयोग करके कैसे हल की जा सकती है।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"सबसे पहले हमें उस मॉडल को चुनना होगा जिसे हम उपयोग करने जा रहे हैं। कुछ बिल्ट-इन मॉडल्स के अलावा, Huggingface में एक [ऑनलाइन मॉडल रिपॉजिटरी](https://huggingface.co/models) भी है, जहां आपको समुदाय द्वारा बनाए गए और भी कई प्री-ट्रेंड मॉडल मिल सकते हैं। इन सभी मॉडलों को केवल मॉडल का नाम देकर लोड और उपयोग किया जा सकता है। मॉडल के लिए आवश्यक सभी बाइनरी फाइल्स स्वचालित रूप से डाउनलोड हो जाएंगी।\n",
|
||||
"\n",
|
||||
"कुछ स्थितियों में आपको अपने खुद के मॉडल लोड करने की आवश्यकता हो सकती है। ऐसे मामलों में, आप उस डायरेक्टरी को निर्दिष्ट कर सकते हैं जिसमें सभी संबंधित फाइल्स मौजूद हों, जैसे कि टोकनाइज़र के पैरामीटर, `config.json` फाइल जिसमें मॉडल पैरामीटर हों, बाइनरी वेट्स आदि।\n",
|
||||
"\n",
|
||||
"मॉडल के नाम से, हम मॉडल और टोकनाइज़र दोनों को इंस्टैंशिएट कर सकते हैं। चलिए टोकनाइज़र से शुरू करते हैं:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 2,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"import transformers\n",
|
||||
"\n",
|
||||
"# To load the model from Internet repository using model name. \n",
|
||||
"# Use this if you are running from your own copy of the notebooks\n",
|
||||
"bert_model = 'bert-base-uncased' \n",
|
||||
"\n",
|
||||
"# To load the model from the directory on disk. Use this for Microsoft Learn module, because we have\n",
|
||||
"# prepared all required files for you.\n",
|
||||
"#bert_model = './bert'\n",
|
||||
"\n",
|
||||
"tokenizer = transformers.BertTokenizer.from_pretrained(bert_model)\n",
|
||||
"\n",
|
||||
"MAX_SEQ_LEN = 128\n",
|
||||
"PAD_INDEX = tokenizer.convert_tokens_to_ids(tokenizer.pad_token)\n",
|
||||
"UNK_INDEX = tokenizer.convert_tokens_to_ids(tokenizer.unk_token)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"`tokenizer` ऑब्जेक्ट में `encode` फ़ंक्शन होता है जिसे सीधे टेक्स्ट को एन्कोड करने के लिए उपयोग किया जा सकता है:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 3,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"[101, 23435, 12314, 2003, 1037, 2307, 7705, 2005, 17953, 2361, 102]"
|
||||
]
|
||||
},
|
||||
"execution_count": 3,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"tokenizer.encode('Tensorflow is a great framework for NLP')"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"हम टोकनाइज़र का उपयोग अनुक्रम को इस प्रकार एन्कोड करने के लिए भी कर सकते हैं जो मॉडल को पास करने के लिए उपयुक्त हो, जैसे `token_ids`, `input_mask` फ़ील्ड्स आदि शामिल करना। हम यह भी निर्दिष्ट कर सकते हैं कि हम Tensorflow टेन्सर चाहते हैं, इसके लिए `return_tensors='tf'` तर्क प्रदान कर सकते हैं:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"{'input_ids': <tf.Tensor: shape=(1, 5), dtype=int32, numpy=array([[ 101, 7592, 1010, 2045, 102]], dtype=int32)>, 'token_type_ids': <tf.Tensor: shape=(1, 5), dtype=int32, numpy=array([[0, 0, 0, 0, 0]], dtype=int32)>, 'attention_mask': <tf.Tensor: shape=(1, 5), dtype=int32, numpy=array([[1, 1, 1, 1, 1]], dtype=int32)>}"
|
||||
]
|
||||
},
|
||||
"execution_count": 4,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"tokenizer(['Hello, there'],return_tensors='tf')"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"हमारे मामले में, हम एक पूर्व-प्रशिक्षित BERT मॉडल का उपयोग करेंगे जिसे `bert-base-uncased` कहा जाता है। *Uncased* का मतलब है कि यह मॉडल केस-सेंसिटिव नहीं है।\n",
|
||||
"\n",
|
||||
"मॉडल को प्रशिक्षित करते समय, हमें टोकनाइज़ किए गए अनुक्रम को इनपुट के रूप में प्रदान करना होता है, और इसलिए हम डेटा प्रोसेसिंग पाइपलाइन डिज़ाइन करेंगे। चूंकि `tokenizer.encode` एक Python फ़ंक्शन है, हम इसे पिछले यूनिट की तरह ही उपयोग करेंगे, जिसमें इसे `py_function` का उपयोग करके कॉल किया जाएगा:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 31,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"def process(x):\n",
|
||||
" return tokenizer.encode(x.numpy().decode('utf-8'),return_tensors='tf',padding='max_length',max_length=MAX_SEQ_LEN,truncation=True)[0]\n",
|
||||
"\n",
|
||||
"def process_fn(x):\n",
|
||||
" s = x['title']+' '+x['description']\n",
|
||||
" e = tf.py_function(process,inp=[s],Tout=(tf.int32))\n",
|
||||
" e.set_shape(MAX_SEQ_LEN)\n",
|
||||
" return e,x['label']"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"अब हम `BertForSequenceClassification` पैकेज का उपयोग करके वास्तविक मॉडल लोड कर सकते हैं। यह सुनिश्चित करता है कि हमारे मॉडल में पहले से ही वर्गीकरण के लिए आवश्यक संरचना है, जिसमें अंतिम वर्गीकर्ता भी शामिल है। आपको एक चेतावनी संदेश दिखाई देगा जिसमें कहा जाएगा कि अंतिम वर्गीकर्ता के वज़न प्रारंभिक नहीं किए गए हैं, और मॉडल को पूर्व-प्रशिक्षण की आवश्यकता होगी - यह पूरी तरह से ठीक है, क्योंकि यही हम करने वाले हैं!\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 32,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"model = transformers.TFBertForSequenceClassification.from_pretrained(bert_model,num_labels=4,output_attentions=False)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 33,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Model: \"tf_bert_for_sequence_classification_1\"\n",
|
||||
"_________________________________________________________________\n",
|
||||
"Layer (type) Output Shape Param # \n",
|
||||
"=================================================================\n",
|
||||
"bert (TFBertMainLayer) multiple 109482240 \n",
|
||||
"_________________________________________________________________\n",
|
||||
"dropout_75 (Dropout) multiple 0 \n",
|
||||
"_________________________________________________________________\n",
|
||||
"classifier (Dense) multiple 3076 \n",
|
||||
"=================================================================\n",
|
||||
"Total params: 109,485,316\n",
|
||||
"Trainable params: 109,485,316\n",
|
||||
"Non-trainable params: 0\n",
|
||||
"_________________________________________________________________\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"model.summary()"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"जैसा कि आप `summary()` से देख सकते हैं, मॉडल में लगभग 110 मिलियन पैरामीटर हैं! संभवतः, यदि हम अपेक्षाकृत छोटे डेटासेट पर सरल वर्गीकरण कार्य करना चाहते हैं, तो हम BERT बेस लेयर को प्रशिक्षित नहीं करना चाहेंगे:\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 34,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"Model: \"tf_bert_for_sequence_classification_1\"\n",
|
||||
"_________________________________________________________________\n",
|
||||
"Layer (type) Output Shape Param # \n",
|
||||
"=================================================================\n",
|
||||
"bert (TFBertMainLayer) multiple 109482240 \n",
|
||||
"_________________________________________________________________\n",
|
||||
"dropout_75 (Dropout) multiple 0 \n",
|
||||
"_________________________________________________________________\n",
|
||||
"classifier (Dense) multiple 3076 \n",
|
||||
"=================================================================\n",
|
||||
"Total params: 109,485,316\n",
|
||||
"Trainable params: 3,076\n",
|
||||
"Non-trainable params: 109,482,240\n",
|
||||
"_________________________________________________________________\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"model.layers[0].trainable = False\n",
|
||||
"model.summary()"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"अब हम प्रशिक्षण शुरू करने के लिए तैयार हैं!\n",
|
||||
"\n",
|
||||
"> **नोट**: पूर्ण-स्तरीय BERT मॉडल का प्रशिक्षण करना बहुत समय लेने वाला हो सकता है! इसलिए हम इसे केवल पहले 32 बैचों के लिए प्रशिक्षित करेंगे। यह केवल यह दिखाने के लिए है कि मॉडल प्रशिक्षण कैसे सेट किया जाता है। यदि आप पूर्ण-स्तरीय प्रशिक्षण आज़माने में रुचि रखते हैं - तो बस `steps_per_epoch` और `validation_steps` पैरामीटर हटा दें, और इंतजार करने के लिए तैयार रहें!\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": 30,
|
||||
"metadata": {},
|
||||
"outputs": [
|
||||
{
|
||||
"name": "stdout",
|
||||
"output_type": "stream",
|
||||
"text": [
|
||||
"32/32 [==============================] - 142s 4s/step - loss: 1.3896 - acc: 0.2500 - val_loss: 1.3863 - val_acc: 0.2480\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"data": {
|
||||
"text/plain": [
|
||||
"<tensorflow.python.keras.callbacks.History at 0x7f1d40a4b6a0>"
|
||||
]
|
||||
},
|
||||
"execution_count": 30,
|
||||
"metadata": {},
|
||||
"output_type": "execute_result"
|
||||
}
|
||||
],
|
||||
"source": [
|
||||
"model.compile('adam','sparse_categorical_crossentropy',['acc'])\n",
|
||||
"tf.get_logger().setLevel('ERROR')\n",
|
||||
"model.fit(ds_train.map(process_fn).batch(32),validation_data=ds_test.map(process_fn).batch(32),steps_per_epoch=32,validation_steps=2)"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"यदि आप iterations की संख्या बढ़ाते हैं, पर्याप्त समय तक प्रतीक्षा करते हैं, और कई epochs तक प्रशिक्षण करते हैं, तो आप उम्मीद कर सकते हैं कि BERT classification हमें सबसे अच्छी सटीकता प्रदान करेगा! इसका कारण यह है कि BERT पहले से ही भाषा की संरचना को काफी अच्छी तरह समझता है, और हमें केवल अंतिम classifier को fine-tune करने की आवश्यकता होती है। हालांकि, क्योंकि BERT एक बड़ा मॉडल है, पूरा प्रशिक्षण प्रक्रिया काफी समय लेती है और इसके लिए गंभीर computational शक्ति की आवश्यकता होती है! (GPU, और अधिमानतः एक से अधिक).\n",
|
||||
"\n",
|
||||
"> **Note:** हमारे उदाहरण में, हमने सबसे छोटे pre-trained BERT मॉडल में से एक का उपयोग किया है। बड़े मॉडल उपलब्ध हैं जो संभवतः बेहतर परिणाम प्रदान कर सकते हैं।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## मुख्य बातें\n",
|
||||
"\n",
|
||||
"इस यूनिट में, हमने **ट्रांसफॉर्मर्स** पर आधारित हाल ही की मॉडल आर्किटेक्चर देखी हैं। हमने इन्हें अपने टेक्स्ट वर्गीकरण कार्य के लिए लागू किया है, लेकिन इसी तरह, BERT मॉडल का उपयोग एंटिटी एक्सट्रैक्शन, प्रश्न उत्तर देने और अन्य NLP कार्यों के लिए भी किया जा सकता है।\n",
|
||||
"\n",
|
||||
"ट्रांसफॉर्मर मॉडल NLP में वर्तमान में सबसे उन्नत तकनीक का प्रतिनिधित्व करते हैं, और अधिकांश मामलों में, यह वह पहला समाधान होना चाहिए जिसके साथ आप कस्टम NLP समाधान लागू करते समय प्रयोग करना शुरू करें। हालांकि, यदि आप उन्नत न्यूरल मॉडल बनाना चाहते हैं, तो इस मॉड्यूल में चर्चा किए गए पुनरावर्ती न्यूरल नेटवर्क के मूलभूत सिद्धांतों को समझना अत्यंत महत्वपूर्ण है।\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"execution_count": null,
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": []
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"\n---\n\n**अस्वीकरण**: \nयह दस्तावेज़ AI अनुवाद सेवा [Co-op Translator](https://github.com/Azure/co-op-translator) का उपयोग करके अनुवादित किया गया है। जबकि हम सटीकता सुनिश्चित करने का प्रयास करते हैं, कृपया ध्यान दें कि स्वचालित अनुवाद में त्रुटियां या अशुद्धियां हो सकती हैं। मूल भाषा में उपलब्ध मूल दस्तावेज़ को प्रामाणिक स्रोत माना जाना चाहिए। महत्वपूर्ण जानकारी के लिए, पेशेवर मानव अनुवाद की सिफारिश की जाती है। इस अनुवाद के उपयोग से उत्पन्न किसी भी गलतफहमी या गलत व्याख्या के लिए हम जिम्मेदार नहीं हैं।\n"
|
||||
]
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
"interpreter": {
|
||||
"hash": "0cb620c6d4b9f7a635928804c26cf22403d89d98d79684e4529119355ee6d5a5"
|
||||
},
|
||||
"kernelspec": {
|
||||
"display_name": "py38_tensorflow",
|
||||
"language": "python",
|
||||
"name": "python3"
|
||||
},
|
||||
"language_info": {
|
||||
"codemirror_mode": {
|
||||
"name": "ipython",
|
||||
"version": 3
|
||||
},
|
||||
"file_extension": ".py",
|
||||
"mimetype": "text/x-python",
|
||||
"name": "python",
|
||||
"nbconvert_exporter": "python",
|
||||
"pygments_lexer": "ipython3",
|
||||
"version": "3.8.12"
|
||||
},
|
||||
"coopTranslator": {
|
||||
"original_hash": "ab59c532409774988ab875f2260e8e53",
|
||||
"translation_date": "2025-08-31T15:19:22+00:00",
|
||||
"source_file": "lessons/5-NLP/18-Transformers/TransformersTF.ipynb",
|
||||
"language_code": "hi"
|
||||
}
|
||||
},
|
||||
"nbformat": 4,
|
||||
"nbformat_minor": 4
|
||||
}
|
||||
Some files were not shown because too many files have changed in this diff Show More
Loading…
Reference in New Issue