Machine Learning with PySpark: With Natural Language Processing and Recommender Systems

1h 50m
Pramod Singh
Apress
2019

Build machine learning models, natural language processing applications, and recommender systems with PySpark to solve various business challenges. This book starts with the fundamentals of Spark and its evolution and then covers the entire spectrum of traditional machine learning algorithms along with natural language processing and recommender systems using PySpark.

Machine Learning with PySpark shows you how to build supervised machine learning models such as linear regression, logistic regression, decision trees, and random forest. You’ll also see unsupervised machine learning models such as K-means and hierarchical clustering. A major portion of the book focuses on feature engineering to create useful features with PySpark to train the machine learning models. The natural language processing section covers text processing, text mining, and embedding for classification.

After reading this book, you will understand how to use PySpark’s machine learning library to build and train various machine learning models. Additionally you’ll become comfortable with related PySpark components, such as data ingestion, data processing, and data analysis, that you can use to develop data-driven intelligent applications.

What You Will Learn

Build a spectrum of supervised and unsupervised machine learning algorithms
Implement machine learning algorithms with Spark MLlib libraries
Develop a recommender system with Spark MLlib libraries
Handle issues related to feature engineering, class balance, bias and variance, and cross validation for building an optimal fit model

Who This Book Is For

Data science and machine learning professionals.

About the Author

Pramod Singh is an established data scientist with over eight years of experience in data and solving business challenges. He has worked in organizations such as Infosys, Tally and SapientRazorfish. Also, president of a data science meet-up group and regular speaker at various webinars. Recently spoke at major conference: GIDS 2018 and presented a session on “Sequence Embedding in Spark” which was well received. He has an online Udemy course on machine learning.

In this Book

Evolution of Data
Introduction to Machine Learning
Data Processing
Linear Regression
Logistic Regression
Random Forests
Recommender Systems
Clustering
Natural Language Processing

FREE ACCESS

Book Pythonic AI: A Beginner's Guide to Building AI Applications in Python

Journey Prompt Engineering for Programmers to Learn Python

(1)

Book Machine Learning: A Constraint-Based Approach, Second Edition

Get Started

Sharpen your skills. Upgrade your career. Find the right learning path for you, based on your role and skills. Take part in hands-on practice, study for a certification, and much more - all personalized for you.

*Not included: Compliance, Leadership Development Program content, and Engineering books

Your content + our content + our platform = a path to learning success

Using our learning experience platform, Percipio, your learners can engage in custom learning paths that can feature curated content from all sources.

Learn More

Aspire to something bigger

Aspire Journeys are guided learning paths that set you in motion for career success.

Browse Aspire Journeys

Explore a world of live learning with Global Knowledge

Choose from convenient delivery formats to get the training you and your team need - where, when and how you want it.

Browse Live Learning

IT Skills & Salary Report

ESG Impact Report

Machine Learning with PySpark: With Natural Language Processing and Recommender Systems

In this Book

YOU MIGHT ALSO LIKE