Apache Spark and Scala Certification Training

BY
Simplilearn

Enrol to become proficient in the primary concepts and principles of the programming languages Spark and Scala.

Mode

Online

Quick Facts

particular details
Medium of instructions English
Mode of learning Self study, Virtual Classroom
Mode of Delivery Video and Text Based

Course overview

The Apache Spark and Scala online training is designed to get students comfortable with the industry-standard platforms that are Apache Spark and Scala. This  Apache Spark and Scala Certification Training course by Edureka takes the students through a variety of Spark concepts such as Spark Streaming, machine learning programming, GraphX programming, Spark SQL, and Shell Scripting in Spark. 

Expand your skill-set of the Big Data Hadoop frameworks with the Apache Spark and Scala certification training course, which is developed to help candidates attain expertise in crucial Apache Spark skills, providing the candidates with a competitive edge for a career in Big Data.

Moreover, the Apache Spark and Scala training course also help the students in building proficiency in the Scala programming language and Scala topics like data types, basic literals, operators, functions, objects, classes, lists, maps, and data structures.

In addition, the Apache Spark and Scala course will make the candidates ready for any Big Data-based industry role that requires proficiency in Spark and Scala.  Apache Spark and Scala Certification Training course by Edureka also helps the students prepare for the CCA175 Spark and Hadoop Developer exam to become a Spark Developer.

The highlights

  • Online classroom learning
  • Online self-paced learning
  • Certificate of completion from Simplilearn
  • Lifelong Certificate validity

Program offerings

  • Scala programming language
  • Spark installation process
  • Graphx programming
  • Free practice test
  • 24*7 learning support

Course and certificate fees

certificate availability

Yes

certificate providing authority

Simplilearn

Who it is for

The Apache Spark and Scala Certification is a perfect fit for professionals such as –

  • IT developers and testers: Developers interested in Scala programming
  • Professionals working in the Analytics domain who wish to upskill
  • Professionals currently working as Data scientists 
  • Students currently pursuing their UG/PG in a related field 

Eligibility criteria

Skills

There are no prerequisites to join the course. However, candidates who wish to join Apache Spark and Scala should have a fundamental knowledge of any programming language, database, SQL, and query language for databases. In addition, they should have a working knowledge of Linux- or Unix-based systems.

What you will learn

Knowledge of apache spark

Candidates will also gain proficiency in the Scala programming language. In the Apache Spark and Scala course, candidates will learn the following skillset:

  • Gain proficiency in the Scala programming language
  • Learn how to install Spark on their system of choice
  • Tune the settings and preferences in Spark to optimize workflow
  • know the various streaming features of Spark like Integration, Combination, Load balancing, resource usage, and recovery from stragglers and failures.
  • Gain knowledge about the machine learning package recently added to the Spark library
  • Use Spark ML is to provide high-level APIs and create machine learning pipelines
  • Attain a working knowledge of Spark SQL
  • Understand the Spark SQL architecture, handle various data formats, implement Data Frame operations, and process DataFrames.
  • Learn about the resilient Distributed Dataset (RDDs) and its applications
  • Learn how to create RDDs, RDD operations, and RDD extensions

The syllabus

Apache Spark & Scala

Module 1 - Course Overview
  • Introduction
  • Course Objectives
  • Course Overview
  • Target Audience
  • Course Prerequisites
  • Value to the Professionals
  • Value to the Professionals (contd.)
  • Value to the Professionals (contd.)
  • Lessons Covered
  • Conclusion
Module 2 - Introduction to Spark
  • Introduction
  • Objectives
  • Evolution of Distributed Systems
  • Need of New Generation Distributed Systems
  • Limitations of MapReduce in Hadoop
  • Limitations of MapReduce in Hadoop (contd.)
  • Batch vs. Real-Time Processing
  • PairRDD Methods – Others
  • Application of In-Memory Processing
  • Introduction to Apache Spark
  • Components of a Spark Project
  • History of Spark
  • Language Flexibility in Spark
  • Spark Execution Architecture
  • Automatic Parallelization of Complex Flows
  • Automatic Parallelization of Complex Flows – Important Points
  • APIs That Match User Goals
  • Apache Spark – A Unified Platform of Big Data Apps
  • More Benefits of Apache Spark
  • Running Spark in Different Modes
  • Installing Spark as a Standalone Cluster – Configurations
  • Installing Spark as a Standalone Cluster – Configurations
  • Demo – Install Apache Spark
  • Demo – Install Apache Spark
  • Overview of Spark on a Cluster
  • Tasks of Spark on a Cluster
  • Companies Using Spark – Use Cases
  • Hadoop Ecosystem vs. Apache Spark
  • Hadoop Ecosystem vs. Apache Spark (contd.)
  • Quiz
  • Summary
  • Summary (contd.)
  • Conclusion
Module 3 - Introduction to Programming in Scala
  • Introduction
  • Objectives
  • Introduction to Scala
  • Features of Scala
  • Basic Data Types
  • Basic Literals
  • Basic Literals (contd.)
  • Basic Literals (contd.)
  • Introduction to Operators
  • Types of Operators
  • Use Basic Literals and the Arithmetic Operator
  • Demo – Use Basic Literals and the Arithmetic Operator
  • Use the Logical Operator
  • Demo – Use the Logical Operator
  • Introduction to Type Inference
  • Type Inference for Recursive Methods
  • Type Inference for Polymorphic Methods and Generic Classes
  • Unreliability on Type Inference Mechanism
  • Mutable Collection vs. Immutable Collection
  • Functions
  • Anonymous Functions
  • Objects
  • Classes
  • Use Type Inference, Functions, Anonymous Function, and Class
  • Demo – Use Type Inference, Functions, Anonymous Function and Class
  • Traits as Interfaces
  • Traits – Example
  • Collections
  • Types of Collections
  • Types of Collections (contd.)
  • Lists
  • Perform Operations on Lists
  • Demo – Use Data Structures
  • Maps
  • Maps – Operations
  • Pattern Matching
  • Implicits
  • Implicits (contd.)
  • Streams
  • Use Data Structures
  • Demo – Perform Operations on Lists
  • Quiz
  • Summary
  • Summary (contd.)
  • Conclusion
Module 4 - Using RDD for Creating Applications in Spark
  • Introduction
  • Objectives
  • RDDs API
  • Features of RDDs
  • Creating RDDs
  • Creating RDDs – Referencing an External Dataset
  • Referencing an External Dataset – Text Files
  • Referencing an External Dataset – Text Files (contd.)
  • Referencing an External Dataset – Sequence Files
  • Referencing an External Dataset – Other Hadoop Input Formats
  • Creating RDDs – Important Points
  • RDD Operations
  • RDD Operations – Transformations
  • Features of RDD Persistence
  • Storage Levels of RDD Persistence
  • Choosing the Correct RDD Persistence Storage Level
  • Invoking the Spark Shell
  • Importing Spark Classes
  • Creating the SparkContext
  • Loading a File in Shell
  • Performing Some Basic Operations on Files in Spark Shell RDDs
  • Packaging a Spark Project with SBT
  • Running a Spark Project with SBT
  • Demo – Build a Scala Project
  • Build a Scala Project
  • Demo – Build a Spark Java Project
  • Build a Spark Java Project
  • Shared Variables – Broadcast
  • Shared Variables – Accumulators
  • Writing a Scala Application
  • Demo – Run a Scala Application
  • Run a Scala Application
  • Demo – Write a Scala Application Reading the Hadoop Data
  • Write a Scala Application Reading the Hadoop Data
  • Demo – Run a Scala Application Reading the Hadoop Data
  • Run a Scala Application Reading the Hadoop Data
  • Scala RDD Extensions
  • DoubleRDD Methods
  • PairRDD Methods – Join
  • PairRDD Methods – Others
  • Java PairRDD Methods
  • Java PairRDD Methods (contd.)
  • General RDD Methods
  • General RDD Methods (contd.)
  • Java RDD Methods
  • Java RDD Methods (contd.)
  • Common Java RDD Methods
  • Spark Java Function Classes
  • Method for Combining JavaPairRDD Functions
  • Transformations in RDD
  • Other Methods
  • Actions in RDD
  • Key-Value Pair RDD in Scala
  • Key-Value Pair RDD in Java
  • Using MapReduce and Pair RDD Operations
  • Reading Text File from HDFS
  • Reading Sequence File from HDFS
  • Writing Text Data to HDFS
  • Writing Sequence File to HDFS
  • Using GroupBy
  • Using GroupBy (contd.)
  • Demo – Run a Scala Application Performing GroupBy Operation
  • Run a Scala Application Performing GroupBy Operation
  • Demo – Run a Scala Application Using the Scala Shell
  • Run a Scala Application Using the Scala Shell
  • Demo – Write and Run a Java Application
  • Write and Run a Java Application
  • Quiz
  • Summary
  • Summary (contd.)
  • Conclusion
Module 5 - Running SQL queries Using SparkSQL
  • Introduction
  • Objectives
  • Importance of Spark SQL
  • Benefits of Spark SQL
  • DataFrames
  • SQLContext
  • SQLContext (contd.)
  • Creating a DataFrame
  • Using DataFrame Operations
  • Using DataFrame Operations (contd.)
  • Demo – Run Spark SQL with a DataFrame
  • Run Spark SQL with a DataFrame
  • Interoperating with RDDs
  • Using the Reflection-Based Approach
  • Using the Reflection-Based Approach (contd.)
  • Using the Programmatic Approach
  • Using the Programmatic Approach (contd.)
  • Demo – Run Spark SQL Programmatically
  • Run Spark SQL Programmatically
  • Data Sources
  • Save Modes
  • Saving to Persistent Tables
  • Parquet Files
  • Partition Discovery
  • Schema Merging
  • JSON Data
  • Hive Table
  • DML Operation – Hive Queries
  • Demo – Run Hive Queries Using Spark SQL
  • Run Hive Queries Using Spark SQL
  • JDBC to Other Databases
  • Supported Hive Features
  • Supported Hive Features (contd.)
  • Supported Hive Data Types
  • Case Classes
  • Case Classes (contd.)
  • Quiz
  • Summary
  • Summary (contd.)
  • Conclusion
Module 6 - Spark Streaming


  • Introduction
  • Objectives
  • Introduction to Spark Streaming
  • Working of Spark Streaming
  • Features of Spark Streaming
  • Streaming Word Count
  • Micro Batch
  • DStreams
  • DStreams (contd.)
  • Input DStreams and Receivers
  • Input DStreams and Receivers (contd.)
  • Basic Sources
  • Advanced Sources
  • Advanced Sources – Twitter
  • Transformations on DStreams
  • Transformations on DStreams (contd.)
  • Output Operations on DStreams
  • Design Patterns for Using ForeachRDD
  • DataFrame and SQL Operations
  • DataFrame and SQL Operations (contd.)
  • Checkpointing
  • Enabling Checkpointing
  • Socket Stream
  • File Stream
  • Stateful Operations
  • Window Operations
  • Types of Window Operations
  • Types of Window Operations (contd.)
  • Join Operations – Stream-Dataset Joins
  • Join Operations – Stream-Stream Joins
  • Monitoring Spark Streaming Application
  • Performance Tuning – High Level
  • Performance Tuning – Detail Level
  • Demo – Capture and Process the Netcat Data
  • Capture and Process the Netcat Data
  • Demo – Capture and Process the Flume Data
  • Capture and Process the Flume Data
  • Demo – Capture the Twitter Data
  • Capture the Twitter Data
  • Quiz
  • Summary
  • Summary (contd.)
  • Conclusion
Module 7 - Spark ML Programming
  • Introduction
  • Objectives
  • Introduction to Machine Learning
  • Common Terminologies in Machine Learning
  • Applications of Machine Learning
  • Machine Learning in Spark
  • Spark ML API
  • DataFrames
  • Transformers and Estimators
  • Pipeline
  • Working of a Pipeline
  • Working of a Pipeline (contd.)
  • DAG Pipelines
  • Runtime Checking
  • Parameter Passing
  • General Machine Learning Pipeline – Example
  • General Machine Learning Pipeline – Example (contd.)
  • Model Selection via Cross-Validation
  • Supported Types, Algorithms, and Utilities
  • Data Types
  • Feature Extraction and Basic Statistics
  • Clustering
  • K-Means
  • K-Means (contd.)
  • Demo – Perform Clustering Using K-Means
  • Perform Clustering Using K-Means
  • Gaussian Mixture
  • Power Iteration Clustering (PIC)
  • Latent Dirichlet Allocation (LDA)
  • Latent Dirichlet Allocation (LDA) (contd.)
  • Collaborative Filtering
  • Classification
  • Classification (contd.)
  • Regression
  • Example of Regression
  • Demo – Perform Classification Using Linear Regression
  • Perform Classification Using Linear Regression
  • Demo – Run Linear Regression
  • Run Linear Regression
  • Demo – Perform Recommendation Using Collaborative Filtering
  • Perform Recommendation Using Collaborative Filtering
  • Demo – Run Recommendation System
  • Run Recommendation System
  • Quiz
  • Summary
  • Summary (contd.)
  • Conclusion
Module 9 - Spark GraphX Programming
  • Introduction
  • Objectives
  • Introduction to Graph-Parallel System
  • Limitations of Graph-Parallel System
  • Introduction to GraphX
  • Introduction to GraphX (contd.)
  • Importing GraphX
  • The Property Graph
  • The Property Graph (contd.)
  • Features of the Property Graph
  • Creating a Graph
  • Demo – Create a Graph Using GraphX
  • Create a Graph Using GraphX
  • Triplet View
  • Graph Operators
  • List of Operators
  • List of Operators (contd.)
  • Property Operators
  • Structural Operators
  • Subgraphs
  • Join Operators
  • Demo – Perform Graph Operations Using GraphX
  • Perform Graph Operations Using GraphX
  • Demo – Perform Subgraph Operations
  • Perform Subgraph Operations
  • Neighborhood Aggregation
  • mapReduceTriplets
  • Demo – Perform MapReduce Operations
  • Perform MapReduce Operations
  • Counting Degree of Vertex
  • Collecting Neighbors
  • Caching and Uncaching
  • Graph Builders
  • Vertex and Edge RDDs
  • Graph System Optimizations
  • Built-in Algorithms
  • Quiz
  • Summary
  • Summary (contd.)
  • Conclusion

Admission details

You can enrol for the  Apache Spark and Scala Certification Training course by Edureka on the edureka website by making online payment using any of the following options: 

  • Visa Credit or Debit Card
  • MasterCard
  • American Express
  • Diner’s Club
  • PayPal

Once payment is received, you will automatically receive a payment receipt and access information via email.


Filling the form

Follow the steps mentioned below to enroll for Apache Spark and Scala Certification Training

Step 1 - Click on the link - https://www.simplilearn.com/big-data-and-analytics/apache-spark-scala-certification-training

Step 2 - Click Enroll now option mentioned top on the right side. You will be redirected to a new page

Step 3 - If applicants have a coupon then they need to apply this or click on the Proceed button. 

Step 4 -  Provide necessary details i.e name, email, and contact number, and click on proceed

Step 5 - Pay the fee and save the receipt of the transaction for future reference.

Evaluation process

To become a certified Spark Developer, you need to pass the CCA175 Spark and Hadoop Developer exam. A total of 8-17 hands-on performance tasks will be assigned to the candidates. Candidates must score a minimum of 70% marks to pass the exam successfully.

In case a candidate fails the exam, they can reattempt the exam any number of times. However, once candidates clear the exam successfully, they are not allowed to retake it.

How it helps

With the Apache Spark market about to reach a market share of a whopping 4.2 Billion USD, now is an opportune time for interested candidates to train themselves in the language of Spark and Scala. With a growth rate of a respectable 67%, professionals will find the market strewn with lucrative job opportunities in sectors like manufacturing, retail, and healthcare.

Once they complete the Apache Spark and Scala Certification course, they can work as a Big Data Architect, Big Data Engineer, or Big Data Developer in varied sectors like IT, finance, healthcare, and many more. Besides, many companies like Amazon, Accenture, Visa, and more are always on the lookout to hire certified Big Data professionals.

FAQs

What is the validity of the Certification?

The Apache Spark and Scala Certification has lifelong validity.

How many marks do I need to pass the CCA175 exam?

In the CCA175 exam, candidates need to perform 8-12 hands-on tasks successfully. A candidate must score minimum 70% to pass the exam.

What does the CCA175 Hadoop certification exam cost?

The CCA175 Spark and Hadoop Developer exam costs 295 USD.

How many attempts can I take to pass the CCA175 exam?

Candidates can take the exam any number of times as they want until you pass. However, retakes are not allowed after the successful completion of a test.

How can I become a Spark Developer?

To become a certified Spark Developer, you need to pass the CCA175 Spark and Hadoop Developer exam and be proficient in the cleaning, analysis, and transformation of Spark. 

Which companies hires a certified Spark Developer?

A certified Spark Developer is usually hired by companies such as Amazon, Accenture, Visa, Hewlett Packard Enterprises, Deloitte, etc.

Are there any prerequisites for the course?

To apply for the Apache Spark and Scala certification training course, interested candidates must have a fundamental knowledge of any programming language like Python, Java, etc. They should also have a basic understanding of SQL, database, and query languages for database. In addition, they should have a working knowledge of Linux and Unix-based systems.

Will I receive practice tests for the certification exam?

Yes, you will receive one practise test for the CCA175 Spark and Hadoop certification exam as part of the course curriculum.

Trending Courses

Popular Courses

Popular Platforms

Learn more about the Courses