Big Data Introduction training course

Leverage big data analysis tools and techniques to foster better business decision-making

JBI training course London UK

"Our tailored course provided a well rounded introduction and also covered some intermediate level topics that we needed to know. Clive gave us some best practice ideas and tips to take away. Fast paced but the instructor never lost any of the delegates"

Brian Leek, Data Analyst, May 2022

Public Courses

14/09/26 - 3 days
£1600 +VAT
26/10/26 - 3 days
£1600 +VAT
07/12/26 - 3 days
£1600 +VAT

Customised Courses

* Train a team
* Tailor content
* Flex dates
From £1200 / day
EDF logo Capita logo Sky logo NHS logo RBS logo BBC logo CISCO logo
JBI training course London UK

  • Gain an introduction to Big Data
  • Learn how to define Big Data
  • Select the correct Big Data stores for disparate data sets
  • Process large data sets using Hadoop to extract value
  • Store, manage and analyse unstructured data
  • Leverage Big Data analysis tools and techniques to foster better business decision-making
  • Query large data sets in near real-time with Pig and Hive
  • Plan and implement a Big Data strategy for your organisation

Introduction to Big Data

  • Defining Big Data
  • The four dimensions of Big Data: volume, velocity, variety, veracity
  • Introducing the Storage, MapReduce and Query Stack
  • Delivering business benefit from Big Data
  • Establishing the business importance of Big Data
  • Addressing the challenge of extracting useful data
  • Integrating Big Data with traditional data

Storing Big Data

  • Analysing your data characteristics
  • Selecting data sources for analysis
  • Eliminating redundant data
  • Establishing the role of NoSQL

Overview of Big Data stores

  •  Data models: key value, graph, document, column–family
  •  Hadoop Distributed File System
  •  HBase
  •  Hive
  •  Cassandra
  •  Hypertable
  •  Amazon S3
  •  BigTable
  •  DynamoDB
  •  MongoDB
  •  Redis
  •  Riak
  •  Neo4J

Selecting Big Data stores

  •  Choosing the correct data stores based on your data characteristics
  •  Moving code to data
  •  Implementing polyglot data store solutions
  •  Aligning business goals to the appropriate data store

Processing Big Data

  • Integrating disparate data stores
  • Mapping data to the programming framework
  • Connecting and extracting data from storage
  • Transforming data for processing
  • Subdividing data in preparation for Hadoop MapReduce
  • Employing Hadoop MapReduce
  • Creating the components of Hadoop MapReduce jobs
  • Distributing data processing across server farms
  • Executing Hadoop MapReduce jobs
  • Monitoring the progress of job flows
  • The building blocks of Hadoop MapReduce
  •  Distinguishing Hadoop daemons
  •  Investigating the Hadoop Distributed File System
  •  Selecting appropriate execution modes: local, pseudo–distributed and fully distributed
  • Handling streaming data
  • Comparing real–time processing models
  •  Leveraging Storm to extract live events
  •  Lightning–fast processing with Spark and Shark

Tools and Techniques to Analyse Big Data

  • Abstracting Hadoop MapReduce jobs with Pig
  • Communicating with Hadoop in Pig Latin
  •  Executing commands using the Grunt Shell
  •  Streamlining high–level processing
  • Performing ad hoc Big Data querying with Hive
  • Persisting data in the Hive MegaStore
  • Performing queries with HiveQL
  • Investigating Hive file formats
  • Creating business value from extracted data
  • Mining data with Mahout
  •  Visualising processed results with reporting tools
  • Querying in real time with Impala

Developing a Big Data Strategy

  • Defining a Big Data strategy for your organisation
  •     Establishing your Big Data needs
  •     Meeting business goals with timely data
  •     Evaluating commercial Big Data tools
  •     Managing organisational expectations
  • Enabling analytic innovation
  •     Focusing on business importance
  •     Framing the problem
  •     Selecting the correct tools
  •     Achieving timely results

Implementing a Big Data Solution

  •     Selecting suitable vendors and hosting options
  •     Balancing costs against business value
  •     Keeping ahead of the curve
  •  
JBI training course London UK

IT professionals looking to learn about how to implement and  enhance a corporate big data environment and looking to get a better elementary practical skills relating to Big Data


5 star

4.8 out of 5 average

"Our tailored course provided a well rounded introduction and also covered some intermediate level topics that we needed to know. Clive gave us some best practice ideas and tips to take away. Fast paced but the instructor never lost any of the delegates"

Brian Leek, Data Analyst, May 2022



“JBI  did a great job of customizing their syllabus to suit our business  needs and also bringing our team up to speed on the current best practices. Our teams varied widely in terms of experience and  the Instructor handled this particularly well - very impressive”

Brian F, Team Lead, RBS, Data Analysis Course, 20 April 2022

 

 

JBI training course London UK

Certification


Every delegate will be entitled to a certificate of achievement on completion of the course.

If you are missing your certificate - please use the link below to apply - you can also use this link to sign up for the JBI Training newsletter to receive technology tips directly from our instructors - Analytics, AI, ML, DevOps, Web, Backend and Security.
 



In this Introduction to Big Data training course, you will learn ways of storing data that allow for efficient processing and analysis. You will also gain the skills you need to store, manage, process and analyse massive amounts of unstructured data to create an appropriate data lake.

Get introduced to leading products such as Hadoop to learn how to apply Big Data in the real world.

JBI Training offers three courses in this group. Big Data Introduction is a three-day course covering the big data landscape, core concepts, and the Hadoop ecosystem for professionals new to the field. Hadoop Administration is a three-day course for IT professionals responsible for deploying and managing Hadoop clusters. SQL is a two-day course covering structured query language for working with relational databases — a foundational skill for anyone working with data at any scale. All courses are available as scheduled classroom sessions in London, as live online instructor-led training, or as customised onsite programmes for data and IT teams.
Hadoop is an open-source distributed computing framework designed to store and process very large datasets across clusters of commodity hardware. It consists of two core components — HDFS (Hadoop Distributed File System) for distributed storage, and MapReduce for distributed batch processing — along with a broader ecosystem of tools including Hive, Pig, HBase, Sqoop, and Flume. JBI's three-day Big Data Introduction course covers the big data landscape and use cases, the Hadoop architecture and ecosystem, working with HDFS, an introduction to MapReduce, querying data with Hive, and an overview of how Hadoop relates to modern cloud-based big data platforms such as Azure Databricks, AWS EMR, and Google Dataproc. It is suited to data engineers, developers, and IT professionals who are new to big data technologies.
The Hadoop Administration course is a three-day programme for IT professionals and system administrators responsible for deploying, configuring, and managing Hadoop clusters. It covers planning and installing a Hadoop cluster, configuring HDFS and YARN, managing cluster nodes and resources, monitoring cluster health and performance, configuring Hadoop security using Kerberos and ACLs, backup and disaster recovery for Hadoop data, capacity planning and tuning, and troubleshooting common cluster issues. It is suited to system administrators, infrastructure engineers, and DevOps professionals who are responsible for the operational reliability of Hadoop-based data infrastructure.
SQL (Structured Query Language) is the foundational language for querying and manipulating data in relational databases, and it remains one of the most widely used skills across data engineering, analytics, and application development. It is included in this group because SQL skills are directly relevant to big data work — Hive uses a SQL-like syntax for querying data in Hadoop, Spark SQL enables SQL-based queries on large distributed datasets, and modern cloud data warehouses such as Azure Synapse, BigQuery, and Snowflake are all queried primarily using SQL. JBI's two-day SQL course covers database design principles, SELECT queries, filtering and sorting, joins, aggregations, subqueries, and data modification — providing the foundational data querying skills needed across virtually every data platform.
Yes on both counts. While Hadoop remains in active use in many large enterprise and financial services environments, the big data landscape has evolved significantly with the rise of cloud-native platforms such as Azure Databricks, AWS EMR, Google Dataproc, and Apache Spark. JBI's Big Data training content is reviewed to reflect this context, ensuring delegates understand both the Hadoop ecosystem and how it relates to modern cloud-based alternatives. All courses can be delivered as customised onsite or online programmes for corporate data engineering and IT teams, with content tailored to the organisation's specific data infrastructure and platform environment. JBI has delivered big data and data engineering training for organisations including the BBC, NHS, RBS, Sky, EDF, and Cisco.

CONTACT
+44 (0)20 8446 7555

[email protected]

 

Copyright © 2026 JBI Training. All Rights Reserved.
JB International Training Ltd  -  Company Registration Number: 08458005
Registered Address: Wohl Enterprise Hub, 2B Redbourne Avenue, London, N3 2BS

Modern Slavery Statement & Corporate Policies | Terms & Conditions | Contact Us

POPULAR

AI training courses                                                                        CoPilot training course

Threat modelling training course   Python for data analysts training course

Power BI training course                                   Machine Learning training course

Spring Boot Microservices training course              Terraform training course

Data Storytelling training course                                               C++ training course

Power Automate training course                               Clean Code training course