Hadoop Administration training course

Install, configure, and manage the Apache Hadoop platform and its associated ecosystem, and build a Hadoop solution for Big Data

JBI training course London UK

"Our tailored course provided a well rounded introduction and also covered some intermediate level topics that we needed to know. Clive gave us some best practice ideas and tips to take away. Fast paced but the instructor never lost any of the delegates"

Brian Leek, Data Analyst, May 2022

Public Courses

14/09/26 - 3 days
£1600 +VAT
26/10/26 - 3 days
£1600 +VAT
07/12/26 - 3 days
£1600 +VAT

Customised Courses

* Train a team
* Tailor content
* Flex dates
From £1200 / day
EDF logo Capita logo Sky logo NHS logo RBS logo BBC logo CISCO logo
JBI training course London UK

  • Gain an introduction to data storage and processing 
  • Architect a Hadoop solution to meet your business requirements
  • Install and build a Hadoop cluster capable of processing large data
  • Configure and tune the Hadoop environment to ensure high throughput and availability
  • Allocate, distribute and manage resources
  • Monitor the file system, job progress and overall cluster performance

Introduction to Data Storage and Processing

  • Installing the Hadoop Distributed File System (HDFS)
  •     Defining key design assumptions and architecture
  •     Configuring and setting up the file system
  •     Issuing commands from the console
  •     Reading and writing files
  • Setting the stage for MapReduce
  •     Reviewing the MapReduce approach
  •     Introducing the computing daemons
  •     Dissecting a MapReduce job

Defining Hadoop Cluster Requirements

  • Planning the architecture
  •     Selecting appropriate hardware
  •     Designing a scalable cluster
  • Building the cluster
  •     Installing Hadoop daemons
  •     Optimising the network architecture

Configuring a Cluster

  • Preparing HDFS
  •     Setting basic configuration parameters
  •     Configuring block allocation, redundancy and replication
  • Deploying MapReduce
  •     Installing and setting up the MapReduce environment
  •     Delivering redundant load balancing via Rack Awareness

Maximising HDFS Robustness

  • Creating a fault–tolerant file system
  •     Isolating single points of failure
  •     Maintaining High Availability
  •     Triggering manual failover
  •     Automating failover with Zookeeper
  • Leveraging NameNode Federation
  •     Extending HDFS resources
  •     Managing the namespace volumes
  • Introducing YARN
  •     Critiquing the YARN architecture
  •     Identifying the new daemons

Managing Resources and Cluster Health

  • Allocating resources
  •     Setting quotas to constrain HDFS utilisation
  •     Prioritising access to MapReduce using schedulers
  • Maintaining HDFS
  •     Starting and stopping Hadoop daemons
  •     Monitoring HDFS status
  •     Adding and removing data nodes
  • Administering MapReduce
  •     Managing MapReduce jobs
  •     Tracking progress with monitoring tools
  •     Commissioning and decommissioning compute nodes

Maintaining a Cluster

  • Employing the standard built–in tools
  •     Managing and debugging processes using JVM metrics
  •     Performing Hadoop status checks
  • Tuning with supplementary tools
  •     Assessing performance with Ganglia
  •     Benchmarking to ensure continued performance

Extending Hadoop

  • Simplifying information access
  •     Enabling SQL–like querying with Hive
  •     Installing Pig to create MapReduce jobs
  • Integrating additional elements of the ecosystem
  •     Imposing a tabular view on HDFS with HBase
  •     Configuring Oozie to schedule workflows

Implementing Data Ingress and Egress

  • Facilitating generic input/output
  •     Moving bulk data into and out of Hadoop
  •     Transmitting HDFS data over HTTP with WebHDFS
  • Acquiring application–specific data
  •     Collecting multi–sourced log files with Flume
  •     Importing and exporting relational information with Sqoop
  • Planning for Backup, Recovery and Security
  •     Coping with inevitable hardware failures
  •     Securing your Hadoop cluster
JBI training course London UK

IT professionals looking to learn about how to architect and administer Apache Hadoop and clusters for Big Data


5 star

4.8 out of 5 average

"Our tailored course provided a well rounded introduction and also covered some intermediate level topics that we needed to know. Clive gave us some best practice ideas and tips to take away. Fast paced but the instructor never lost any of the delegates"

Brian Leek, Data Analyst, May 2022



“JBI  did a great job of customizing their syllabus to suit our business  needs and also bringing our team up to speed on the current best practices. Our teams varied widely in terms of experience and  the Instructor handled this particularly well - very impressive”

Brian F, Team Lead, RBS, Data Analysis Course, 20 April 2022

 

 

JBI training course London UK

Certification


Every delegate will be entitled to a certificate of achievement on completion of the course.

If you are missing your certificate - please use the link below to apply - you can also use this link to sign up for the JBI Training newsletter to receive technology tips directly from our instructors - Analytics, AI, ML, DevOps, Web, Backend and Security.
 



In this Hadoop architecture and administration training course, you will gain the skills to install, configure and manage the Apache Hadoop platform and its associated ecosystem.

You will learn how to build a Hadoop solution that satisfies your business requirements.

JBI Training offers three courses in this group. Big Data Introduction is a three-day course covering the big data landscape, core concepts, and the Hadoop ecosystem for professionals new to the field. Hadoop Administration is a three-day course for IT professionals responsible for deploying and managing Hadoop clusters. SQL is a two-day course covering structured query language for working with relational databases — a foundational skill for anyone working with data at any scale. All courses are available as scheduled classroom sessions in London, as live online instructor-led training, or as customised onsite programmes for data and IT teams.
Hadoop is an open-source distributed computing framework designed to store and process very large datasets across clusters of commodity hardware. It consists of two core components — HDFS (Hadoop Distributed File System) for distributed storage, and MapReduce for distributed batch processing — along with a broader ecosystem of tools including Hive, Pig, HBase, Sqoop, and Flume. JBI's three-day Big Data Introduction course covers the big data landscape and use cases, the Hadoop architecture and ecosystem, working with HDFS, an introduction to MapReduce, querying data with Hive, and an overview of how Hadoop relates to modern cloud-based big data platforms such as Azure Databricks, AWS EMR, and Google Dataproc. It is suited to data engineers, developers, and IT professionals who are new to big data technologies.
The Hadoop Administration course is a three-day programme for IT professionals and system administrators responsible for deploying, configuring, and managing Hadoop clusters. It covers planning and installing a Hadoop cluster, configuring HDFS and YARN, managing cluster nodes and resources, monitoring cluster health and performance, configuring Hadoop security using Kerberos and ACLs, backup and disaster recovery for Hadoop data, capacity planning and tuning, and troubleshooting common cluster issues. It is suited to system administrators, infrastructure engineers, and DevOps professionals who are responsible for the operational reliability of Hadoop-based data infrastructure.
SQL (Structured Query Language) is the foundational language for querying and manipulating data in relational databases, and it remains one of the most widely used skills across data engineering, analytics, and application development. It is included in this group because SQL skills are directly relevant to big data work — Hive uses a SQL-like syntax for querying data in Hadoop, Spark SQL enables SQL-based queries on large distributed datasets, and modern cloud data warehouses such as Azure Synapse, BigQuery, and Snowflake are all queried primarily using SQL. JBI's two-day SQL course covers database design principles, SELECT queries, filtering and sorting, joins, aggregations, subqueries, and data modification — providing the foundational data querying skills needed across virtually every data platform.
Yes on both counts. While Hadoop remains in active use in many large enterprise and financial services environments, the big data landscape has evolved significantly with the rise of cloud-native platforms such as Azure Databricks, AWS EMR, Google Dataproc, and Apache Spark. JBI's Big Data training content is reviewed to reflect this context, ensuring delegates understand both the Hadoop ecosystem and how it relates to modern cloud-based alternatives. All courses can be delivered as customised onsite or online programmes for corporate data engineering and IT teams, with content tailored to the organisation's specific data infrastructure and platform environment. JBI has delivered big data and data engineering training for organisations including the BBC, NHS, RBS, Sky, EDF, and Cisco.

CONTACT
+44 (0)20 8446 7555

[email protected]

 

Copyright © 2026 JBI Training. All Rights Reserved.
JB International Training Ltd  -  Company Registration Number: 08458005
Registered Address: Wohl Enterprise Hub, 2B Redbourne Avenue, London, N3 2BS

Modern Slavery Statement & Corporate Policies | Terms & Conditions | Contact Us

POPULAR

AI training courses                                                                        CoPilot training course

Threat modelling training course   Python for data analysts training course

Power BI training course                                   Machine Learning training course

Spring Boot Microservices training course              Terraform training course

Data Storytelling training course                                               C++ training course

Power Automate training course                               Clean Code training course