IBM Storage Ceph
IBM
© Copyright IBM Corp. 2024.
US Government Users Restricted Rights - Use, duplication or disclosure restricted by GSA ADP Schedule Contract with IBM Corp.
Tables of Contents
IBM Storage Ceph
1
Summary of changes
1
Compatibility matrix
1
Release notes for 6.1
2
2
7
15
15
16
Asynchronous updates
Release notes for 6.1z7
Enhancements
Bug fixes
Release notes for 6.1z6
Release notes for 6.1z5
Enhancements
Bug fixes
Release notes for 6.1z4
Enhancements
Bug fixes
Release notes for 6.1z3
Enhancements
Bug fixes
Release notes for 6.1z2
Enhancements
Bug fixes
16
16
16
17
20
20
20
20
21
22
22
24
24
25
27
27
28
Introduction to IBM Storage Ceph
30
Overview
32
32
32
33
33
33
34
34
35
35
38
39
39
39
40
41
42
42
42
43
44
44
45
45
45
45
46
46
46
46
47
47
48
48
48
49
49
50
51
51
51
Enhancements
Bug fixes
Known issues
Technology previews
Sources
Architecture
Ceph architecture
Ceph client components
Prerequisites
Ceph client native protocol
Ceph client object watch and notify
Ceph client Mandatory Exclusive Locks
Ceph client object map
Ceph client data striping
Ceph on-wire encryption
Core Ceph components
Prerequisites
Pools
Ceph authentication
Placement groups
CRUSH ruleset
Input/output operations
Replication
Erasure coding
ObjectStore
BlueStore
Self management operations
Heartbeat
Peering
Rebalancing and recovery
Data integrity
High availability
Clustering the Ceph Monitor
Data security and hardening
Introduction to IBM Storage Ceph
Supporting Software
Threat and Vulnerability Management
Threat Actors
Security Zones
Connecting Security Zones
Security-Optimized Architecture
Encryption and Key Management
SSH
SSL Termination
Messenger v2 protocol
Encryption in transit
Compression modes of messenger v2 protocol
Encryption at Rest
Enabling key rotation
Identity and Access Management
Ceph Storage Cluster User Access
Ceph Object Gateway User Access
Ceph Object Gateway LDAP or AD authentication
Ceph Object Gateway OpenStack Keystone authentication
Infrastructure Security
Administration
Network Communication
Hardening the Network Service
Reporting
Auditing Administrator Actions
Data Retention
Ceph Storage Cluster
Ceph Block Device
Ceph File System
Ceph Object Gateway
Federal Information Processing Standard (FIPS)
Summary
Planning
Considerations and recommendations
Basic considerations
Workload considerations
Network considerations for IBM Storage Ceph
Considerations for using a RAID controller with OSD hosts
Tuning considerations for the Linux kernel when running Ceph
Colocation
Operating system requirements
Accessing Red Hat entitlements from IBM Storage Ceph
Minimum hardware considerations
Hardware
Executive summary
General principles for selecting hardware
Identify performance use case
Consider storage density
Identical hardware configuration
Network considerations for IBM Storage Ceph
Avoid using RAID and SAN solutions
Summary of common mistakes when selecting hardware
Reference
Optimize workload performance domains
Server and rack solutions
Minimum hardware recommendations for containerized Ceph
Recommended minimum hardware requirements for the IBM Storage Ceph Dashboard
Storage Strategies
Overview
What are storage strategies?
Configuring storage strategies
Crush admin overview
CRUSH introduction
Dynamic data placement
CRUSH failure domain
CRUSH performance domain
Using different device classes
CRUSH hierarchy
CRUSH location
Adding a bucket
Moving a bucket
Removing a bucket
CRUSH Bucket algorithms
Ceph OSDs in CRUSH
Viewing OSDs in CRUSH
Adding an OSD to CRUSH
Moving an OSD within a CRUSH Hierarchy
Removing an OSD from a CRUSH Hierarchy
Device class
Setting a device class
Removing a device class
Renaming a device class
Listing a device class
Listing OSDs of a device class
52
52
52
53
53
54
54
54
55
55
55
55
56
57
57
58
58
58
58
58
59
59
59
59
60
61
63
63
64
64
69
70
70
71
71
72
72
72
72
72
73
73
73
74
75
76
77
77
78
78
79
79
79
80
81
81
82
82
83
83
83
84
84
84
85
87
87
87
88
88
88
88
89
89
Listing CRUSH Rules by Class
CRUSH weights
Setting CRUSH weights of OSDs
Setting a Bucket’s OSD Weights
Set an OSD’s in Weight
Setting the OSDs weight by utilization
Setting an OSD’s Weight by PG distribution
Recalculating a CRUSH Tree’s weights
Primary affinity
CRUSH rules
Listing CRUSH rules
Dumping CRUSH rules
Adding CRUSH rules
Creating CRUSH rules for replicated pools
Creating CRUSH rules for erasure coded pools
Removing CRUSH rules
CRUSH tunables overview
CRUSH tuning
CRUSH tuning, the hard way
CRUSH legacy values
Edit a CRUSH map
Getting the CRUSH map
Decompiling the CRUSH map
Compiling the CRUSH map
Setting a CRUSH map
CRUSH storage strategies examples
Placement Groups
About placement groups
Placement group states
Placement group tradeoffs
Data durability
Data distribution
Resource usage
Placement group count
Placement group calculator
Configuring default placement group count
Placement group count for small clusters
Calculating placement group count
Maximum placement group count
Auto-scaling placement groups
Placement group auto-scaling
Placement group splitting and merging
Setting placement group auto-scaling modes
Setting minimum and maximum number of placement groups for pools
Viewing placement group scaling recommendations
Setting placement group auto-scaling
Updating noautoscale flag
Specifying target pool size
Specifying target size using the absolute size of the pool
Specifying target size using the total cluster capacity
Placement group command line interface
Setting number of placement groups in a pool
Getting number of placement groups in a pool
Getting statistics for placement groups
Getting statistics for stuck placement groups
Getting placement group maps
Getting a placement group statistics
Scrubbing placement groups
Marking unfound objects
Pools overview
Pools and storage strategies overview
Listing pool
Creating a pool
Setting pool quota
Deleting a pool
Renaming a pool
Migrating a pool
Viewing pool statistics
Setting pool values
Getting pool values
Enabling a client application
Disabling a client application
89
89
89
90
90
90
91
91
91
91
93
94
94
94
94
95
95
95
96
96
96
96
96
97
97
97
98
98
99
101
101
101
102
102
102
102
102
103
103
103
104
104
105
106
106
107
108
108
108
108
109
109
109
109
110
110
110
110
110
111
112
112
112
114
114
114
114
115
115
115
116
116
Setting application metadata
Removing application metadata
Setting the number of object replicas
Getting the number of object replicas
Pool values
Erasure code pools overview
Creating a sample erasure-coded pool
Erasure code profiles
Setting OSD erasure-code-profile
Removing OSD erasure-code-profile
Getting OSD erasure-code-profile
Listing OSD erasure-code-profile
Erasure Coding with Overwrites
Erasure Code Plugins
Creating a new erasure code profile using jerasure erasure code plugin
Controlling CRUSH Placement
Installing
Installing Pro Edition for free
Initial installation
cephadm utility
How cephadm works
cephadm-ansible playbooks
Registering the IBM Storage Ceph nodes
Configuring Ansible inventory location
Creating an Ansible user with sudo access
Configuring SSH
Configuring a different SSH user
Enabling password-less SSH for Ansible
Enabling SSH login as root user on Red Hat Enterprise Linux 9
Running the preflight playbook
Bootstrapping a new storage cluster
Recommended cephadm bootstrap command options
Using a JSON file to protect login information
Bootstrapping a storage cluster using a service configuration file
Bootstrapping the storage cluster as a non-root user
Bootstrap command options
Obtaining entitlement key
Distributing SSH keys
Disconnected installation
Configuring a private registry for a disconnected installation
Running the preflight playbook for a disconnected installation
Performing a disconnected installation
Changing configurations of custom container images for disconnected installations
Adding hosts in disconnected deployments
Launching the cephadm shell
cephadm commands
Verifying the cluster installation
Adding hosts
Using the addr option to identify hosts
Adding multiple hosts
Removing hosts
Labeling hosts
Adding a label to a host
Removing a label from a host
Using host labels to deploy daemons on specific hosts
Adding Monitor service
Deploying Ceph monitor nodes using host labels
Adding Ceph Monitor nodes by IP address or network name
Setting up the admin node
Removing the admin label from a host
Adding Manager service
Adding OSDs
Purging the Ceph storage cluster
Deploying client nodes
Managing an IBM Storage Ceph cluster using cephadm-ansible modules
cephadm-ansible modules
cephadm-ansible modules options
Bootstrapping a storage cluster using the cephadm_ansible modules
Adding or removing hosts using the ceph_orch_host module
Setting configuration options using the ceph_config module
Applying a service specification using the ceph_orch_apply module
Managing Ceph daemon states using the ceph_orch_daemon module
Comparison between Ceph Ansible and Cephadm
What to do next? Day 2
116
117
117
117
117
120
121
121
122
123
124
124
124
124
124
126
126
126
127
127
128
129
129
130
131
132
132
133
134
134
135
137
137
138
139
140
141
141
141
141
144
145
146
147
147
148
150
151
152
153
154
155
155
156
156
157
158
159
159
160
161
161
162
163
164
165
165
166
167
170
171
172
173
174
Upgrading
174
174
175
175
177
179
180
181
181
184
184
184
186
Configuring
186
187
187
188
189
189
190
190
191
191
192
192
192
193
194
195
195
197
197
198
198
199
199
200
200
200
200
201
201
201
201
202
202
203
203
204
204
205
205
205
205
206
206
206
206
207
207
207
207
207
207
208
209
209
210
210
211
214
214
221
222
224
230
232
234
Upgrading an IBM Storage Ceph cluster using cephadm
Compatibility considerations between Ceph and podman versions
Upgrading the IBM Storage Ceph cluster
Crossgrading from Red Hat Ceph Storage 6.1 to IBM Storage Ceph 6.1
Upgrading cluster in a disconnected environment
Upgrading a host operating system from RHEL 8 to RHEL 9
Upgrading IBM Storage Ceph 5 to IBM Storage Ceph 6 involving RHEL 8 to RHEL 9 upgrades
Upgrading IBM Storage Ceph 5 to 6 involving RHEL 8 to RHEL 9 upgrades with stretch mode enabled
Staggered upgrade
Staggered upgrade options
Performing a staggered upgrade
Monitoring and managing upgrade
Ceph configuration
Configuration database
Using the Ceph metavariables
Viewing the Ceph configuration at runtime
Viewing a specific configuration at runtime
Setting a specific configuration at runtime
OSD Memory Target
Setting the OSD memory target
Automatically tuning OSD memory
MDS Memory Cache Limit
Ceph network configuration
Network configuration for Ceph
Network messenger
Configuring a public network
Configuring a private network
Configuring multiple public networks to the cluster
Verifying firewall rules are configured for default Ceph ports
Firewall settings for Ceph Monitor node
Firewall settings for Ceph OSDs
Ceph Monitor configuration
Ceph Monitor configuration
Viewing the Ceph Monitor configuration database
Cluster maps
Ceph Monitor quorum
Ceph Monitor consistency
Bootstrap the Ceph Monitor
Minimum configuration for a Ceph Monitor
Unique identifier for Ceph
Ceph Monitor data store
Storage capacity
Ceph heartbeat
Ceph Monitor synchronization role
Time synchronization
Ceph authentication configuration
Cephx authentication
Enabling Cephx
Disabling Cephx
Cephx user keyrings
Cephx daemon keyrings
Cephx message signatures
Pools, placement groups, and CRUSH configuration
Pools placement groups and CRUSH
Ceph Object Storage Daemon (OSD) configuration
Ceph OSD configuration
Scrubbing the OSD
Backfilling an OSD
OSD recovery
Ceph Monitor and OSD interaction configuration
Ceph Monitor and OSD interaction
OSD heartbeat
Reporting an OSD as down
Reporting a peering failure
OSD reporting status
Debugging and logging configuration
General configuration options
Network configuration options
Ceph firewall ports
Ceph Monitor configuration options
Cephx configuration options
Pools, placement groups, and CRUSH configuration options
Object Storage Daemon (OSD) configuration options
Ceph Monitor and OSD configuration options
Debugging and logging configuration options
Scrubbing options
BlueStore configuration options
237
Administering
237
237
237
238
238
238
239
240
240
243
244
245
245
246
246
246
247
248
249
250
252
255
255
256
256
258
258
258
258
259
259
259
259
260
260
260
260
261
261
261
262
262
264
264
266
267
267
268
268
268
270
270
271
271
272
272
273
273
273
274
274
275
276
277
277
278
279
281
282
283
283
Administration
Understanding process management
Process management
Starting, stopping, and restarting all Ceph daemons using the systemctl command
Starting, stopping, and restarting all Ceph services
Viewing log files of Ceph daemons
Powering down and rebooting the cluster
Powering down and rebooting that uses Ceph Orchestrator
Powering down and rebooting that uses systemctl commands
Monitoring a Ceph cluster
High-level monitoring
Checking storage cluster health
Watching storage cluster events
How Ceph calculates data usage
Understanding storage clusters usage stats
Checking storage cluster status
Understanding OSD usage stats
Understanding Ceph OSD status
Checking Ceph Monitor status
Using Ceph administration socket
Low-level monitoring
Monitoring placement group sets
Ceph OSD peering
Placement Group States
Placement Group creating state
Placement group peering state
Placement group active state
Placement Group clean state
Placement Group degraded state
Placement Group recovering state
Back fill state
Placement Group remapped state
Placement Group stale state
Placement Group misplaced state
Placement Group incomplete state
Identifying stuck Placement Groups
Finding object’s location
Stretch clusters for Ceph storage
Stretch mode
Setting crush location for daemons
Setting the crush location during bootstrap
Setting the crush location for daemons manually
Entering stretch mode
Adding OSD hosts in stretch mode
Override Ceph behavior
Setting and unsetting override options
Override use cases
Ceph user management
Ceph user management background
Managing Ceph users
Listing Ceph users
Displaying Ceph user information
Adding Ceph user
Modifying Ceph user
Deleting Ceph user
Printing Ceph user key
Using ceph-volume utility
Ceph volume lvm plugin
Why does ceph-volume replace ceph-disk?
Preparing OSDs
Listing devices
Activating OSDs
Deactivating OSDs
Creating Ceph OSDs
Migrating BlueFS data
Expanding BlueFS DB device
Using batch mode
Zapping data
Ceph performance benchmark
Performance baseline
Benchmarking Ceph performance
Benchmarking Ceph block performance
Benchmarking CephFS performance
Benchmarking Ceph Object Gateway performance
Ceph performance counters
Access to Ceph performance counters
Display Ceph performance counters
Dump Ceph performance counters
Average count and sum
Ceph Monitor metrics
Ceph OSD metrics
Ceph Object Gateway metrics
mClock OSD scheduler
Comparison of mClock OSD scheduler with WPQ OSD scheduler
Allocation of input and output resources
Factors impacting mClock operation queues
mClock configuration
mClock clients
mClock profiles
mClock profile types
Changing mClock profiles
Switching between built-in and custom profiles
Switching temporarily between mClock profiles
Degraded and misplaced object recovery rate with mClock profiles
Modifying backfills and recovery options
Ceph OSD capacity determination
Verifying the capacity of an OSD
Manually benchmarking OSDs
Determining BlueStore throttle values
Specifying maximum OSD capacity
mClock configuration options
BlueStore
BlueStore features
BlueStore devices
BlueStore caching
Sizing considerations
Tuning BlueStore
Resharding RocksDB database
BlueStore fragmentation tool
Checking for fragmentation
Ceph BlueStore BlueFS
Viewing the bluefs_buffered_io setting
Viewing Ceph BlueFS statistics for Ceph OSDs
Cephadm troubleshooting
Pause or disable cephadm
Per service and per daemon event
Check cephadm logs
Gather log files
Collect systemd status
List all downloaded container images
Manually run containers
CIDR network error
Access admin socket
Manually deploying a mgr daemon
Cephadm operations
Monitor cephadm log messages
Ceph daemon logs
Data location
Cephadm health checks
Cephadm operations health checks
Cephadm configuration health checks
Using cephadm-ansible modules
cephadm-ansible module options
Bootstrapping storage cluster
Adding or removing hosts
Setting configuration options
Applying a service specification
Managing Ceph daemon states
Operations
Introduction to the Ceph Orchestrator
Use of the Ceph Orchestrator
Managing services
Checking service status
283
285
285
286
286
287
287
288
288
288
290
294
295
296
296
297
298
298
298
298
301
302
303
303
304
305
305
306
306
307
308
311
311
311
312
312
312
313
315
316
317
317
318
319
319
319
320
320
321
321
321
322
322
322
323
323
324
324
325
325
325
326
327
328
329
332
333
334
335
336
336
337
337
Checking daemon status
Placement specification of the Ceph Orchestrator
Deploying the Ceph daemons using the command line interface
Deploying the Ceph daemons on a subset of hosts using the command line interface
Service specification of the Ceph Orchestrator
Deploying the Ceph daemons using the service specification
Managing hosts
Adding hosts
Adding multiple hosts
Listing hosts
Adding a label to a host
Removing a label from a host
Removing hosts
Placing hosts in the maintenance mode
Managing monitors
Ceph Monitors
Configuring monitor election strategy
Deploying the Ceph monitor daemons using the command line interface
Deploying the Ceph monitor daemons using the service specification
Deploying the monitor daemons on specific network
Removing the monitor daemons
Removing a Ceph Monitor from an unhealthy storage cluster
Managing manager daemons
Deploying the manager daemons
Removing the manager daemons
Using Ceph Manager modules
Using the Ceph Manager balancer module
Using the Ceph Manager alerts module
Using the Ceph manager crash module
Telemetry module
Managing OSDs
Ceph OSDs
Ceph OSD node configuration
Automatically tuning OSD memory
Listing devices for Ceph OSD deployment
Zapping devices for Ceph OSD deployment
Deploying Ceph OSDs on all available devices
Deploying Ceph OSDs on specific devices and hosts
Advanced service specifications and filters for deploying OSDs
Deploying Ceph OSDs using advanced service specifications
Removing the OSD daemons
Replacing the OSDs
Replacing the OSDs with pre-created LVM
Replacing the OSDs in a non-colocated scenario
Stopping the removal of the OSDs
Activating the OSDs
Observing the data migration
Recalculating the placement groups
Managing the monitoring stack
Deploying the monitoring stack
Removing the monitoring stack
Basic client setup
Configuring file setup on client machines
Setting-up keyring on client machines
Managing the MDS service
Deploying the MDS service using the command line interface
Deploying the MDS service using the service specification
Removing the MDS service
Managing the Ceph Object Gateway
Deploying the Ceph Object Gateway using the command line interface
Deploying the Ceph Object Gateway using the service specification
Deploying a multi-site Ceph Object Gateway
Removing the Ceph Object Gateway
Managing the NFS-Ganesha gateway (Technology Preview)
Creating the NFS-Ganesha cluster
Deploying the NFS-Ganesha gateway using the command line interface
Deploying the NFS-Ganesha gateway using the service specification
Implementing HA for CephFS/NFS service (Technology Preview)
Upgrading a standalone CephFS/NFS cluster for HA
Deploying HA for CephFS/NFS using a specification file
Updating the NFS-Ganesha cluster
Viewing the NFS-Ganesha cluster information
Fetching the NFS-Ganesha cluster logs
338
338
339
340
341
341
342
343
344
345
155
156
346
347
348
348
349
349
350
351
352
352
353
354
355
355
356
358
360
362
364
364
365
365
366
366
367
368
369
370
373
374
375
376
379
380
381
381
381
382
383
384
384
385
385
385
387
388
389
389
391
392
395
396
396
397
398
399
400
402
405
406
406
Setting custom NFS-Ganesha configuration
Resetting custom NFS-Ganesha configuration
Deleting the NFS-Ganesha cluster
Removing the NFS-Ganesha gateway
Configuring SNMP traps
Simple network management protocol
Configuring snmptrapd
Deploying the SNMP gateway
Handling a node failure
Considerations before adding or removing a node
Workflow for replacing a node
Replacing the node by using the root and Ceph OSD disks from the failed node
Replacing the node by reinstalling the operating system and using the Ceph OSD disks from the failed node
Replacing the node by reinstalling the operating system and using all new Ceph OSD disks
Performance considerations
Recommendations for adding or removing nodes
Adding a Ceph OSD node
Removing a Ceph OSD node
Simulating a node failure
Handling a data center failure
Avoiding a data center failure
Handling a data center failure
Dashboard
Ceph dashboard overview
Components
Features
Architecture
Installation and access
Network port requirements for Ceph Dashboard
Accessing the Ceph dashboard
Setting message of the day (MOTD)
Expanding the cluster
Toggling Ceph dashboard features
Understanding the landing page of the Ceph dashboard
Changing the dashboard password
Changing the Ceph dashboard password using the command line interface
Setting admin user password for Grafana
Enabling IBM Storage Ceph Dashboard manually
Creating an admin account for syncing users to the Ceph dashboard
Syncing users from Red Hat Sign-On to the Ceph dashboard
Enabling Single Sign-On for the Ceph Dashboard
Disabling Single Sign-On for the Ceph Dashboard
Managing roles
User roles and permissions
Creating roles
Editing roles
Cloning roles
Deleting roles
Managing users
Creating users
Editing users
Deleting users
User capabilities
Access capabilities
Creating user capabilities
Editing user capabilities
Importing user capabilities
Exporting user capabilities
Deleting user capabilities
Managing Ceph daemons
Daemon actions
Monitoring the cluster
Monitoring hosts of the Ceph cluster
Viewing and editing the configuration of the Ceph cluster
Viewing and editing the manager modules of the Ceph cluster
Monitoring monitors of the Ceph cluster
Monitoring services of the Ceph cluster
Monitoring Ceph OSDs
Monitoring HAProxy
Viewing the CRUSH map of the Ceph cluster
Filtering logs of the Ceph cluster
Viewing centralized logs of the Ceph cluster
Monitoring pools of the Ceph cluster
Monitoring Ceph File Systems
Monitoring Ceph Object Gateway daemons
Monitoring Block device images
Managing alerts
407
409
410
410
411
411
412
414
416
416
417
417
417
418
419
419
420
421
421
422
422
423
424
424
425
425
426
427
427
428
429
430
431
432
434
435
435
436
437
438
440
441
442
442
443
444
445
445
446
446
447
448
448
449
450
451
452
453
453
454
454
455
455
456
456
457
457
458
458
459
460
460
462
462
462
463
463
Enabling monitoring stack
Configuring a Grafana certificate
Adding Alertmanager webhooks
Viewing alerts
Creating a silence
Re-creating a silence
Editing a silence
Expiring a silence
Managing NFS Ganesha exports
Configuring NFS Ganesha daemons
Configuring NFS exports with CephFS
Editing NFS Ganesha daemons
Deleting NFS Ganesha daemons
Managing pools
Creating pools
Editing pools
Deleting pools
Managing hosts
Entering maintenance mode
Exiting maintenance mode
Removing hosts
Managing Ceph OSDs
Managing the OSDs
Replacing the failed OSDs
Managing Ceph Object Gateway
Manually adding Ceph Object Gateway login credentials to the dashboard
Creating the Ceph Object Gateway services with SSL using the dashboard
Managing Ceph Object Gateway users
Creating Ceph Object Gateway users
Creating Ceph Object Gateway subusers
Editing Ceph Object Gateway users on the dashboard
Deleting Ceph Object Gateway users
Managing Ceph Object Gateway buckets
Creating Ceph Object Gateway buckets
Editing Ceph Object Gateway buckets
Deleting Ceph object gateway buckets
Monitoring multisite Object Gateway configuration
Managing buckets of a multi-site object configuration
Editing buckets of a multi-site Object Gateway configuration
Deleting buckets of a multisite Object Gateway configuration
Managing block devices
Managing block device images
Creating images
Creating namespaces
Editing images
Copying images
Moving images to trash
Purging trash
Restoring images from trash
Deleting images
Deleting namespaces.
Creating snapshots of images
Renaming snapshots of images
Protecting snapshots of images
Cloning snapshots of images
Copying snapshots of images
Unprotecting snapshots of images
Rolling back snapshots of images
Deleting snapshots of images
Managing mirroring functions
Mirroring view
Editing mode of pools
Adding peer in mirroring
Editing peer in mirroring
Deleting peer in mirroring
Activating and deactivating telemetry
Ceph Object Gateway
Considerations and recommendations
Network considerations for IBM Storage Ceph
Basic IBM Storage Ceph considerations
Colocating Ceph daemons and its advantages
IBM Storage Ceph workload considerations
Ceph Object Gateway considerations
Administrative data storage
Index pool
464
466
467
468
468
468
469
469
470
470
471
473
473
474
474
474
475
476
477
477
478
479
479
481
483
483
484
485
485
487
488
489
489
489
490
491
491
492
492
493
494
494
494
495
496
496
497
497
498
498
499
499
499
500
500
501
502
502
503
503
503
504
504
507
508
508
509
511
511
512
512
514
516
516
517
Data pool
Data extra pool
Developing CRUSH hierarchies
Creating CRUSH roots
Using logical host names in a CRUSH map
Creating CRUSH rules
Ceph Object Gateway multi-site considerations
Considering storage sizing
Considering storage density
Considering disks for the Ceph Monitor nodes
Adjusting backfill and recovery settings
Adjusting the cluster map size
Adjusting scrubbing
Increase rgw_thread_pool_size
Increase objecter_inflight_ops
Tuning considerations for the Linux kernel when running Ceph
Deployment
Deploying the Ceph Object Gateway using the command line interface
Deploying the Ceph Object Gateway using the service specification
Deploying a multi-site Ceph Object Gateway using the Ceph Orchestrator
Removing the Ceph Object Gateway using the Ceph Orchestrator
Using the Ceph Manager rgw module
Deploying Ceph Object Gateway using the rgw module
Deploying Ceph Object Gateway multi-site using the rgw module
Basic configuration
Add a wildcard to the DNS
Beast front-end web server
Beast configuration options
Configuring SSL for Beast
D3N Data Cache
Adding D3N cache directory
Configuring D3N
Adjusting logging and debugging output
Static web hosting
Static web hosting assumptions
Static web hosting requirements
Static web hosting gateway setup
Static web hosting DNS configuration
Creating a static web hosting site
High availability for the Ceph Object Gateway
High availability service
Configuring high availability for the Ceph Object Gateway
Exporting the namespace to NFS-Ganesha (Technology Preview)
Multi-site configuration and administration
Requirements and Assumptions
Pools
Migrating a single site system to multi-site
Establishing a secondary zone
Configuring the archive zone (Technology Preview)
Deleting objects in archive zone
Failover and disaster recovery
Configuring multiple realms in the same storage cluster
Using multi-site sync policies
Multi-site sync policy group state
Retrieving the current policy
Creating a sync policy group
Modifying a sync policy group
Get a sync policy group
Removing a sync policy group
Creating a sync flow
Removing sync flows and zones
Creating or updating a sync group pipe
Modifying or deleting a sync group pipe
Obtaining information about sync operations
Bucket granular sync policies
Disabling policy between buckets
Setting bi-directional policy for buckets
Setting bi-directional policy for zonegroups
Advanced configuration
Multi-site Ceph Object Gateway command line usage
Realms
Creating a realm
Making a Realm the Default
Deleting a Realm
Getting a realm
Listing realms
518
518
518
518
519
520
520
522
522
522
523
523
523
523
523
523
524
389
525
527
530
530
531
532
533
534
535
536
536
537
537
538
539
540
540
540
541
541
542
542
542
543
545
545
546
548
548
549
550
551
552
554
559
559
560
560
561
561
561
562
562
563
564
564
564
565
566
568
569
569
570
570
570
570
570
571
Setting a realm
Listing Realm Periods
Pulling a Realm
Renaming a Realm
Zone Groups
Creating a Zone Group
Making a Zone Group the Default
Renaming a Zone Group
Deleting a Zone Group
Listing Zone Groups
Getting a Zone Group
Setting a Zone Group Map
Setting a Zone Group
Zones
Creating a Zone
Deleting a zone
Modifying a Zone
Listing Zones
Getting a Zone
Setting a Zone
Renaming a zone
Adding a Zone to a Zone Group
Removing a Zone from a Zone Group
Configuring LDAP and Ceph Object Gateway
Install Red Hat Directory Server
Configure the Directory Server firewall
Label ports for SELinux
Configure LDAPS
Check if the gateway user exists
Add a gateway user
Configure the gateway to use LDAP
Using a custom search filter
Add an S3 user to the LDAP server
Export an LDAP token
Test the configuration with an S3 client
Configuring Active Directory and Ceph Object Gateway
Using Microsoft Active Directory
Configuring Active Directory for LDAPS
Check if the gateway user exists
Add a gateway user
Configuring the gateway to use Active Directory
Add an S3 user to the LDAP server
Export an LDAP token
Test the configuration with an S3 client
Ceph Object Gateway and OpenStack Keystone
Roles for Keystone authentication
Keystone authentication and the Ceph Object Gateway
Creating the Swift service
Setting the Ceph Object Gateway endpoints
Verifying Openstack is using the Ceph Object Gateway endpoints
Configuring the Ceph Object Gateway to use Keystone SSL
Configuring the Ceph Object Gateway to use Keystone authentication
Restarting the Ceph Object Gateway daemon
Security
Server side encryption
Setting the default encryption for an existing S3 bucket
Deleting the default bucket encryption
Server-side encryption requests
Configuring server-side encryption
Using HashiCorp Vault
Secret engines for Vault
Authentication for Vault
Namespaces for Vault
Transit engine compatibility support
Creating token policies for Vault
Configuring the Ceph Object Gateway to use SSE-S3 with Vault
Configuring the Ceph Object Gateway to use SSE-KMS with Vault
Creating a key using the kv engine
Creating a key using the transit engine
Uploading an object using AWS and the Vault
Ceph Object Gateway and multi-factor authentication
Multi-factor authentication
Creating a seed for multi-factor authentication
571
571
571
571
571
572
572
572
572
572
573
573
574
575
575
575
576
576
576
577
577
577
577
577
578
578
578
578
579
579
579
580
580
580
581
581
582
582
582
582
583
583
583
584
584
585
585
585
586
586
587
587
588
589
589
590
591
591
591
592
593
594
594
595
595
596
598
600
600
601
601
601
602
Creating a new multi-factor authentication TOTP token
Test a multi-factor authentication TOTP token
Resynchronizing a multi-factor authentication TOTP token
Listing multi-factor authentication TOTP tokens
Display a multi-factor authentication TOTP token
Deleting a multi-factor authentication TOTP token
Administrating Ceph Object Gateway
Creating storage policies
Creating indexless buckets
Configure bucket index resharding
Bucket index resharding
Recovering bucket index
Limitations of bucket index resharding
Configuring bucket index resharding in simple deployments
Configuring bucket index resharding in multi-site deployments
Resharding bucket index dynamically
Resharding bucket index dynamically in multi-site configuration
Resharding bucket index manually
Cleaning stale instances of bucket entries after resharding
Fixing lifecycle policies after resharding
Enabling compression
User management
Multi-tenant namespace
Create a user
Create a subuser
Get user information
Modify user information
Enable and suspend users
Remove a user
Remove a subuser
Rename a user
Create a key
Add and remove access keys
Add and remove admin capabilities
Role management
Creating a role
Getting a role
Listing a role
Updating assume role policy document of a role
Getting permission policy attached to a role
Listing permission policy attached to a role
Deleting policy attached to a role
Deleting a role
Updating the session duration of a role
Quota management
Set user quotas
Enable and disable user quotas
Set bucket quotas
Enable and disable bucket quotas
Get quota settings
Update quota stats
Get user quota usage stats
Quota cache
Reading and writing global quotas
Bucket management
Renaming buckets
Removing buckets
Moving buckets
Moving buckets between non-tenanted users
Moving buckets between tenanted users
Moving buckets from non-tenanted users to tenanted users
Finding orphan and leaky objects
Managing bucket index entries
Bucket notifications
Creating bucket notifications
Bucket lifecycle
Creating a lifecycle management policy
Deleting a lifecycle management policy
Updating a lifecycle management policy
Monitoring bucket lifecycles
Configuring lifecycle expiration window
S3 bucket lifecycle transition within a storage cluster
Transitioning an object from one storage class to another
602
603
603
604
604
605
605
606
607
608
608
608
609
609
610
611
613
614
615
615
616
617
617
618
618
619
619
619
619
619
620
621
622
622
623
623
624
624
625
625
626
626
627
627
628
628
628
628
629
629
629
629
629
629
630
630
631
632
632
632
633
634
635
636
636
638
639
640
641
643
644
645
645
Enabling object lock for S3
Usage
Show usage
Trim usage
Ceph Object Gateway data layout
Object lookup path
Multiple data pools
Bucket and object listing
Object Gateway data layout parameters
Optimize the Ceph Object Gateway's garbage collection
Viewing the garbage collection queue
Adjusting garbage collection settings
Adjusting garbage collection for delete-heavy workloads
Optimize the Ceph Object Gateway's data object storage
Parallel thread processing for bucket life cycles
Optimizing the bucket lifecycle
Transitioning data to Amazon S3 cloud service
Transitioning data to Azure cloud service (Technology Preview)
Testing
Create an S3 user
Create a Swift user
Test S3 access
Test Swift access
Configuration reference
General settings
About pools
Lifecycle settings
Swift settings
Logging settings
Keystone settings
Keystone integration configuration options
LDAP settings
File systems
Ceph File System features and enhancements
File system components
Ceph File System and SELinux
Ceph File System limitations and the POSIX standards
Ceph File System Metadata Server
Metadata Server daemon states
Metadata Server ranks
Metadata Server cache size limits
File system affinity
Managing the MDS service using the Ceph Orchestrator
Deploying the MDS service with the command-line interface
Deploying the MDS service using the service specification
Removing the MDS service using the Ceph Orchestrator
Configuring file system affinity
Configuring multiple active Metadata Server daemons
Configuring the number of standby daemons
Configuring the standby-replay Metadata Server
Ephemeral pinning policies
Manually pinning directory trees to a particular rank
Decreasing the number of active Metadata Server daemons
Deploying the Ceph File System
Layout, quota, snapshot, and network restrictions
Creating Ceph File Systems
Adding an erasure-coded pool to a Ceph File System
Creating client users for a Ceph File System
Mounting the Ceph File System as a kernel client
Mounting the Ceph File System as a FUSE client
Managing Ceph File System volumes, subvolume groups, and subvolumes
Ceph File System volumes
Creating a Ceph File System volume
Listing Ceph File System volumes
Viewing information about a Ceph File System volume
Removing a Ceph File System volume
Ceph File System subvolume groups
Creating a file system subvolume group
Setting and managing quotas on a file system subvolume group
Listing file system subvolume groups
Fetching the absolute path of a file system subvolume group
Listing snapshots of a file system subvolume group
Removing snapshot of a file system subvolume group
Removing a file system subvolume group
Ceph File System subvolumes
Creating a file system subvolume
649
651
651
651
651
652
652
653
653
653
654
654
654
655
655
655
656
661
666
666
667
669
670
670
670
672
672
673
673
673
673
676
677
678
678
679
680
680
681
681
681
682
682
682
683
684
685
685
686
686
687
688
688
689
690
691
692
693
694
696
698
698
699
699
699
700
700
700
701
701
702
702
702
703
703
703
Listing a file system subvolume
Resizing a file system subvolume
Fetching absolute path of a file system subvolume
Fetching metadata of a file system subvolume
Creating snapshot of a file system subvolume
Cloning subvolumes from snapshots
Listing snapshots of a file system subvolume
Fetching metadata of the snapshots of a file system subvolume
Removing a file system subvolume
Removing snapshot of a file system subvolume
Metadata information on Ceph File System subvolumes
Setting custom metadata on the file system subvolume
Getting custom metadata on the file system subvolume
Listing custom metadata on the file system subvolume
Removing custom metadata from the file system subvolume
Administering Ceph File Systems
Using the cephfs-top utility
cephfs-top utility interactive commands
cephfs-top utility options
Using the MDS autoscaler module
Unmounting Ceph File Systems mounted as kernel clients
Unmounting Ceph File Systems mounted as FUSE clients
Mapping directory trees to Metadata Server daemon ranks
Disassociating directory trees from Metadata Server daemon ranks
Adding data pools
Taking down a Ceph File System cluster
Removing a Ceph File System
Using the ceph mds fail command
Client features
Ceph File System client evictions
Blocklist Ceph File System clients
Manually evicting a Ceph File System client
Removing a Ceph File System client from the blocklist
NFS cluster and export management
Creating an NFS cluster
Customizing an NFS configuration
Exporting Ceph File System namespaces over the NFS protocol (Technology Preview)
Modifying the Ceph File System exports
Creating custom Ceph File System exports
Deleting Ceph File System exports
Deleting an NFS cluster
Ceph File System quotas
Viewing quotas
Setting quotas
Removing quotas
File and directory layouts
Setting file and directory layout fields
Viewing file and directory layout fields
Viewing individual layout fields
Removing directory layouts
Ceph File System snapshots
Creating a snapshot for a Ceph File System
Ceph File System snapshot schedules
Adding a snapshot schedule for a Ceph File System
Adding a snapshot schedule for Ceph File System subvolume
Activating snapshot schedule for a Ceph File System
Activating snapshot schedule for a Ceph File System subvolume
Deactivating snapshot schedule for a Ceph File System
Deactivating snapshot schedule for a Ceph File System subvolume
Removing a snapshot schedule for a Ceph File System
Removing a snapshot schedule for a Ceph File System subvolume
Removing snapshot schedule retention policy for a Ceph File System
Removing snapshot schedule retention policy for a Ceph File System subvolume
Ceph File System mirrors
Configuring a snapshot mirror for a Ceph File System
Viewing the mirror status for a Ceph File System
Configuring Ceph File Systems
Health messages
Configuring Metadata Server daemons
Configuring journaler
Configuring Ceph File System mirrors
Ceph Block Devices
Using Ceph Block Devices
Displaying command help
Creating block device pool
Creating images
Listing images
704
704
704
705
706
706
707
708
708
709
709
709
710
710
710
710
711
712
713
713
714
714
714
715
715
716
717
718
718
719
719
719
720
721
721
722
723
725
727
728
728
729
729
730
730
730
731
731
732
732
732
733
733
734
735
737
738
738
738
739
740
740
741
741
742
744
746
746
747
756
757
759
759
760
760
761
761
Retrieving image information
Resizing images
Removing images
Moving images to trash
Schedule automatic trash purge
Enabling and disabling image features
Working with image metadata
Moving images between pools
Migrating pools
rbdmap service
Configuring rbdmap service
Persistent write log cache
Limitations
Enabling
Checking status
Flushing
Discarding
Monitoring performance of Ceph Block Devices
Ceph user and keyring
Live migration of images
Process
Formats
Streams
Preparation
Preparing import-only migration
Execution
Commit
Abort
Image encryption
Encryption format
Encryption load
Supported formats
Adding encryption format to images and clones
Snapshot management
Creating snapshots
Listing snapshots
Roll back snapshots
Deleting snapshot
Purging snapshots
Renaming snapshots
Ceph Block Device layering
Protecting snapshots
Cloning snapshots
Unprotecting snapshots
Listing children of a snapshot
Flattening cloned images
Mirroring Ceph Block Devices
Ceph Block Device mirroring
Overview of journal-based and snapshot-based mirroring
Configuring one-way mirroring
Configuring two-way mirroring
Administration for mirroring Ceph Block Devices
Viewing information on peers
Enabling mirroring on a pool
Disabling mirroring on a pool
Enabling image mirroring
Disabling image mirroring
Image promotion and demotion
Image resynchronization
Adding a storage cluster peer
Removing a storage cluster peer
Getting mirroring status for a pool
Getting mirroring status for a single image
Delaying block device replication
Converting journal-based mirroring to snapshot-based mirroring
Creating an image mirror-snapshot
Scheduling mirror-snapshots
Creating a schedule
Listing all snapshot schedules at a specific level
Removing a schedule
Viewing the status for the next snapshots to be created
Disaster recovery
Recover with one-way mirroring
Recover with two-way mirroring
Failover after an orderly shutdown
Failover after a non-orderly shutdown
761
762
762
763
763
764
765
766
767
767
768
768
769
769
770
770
771
771
772
772
772
773
773
774
775
776
776
777
777
777
778
778
778
780
780
781
781
781
782
782
782
783
784
784
784
785
785
785
787
787
789
791
792
792
793
793
793
794
794
795
795
795
796
796
797
797
798
798
798
799
799
800
800
800
800
801
Prepare for fail back
Fail back to the primary storage cluster
Remove two-way mirroring
Managing ceph-immutable-object-cache daemons
Overview
Configuration
Generic settings
QOS settings
rbd kernel module
Creating a Ceph Block Device and using it from a Linux kernel module client
Creating a Ceph Block Device for a Linux kernel module client using dashboard
Map and mount a Ceph Block Device on Linux using the command line
Mapping a block device
Displaying mapped block devices
Unmapping a block device
Ceph Block Device python module
Ceph Block Device configuration reference
Default options
General options
Caching options
Parent and child read options
Read ahead options
Blocklist options
Journal options
Configuration override options
Input and output options
Developer
Ceph RESTful API
Prerequisites
Versioning for the Ceph API
Authentication and authorization for the Ceph API
Enabling and Securing the Ceph API module
Questions and Answers
Getting information
How Can I View All Cluster Configuration Options?
How Can I View a Particular Cluster Configuration Option?
How Can I View All Configuration Options for OSDs?
How Can I View CRUSH Rules?
How Can I View Information about Monitors?
How Can I View Information About a Particular Monitor?
How Can I View Information about OSDs?
How Can I View Information about a Particular OSD?
How Can I Determine What Processes Can Be Scheduled on an OSD?
How Can I View Information About Pools?
How Can I View Information About a Particular Pool?
How Can I View Information About Hosts?
How Can I View Information About a Particular Host?
Changing Configuration
How Can I Change OSD Configuration Options?
How Can I Change the OSD State?
How Can I Reweight an OSD?
How Can I Change Information for a Pool?
Administering the Cluster
How Can I Run a Scheduled Process on an OSD?
How Can I Create a New Pool?
How can I remove pools?
Ceph Object Gateway administrative API
Administration operations
Administration authentication requests
Creating an administrative user
Get user information
Create a user
Modify a user
Remove a user
Create a subuser
Modify a subuser
Remove a subuser
Add capabilities to a user
Remove capabilities from a user
Create a key
Remove a key
Bucket notifications
Persistent notifications
Creating a topic
Getting topic information
802
803
804
805
805
806
807
808
808
809
809
810
812
812
812
813
814
814
815
816
818
818
819
819
820
822
822
823
823
823
823
824
825
825
825
826
827
828
828
829
830
831
832
832
833
834
835
836
836
836
837
838
838
839
839
840
841
842
842
847
848
849
853
857
857
859
861
862
863
864
866
867
867
867
869
Listing topics
Deleting topics
Using the command-line interface for topic management
Event record
Supported event types
Get bucket information
Check a bucket index
Remove a bucket
Link a bucket
Unlink a bucket
Get a bucket or object policy
Remove an object
Quotas
Get a user quota
Set a user quota
Get a bucket quota
Set a bucket quota
Get usage information
Remove usage information
Standard error responses
Ceph Object Gateway and the S3 API
Prerequisites
S3 limitations
Accessing the Ceph Object Gateway with the S3 API
Prerequisites
S3 authentication
S3 server-side encryption
S3 access control lists
Preparing access to the Ceph Object Gateway using S3
Accessing the Ceph Object Gateway using Ruby AWS S3
Accessing the Ceph Object Gateway using Ruby AWS SDK
Accessing the Ceph Object Gateway using PHP
Secure Token Service
Secure Token Service application programming interfaces
Configuring the Secure Token Service
Creating a user for an OpenID Connect provider
Obtaining a thumbprint of an OpenID Connect provider
Registering the OpenID Connect provider
Creating IAM roles and policies
Accessing S3 resources
Configuring and using STS Lite with Keystone (Technology Preview)
Working around the limitations of using STS Lite with Keystone (Technology Preview)
S3 bucket operations
Prerequisites
S3 create bucket notifications
S3 get bucket notifications
S3 delete bucket notifications
Accessing bucket host names
S3 list buckets
S3 return a list of bucket objects
S3 create a new bucket
S3 put bucket website
S3 get bucket website
S3 delete bucket website
S3 delete a bucket
S3 bucket lifecycle
S3 GET bucket lifecycle
S3 create or replace a bucket lifecycle
S3 delete a bucket lifecycle
S3 get bucket location
S3 get bucket versioning
S3 put bucket versioning
S3 get bucket access control lists
S3 put bucket Access Control Lists
S3 get bucket cors
S3 put bucket cors
S3 delete a bucket cors
S3 list bucket object versions
S3 head bucket
S3 list multipart uploads
S3 bucket policies
S3 get the request payment configuration on a bucket
S3 set the request payment configuration on a bucket
Multi-tenant bucket operations
870
870
871
871
873
873
875
876
877
878
879
880
880
881
881
881
883
883
885
886
887
887
887
887
888
888
889
889
889
890
893
896
898
899
901
902
902
904
904
905
906
908
909
910
910
912
914
914
914
915
917
918
918
919
919
919
920
921
921
921
922
922
922
923
924
924
924
925
926
927
929
931
931
931
S3 Block Public Access
S3 GET PublicAccessBlock
S3 PUT PublicAccessBlock
S3 delete PublicAccessBlock
S3 object operations
Prerequisites
S3 get an object from a bucket
S3 get information on an object
S3 put object lock
S3 get object lock
S3 put object legal hold
S3 get object legal hold
S3 put object retention
S3 get object retention
S3 put object tagging
S3 get object tagging
S3 delete object tagging
S3 add an object to a bucket
S3 delete an object
S3 delete multiple objects
S3 get an object’s Access Control List (ACL)
S3 set an object’s Access Control List (ACL)
S3 copy an object
S3 add an object to a bucket using HTML forms
S3 determine options for a request
S3 initiate a multipart upload
S3 add a part to a multipart upload
S3 list the parts of a multipart upload
S3 assemble the uploaded parts
S3 copy a multipart upload
S3 abort a multipart upload
S3 Hadoop interoperability
S3 select operations (Technology Preview)
Prerequisites
S3 select content from an object
S3 supported select functions
S3 alias programming construct
S3 CSV parsing explained
Ceph Object Gateway and the Swift API
Prerequisites
Swift API limitations
Create a Swift user
Swift authenticating a user
Swift container operations
Prerequisites
Swift container operations
Swift update a container’s Access Control List (ACL)
Swift list containers
Swift list a container’s objects
Swift create a container
Swift delete a container
Swift add or update the container metadata
Swift object operations
Prerequisites
Swift object operations
Swift get an object
Swift create or update an object
Swift delete an object
Swift copy an object
Swift get object metadata
Swift add or update object metadata
Swift temporary URL operations
Swift get temporary URL objects
Swift POST temporary URL keys
Swift multi-tenancy container operations
Ceph RESTful API specifications
Prerequisites
Ceph summary
Authentication
Ceph File System
Storage cluster configuration
CRUSH rules
Erasure code profiles
Feature toggles
932
933
933
934
934
935
935
936
937
938
939
940
940
941
942
942
942
943
943
944
944
945
946
947
947
947
948
948
950
951
952
952
952
952
952
957
958
958
959
959
959
960
961
961
962
962
962
962
964
965
966
966
967
967
967
967
968
969
969
970
970
970
971
971
971
972
972
972
973
974
977
979
980
981
Grafana
Storage cluster health
Host
Logs
Ceph Manager modules
Ceph Monitor
Ceph OSD
Ceph Object Gateway
REST APIs for manipulating a role
Ceph Orchestrator
Pools
Prometheus
RADOS block device
Performance counters
Roles
Services
Settings
Ceph task
Telemetry
Ceph users
S3 common request headers
S3 common response status codes
S3 unsupported header fields
Swift request headers
Swift response headers
Examples using the Secure Token Service APIs
981
982
983
986
986
988
988
994
1001
1003
1004
1005
1007
1017
1020
1022
1023
1025
1025
1026
1028
1029
1029
1029
1029
1030
Troubleshooting
1031
1031
1032
1032
1033
1033
1034
1034
1035
1035
1036
1037
1037
1038
1039
1039
1041
1042
1042
1042
1043
1043
1044
1044
1045
1046
1047
1047
1048
1049
1049
1051
1051
1052
1052
1052
1052
1053
1054
1055
1056
1056
1057
1057
1059
1059
1060
1060
1061
1061
1062
Initial Troubleshooting
Identifying problems
Diagnosing health
Understanding Ceph health
Muting health alerts
Understanding Ceph logs
Generating an sos report
Configuring logging
Ceph subsystems
Configuring logging at runtime
Configuring logging in configuration file
Accelerating log rotation
Creating and collecting operation logs for Ceph Object Gateway
Troubleshooting networking issues
Basic networking troubleshooting
Basic chrony NTP troubleshooting
Troubleshooting Ceph Monitors
Most common Ceph Monitor errors
Ceph Monitor error messages
Common Ceph Monitor error messages in Ceph logs
Ceph Monitor is out of quorum
Clock skew
Ceph Monitor store is getting too big
Understanding Ceph Monitor status
Injecting a monmap
Replacing a failed monitor
Compacting monitor store
Opening port for Ceph Manager
Recovering Ceph Monitor store
Recovering Ceph Monitor store when using BlueStore
Troubleshooting Ceph OSDs
Most common Ceph OSD errors
Ceph OSD error messages
Common Ceph OSD error messages in Ceph logs
Full OSDs
Backfillfull OSDs
Nearfull OSDs
Down OSDs
Flapping OSDs
Slow requests or requests are blocked
Stopping and starting rebalancing
Mounting the OSD data partition
Replacing an OSD drive
Increasing PID count
Deleting data from a full storage cluster
Troubleshooting multi-site Ceph Object Gateway
Error code definitions for Ceph Object Gateway
Syncing a multisite Ceph Object Gateway
Performance counters for multi-site Ceph Object Gateway data sync
Synchronizing data in a multi-site Ceph Object Gateway configuration
Troubleshooting Ceph placement groups
Most common Ceph placement groups errors
Placement group error messages
Ceph subsystems default logging level values
Troubleshooting upgrade error messages
Health messages of a Ceph cluster
1062
1063
1063
1063
1064
1065
1065
1065
1066
1067
1068
1070
1070
1071
1072
1072
1073
1073
1074
1075
1075
1076
1076
1077
1077
1078
1078
1080
1081
1081
1082
1082
1082
1085
1085
1085
Related information
1088
Acknowledgments
1088
Glossary
1089
1089
1089
1089
1090
1091
1091
1092
1092
1092
1093
1093
1093
1093
1093
1094
1094
1094
1095
1096
1096
1096
1097
Stale placement groups
Inconsistent placement groups
Unclean placement groups
Inactive placement groups
Placement groups are down
Unfound objects
Listing placement groups stuck in stale, inactive, or unclean state
Listing placement group inconsistencies
Repairing inconsistent placement groups
Increasing the placement group
Troubleshooting Ceph objects
Troubleshooting high-level object operations
Listing objects
Fixing lost objects
Troubleshooting low-level object operations
Manipulating object’s content
Removing an object
Listing object map
Manipulating object map header
Manipulating object map key
Listing object’s attributes
Manipulating object attribute key
Troubleshooting clusters in stretch mode
Replacing tiebreaker with a monitor in quorum
Replacing tiebreaker with a new monitor
Forcing stretch cluster into recovery or healthy mode
Contacting IBM support for service
Providing information to IBM Support engineers
Generating readable core dump files
Generating readable core dump files in containerized deployments
A
B
C
D
E
F
G
H
I
K
L
M
N
O
P
R
S
T
U
V
W
Z
IBM Storage Ceph
IBM Storage Ceph is a software-defined storage platform engineered for private cloud architectures.
Summary of changes
This topic lists the dates and nature of updates to the published information for IBM Storage Ceph.
Date
28 August 2024
Nature of updates to the published information
Content updated for IBM Storage Ceph 6.1z7 release.
For more information about this release, see Release notes for 6.1z7.
9 May 2024
Content updated for IBM Storage Ceph 6.1z6 release.
For more information about this release, see Release notes for 6.1z6.
29 March 2024
Content updated for IBM Storage Ceph 6.1z5 release.
For a full list of enhancements and fixes in this release, see Release notes for 6.1z5.
All relevant sections in the document are updated.
09 February 2024
Content updated for IBM Storage Ceph 6.1z4 release.
For a full list of enhancements and fixes in this release, see Release notes for 6.1z4.
All relevant sections in the document are updated.
13 December 2023 Content updated for IBM Storage Ceph 6.1z3 release.
For a full list of enhancements and fixes in this release, see Release notes for 6.1z3.
All relevant sections in the document are updated.
13 October 2023
Content updated for IBM Storage Ceph 6.1z2 release.
For a full list of enhancements and fixes in this release, see Release notes for 6.1z2.
All relevant sections in the document are updated.
22 August 2023
The version information was added in IBM Documentation as part of the initial IBM Storage Ceph 6.1 release.
For updates in this release, see Release notes for 6.1.
Compatibility matrix
Use this information for products supported with IBM Storage Ceph 6.1.
Table 1. Support host operating system
Host operating system
Version
Notes
Red Hat Enterprise Linux (RHEL) 9.4, 9.3, 9.2, 8.10, 8.9
Standard lifecycle 9.2 is included in the product.
FIPS mode is supported.
Note: IBM Storage Ceph 6.1 supports Red Hat Enterprise Linux 9 with bootstrap node.
Note: Red Hat Enterprise Linux 9 requires a CPU type x86-64-v2 or higher (v3, v4) which supports SSE4.2 or higher. See Architecture for more information.
Important:
All nodes in the cluster and their clients must use the supported OS version(s) to ensure that the version of the ceph package is the same on all nodes. Using
different versions of the ceph package is not supported.
Ubuntu is not supported as a deploying host operating system.
Table 2. Supported products
Product
Version
Notes
Ansible
Supported in a limited capacity.
Supported for upgrade and conversion to Cephadm and for other minimal
playbooks.
Red Hat
OpenShift
3.x, 4.2 – 4.7
RBD, Cinder, and CephFS drivers are supported.
Red Hat
OpenShift Data
Foundation
See the Red Hat OpenShift Data Foundation Supportability and
Interoperability Checker for detailed external mode version
compatibility.
Red Hat
OpenStack
Platform
17.1
Supports IBM Storage Ceph through both external and director deployed.
Red Hat Satellite
6.x
Only registering with the Content Delivery Network (CDN) is supported.
Registering with Red Hat Network (RHN) is deprecated and not supported.
IBM Storage Ceph 1
Product
Version
Notes
VMware
VMware is supported where each daemon is provided with dedicated
RAM and CPU resources matching bare-metal requirements.
VMware ESX and vSphere are supported to deploy and run IBM
Storage Ceph software on.
When deploying IBM Storage Ceph on VMware, directly attached
storage is strongly recommended.
Note: If using IBM Storage Fusion Data Foundation with disaster
recovery SAN hardware may be used.
Note: Ensure that any bare-metal deployment clusters can economically
scale out.
For more information about hardware requirements, see Minimum hardware
requirements.
For more information about IBM Storage Fusion Data Foundation, see IBM
Storage Fusion Data Foundation within IBM Storage Fusion documentation.
Table 3. Supported client
connector
Client connector
S3A
Version
2.8.x, 3.2.x, and trunk
Table 4. Supported products for IBM Storage Ceph as a backup target
IBM Storage Ceph as a backup target
Version
Notes
CommVault
Cloud Data Management v11
IBM Spectrum Protect Plus
10.1.5
IBM Spectrum Protect server
8.1.8
NetApp AltaVault
4.3.2 and 4.4
Rubrik Cloud Data Management (CDM)
3.2 onwards
Trilio, TrilioVault
3.0
Veeam (object storage)
Veeam Availability Suite 9.5 Update Supported on IBM Storage Ceph object storage with the S3
4
protocol
Veritas NetBackup for Symantec OpenStorage (OST) cloud
backup
7.7 and 8.0
S3 target
Table 5. Independent software vendors
Independent software vendors
Version
IBM Spectrum Discover
2.0.3
WekaIO
3.12.2
Watson X.Data
1.0.3
Release notes for 6.1
IBM Storage Ceph is a hardened, qualified, secure, and supported enterprise software curated from the Ceph open-source project and delivered by IBM.
Enhancements
This section lists all the major updates, enhancements, and new features introduced in this release of IBM Storage Ceph.
Bug fixes
This section describes bugs with significant user impact, which were fixed in this release of IBM Storage Ceph. In addition, the section includes descriptions of fixed
known issues found in previous versions.
Known issues
This section documents known issues found in this release of IBM Storage Ceph.
Technology previews
This section provides an overview of Technology Preview features introduced or updated in this release of IBM Storage Ceph
Sources
Enhancements
This section lists all the major updates, enhancements, and new features introduced in this release of IBM Storage Ceph.
Cephadm utility
Ceph Dashboard
Ceph File System
RADOS
RADOS Block Devices (RBD)
RBD Mirroring
Ceph Object Gateway
Multi-site Ceph Object Gateway
Cephadm utility
public_network parameter can now have configuration options, such as global or mon
2 IBM Storage Ceph
Previously, in cephadm, the public_network parameter was always set as a part of the mon configuration section during a cluster bootstrap without providing any
configuration option to alter this behavior.
With this enhancement, you can specify the configuration options, such as global or mon for the public_network parameter during cluster bootstrap by utilizing
the Ceph configuration file.
The Cephadm commands that are run on the host from the cephadm Manager module now have timeouts
Previously, one of the Cephadm commands would occasionally hang indefinitely, and it was difficult for users to notice and sort the issue.
With this release, timeouts are introduced in the Cephadm commands that are run on the host from the Cephadm mgr module. Users are now alerted with a health
warning about eventual failure if one of the commands hangs. The timeout is configurable with the mgr/cephadm/default_cephadm_command_timeout
setting, and defaults to 900 seconds.
cephadm support for CA signed keys is implemented
Previously, CA signed keys worked as a deployment setup in Red Hat Ceph Storage 5, although their working was accidental, untested, and broken in changes from
Red Hat Ceph Storage 5 to Red Hat Ceph Storage 6.
With this enhancement, cephadm support for CA signed keys is implemented. Users can now use CA signed keys rather than typical pubkeys for SSH authentication
scheme.
Users can now rotate the authentication key for Ceph daemons
For security reasons, some users might desire to occasionally rotate the authentication key used for daemons in the storage cluster.
With this release, the ability to rotate the authentication key for ceph daemons using the ceph orch daemon rotate-key DAEMON_NAME command is
introduced. For MDS, OSD and MGR daemons, this does not require a daemon restart. However, for other daemons, such as Ceph Object Gateway daemons, the
daemon might require restarting to switch to the new key.
Bootstrap logs are now logged to STDOUT
With this release, to reduce potential errors, bootstrap logs are now logged to STDOUT instead of STDERR in successful bootstrap scenarios.
Ceph Object Gateway zonegroup can now be specified in the specification used by the orchestrator
Previously, the orchestrator could handle setting the realm and zone for the Ceph Object Gateway. However, setting the zonegroup was not supported.
With this release, users can specify a rgw_zonegroup parameter in the specification that is used by the orchestrator. Cephadm sets the zonegroup for Ceph Object
Gateway daemons deployed from the specification.
ceph orch daemon add osd now reports if the hostname specified for deploying the OSD is unknown
Previously, since the ceph orch daemon add osd command gave no output, users would not notice if the hostname was incorrect. Due to this, Cephadm would
discard the command.
With this release, the ceph orch daemon add osd command reports to the user if the hostname specified for deploying the OSD on is unknown.
cephadm shell command now reports the image being used for the shell on startup
Previously, users would not always know which image was being used for the shell. This would affect the packages that were used for commands being run within
the shell.
With this release, cephadm shell command reports the image used for the shell on startup. Users can now see the packages being used within the shell, as they
can see the container image being used, and when that image was created as the shell starts up.
Cluster logs under /var/log/ceph are now deleted
With this release, to better clean up the node as part of removing the Ceph cluster from that node, cluster logs under /var/log/ceph are deleted when cephadm
rm-cluster command is run. The cluster logs are removed as long as --keep-logs are not passed to the rm-cluster command.
Note: If the cephadm rm-cluster command is run on a host that is part of a still existent cluster, the host is managed by Cephadm, and the Cephadm mgr
module is still enabled and running, then Cephadm might immediately start deploying new daemons, and more logs could appear.
Better error handling when daemon names are passed to ceph orch restart command
Previously, in cases where the daemon passed to ceph orch restart command was a haproxy or keepalived daemon, it would return a traceback. This made it
unclear to users if they had made a mistake or Cephadm had failed in some other way.
With this release, better error handling is introduced to identify when the users pass a daemon name to ceph orch restart command instead of the expected
service name. Upon encountering a daemon name, Cephadm reports and requests the user to check ceph orch
ls for valid services to pass.
Users can now create Ceph Object Gateway realm, zone, and zonegroup using ceph rgw realm
bootstrap -i rgw_spec.yaml command
With this release, to streamline the process of setting up Ceph Object Gateway on a IBM Storage Ceph cluster, users can create a Ceph Object Gateway realm, zone,
and zonegroup using ceph rgw realm bootstrap -i rgw_spec.yaml command. The specification file should be modeled similar to the one that is used to
deploy Ceph Object Gateway daemons using the orchestrator. The command then creates the realm,zone, and, zonegroup, and passes the specification on to the
orchestrator, which then deploys the Ceph Object Gateway daemons.
rgw_realm: myrealm
rgw_zonegroup: myzonegroup
rgw_zone: myzone
placement:
hosts:
- rgw-host1
- rgw-host2
spec:
rgw_frontend_port: 5500
crush_device_class and location fields are added to OSD specifications and host specifications respectively
With this release, the crush_device_class field is added to the OSD specifications, and the location field, referring to the initial crush location of the host, is added to
host specifications. If a user sets the location field in a host specification, Cephadm runs ceph osd crush add-bucket with the hostname, and the given
location to add it as a bucket in the crush map. For OSDs, they are set with the given crush_device_class in the crush map upon creation.
Note: This is only for OSDs that were created based on the specification with the field set. It does not affect the already deployed OSDs.
Users can enable Ceph Object Gateway manager module
With this release, Ceph Object Gateway Manager module is now available, and can be turned on with ceph mgr module enable rgw command to enable users
to gain access to the functionality of the Ceph Object Gateway manager module, such as ceph rgw realm
IBM Storage Ceph 3
bootstrap, and ceph rgw realm tokens commands.
Users can enable additional metrics for node-exporter daemons
With this release, to enable users to have more customization of their node-exporter deployments, without requiring explicit support for each individual option,
additional metrics are introduced that can now be enabled for node-exporter daemons deployed by Cephadm, using the extra_entrypoint_args field.
service_type: node-exporter
service_name: node-exporter
placement:
label: "node-exporter"
extra_entrypoint_args:
- "--collector.textfile.directory=/var/lib/node_exporter/textfile_collector2"
Users can set the crush location for a Ceph Monitor to replace tiebreaker monitors
With this release, users can set the crush location for a monitor deployed on a host. It should be assigned in the monitor specification file. This is primarily added to
make the replacing of a tiebreaker monitor daemon in stretch clusters deployed by Cephadm, more feasible. Without this change, users would have to manually edit
the files written by Cephadm to deploy the tiebreaker monitor, as the tiebreaker monitor is not allowed to join without declaring its crush location.
service_type: mon
service_name: mon
placement:
hosts:
- host1
- host2
- host3
spec:
crush_locations:
host1:
- datacenter=a
host2:
- datacenter=b
- rack=2
host3:
- datacenter=a
crush_device_class can now be specified per path in an OSD specification
With this release, to allow users more flexibility with crush_device_class settings when deploying OSDs through Cephadm, crush_device_class, you can
specify per path inside an OSD specification. It is also supported to provide these per-path crush_device_classes along with a service-wide
crush_device_class for the OSD service. In cases of service-wide crush_device_class, the setting is considered as default, and the path-specified settings
take priority.
service_type: osd
service_id: osd_using_paths
placement:
hosts:
- Node01
- Node02
crush_device_class: hdd
spec:
data_devices:
paths:
- path: /dev/sdb
crush_device_class: ssd
- path: /dev/sdc
crush_device_class: nvme
- /dev/sdd
db_devices:
paths:
- /dev/sde
wal_devices:
paths:
- /dev/sdf
The Cephadm commands run on the host from the cephadm mgr module now have timeouts
Previously, one of the Cephadm commands would occasionally hang indefinitely, and it was difficult for users to notice and sort the issue.
With this release, timeouts are introduced in the Cephadm commands that are run on the host from the Cephadm mgr module. Users are now alerted with a health
warning about eventual failure if one of the commands hangs. The timeout is configurable with the mgr/cephadm/default_cephadm_command_timeout
setting, and defaults to 900 seconds.
Cephadm now raises a specific health warning UPGRADE_OFFLINE_HOST when the host goes offline during upgrade
Previously, when upgrades failed due to a host going offline, a generic UPGRADE_EXCEPTION health warning would be raised that was too ambiguous for users to
understand.
With this release, when an upgrade fails due to a host being offline, Cephadm raises a specific health warning - UPGRADE_OFFLINE_HOST, and the issue is now
made transparent to the user.
All the Cephadm logs are no longer logged into cephadm.log when --verbose is not passed
Previously, some Cephadm commands, such as gather-facts, would spam the log with massive amounts of command output every time they were run. In some
cases, it was once per minute.
With this release, in Cephadm, all the logs are no longer logged into cephadm.log when --verbose is not passed. The cephadm.log is now easier to read since
most of the spam previously written is no longer present.
Ceph Dashboard
A metric to track slow operation per daemon is added
Previously, tracking slow operations was cumbersome as it required logging parsing.
4 IBM Storage Ceph
With this release, a metric to track slow operations in Ceph daemon is added in Prometheus.
Authx management feature in IBM Storage Ceph Dashboard
With this release, to improve management of Authx users, Authx management feature is added in the IBM Storage Ceph Dashboard UI. Users can now create,
edit, and delete Authx users from the dashboard.
IBM Storage Ceph Dashboard has a new landing page
With this release, the IBM Storage Ceph 3.0 has a new landing page to improve the visualization. The new view is categorized into 5 cards, namely, Details, Status,
Capacity, Inventory, and Cluster utilization. Following are the highlights of the new landing page:
Details such as FSID, Orchestrator, and the Ceph versions.
Status such as graphical display of the health of the cluster.
Capacity shows the overall capacity of the cluster.
Inventory holds all the logical entries of the cluster.
Cluster utilization includes charts such as Used Capacity, IOPS, Latency Client Throughput, and Recovery throughput.
Recovery throughput metrics are displayed in the landing page
With this release, recovery throughput metrics are added to the Cluster utilization card and displayed in the landing page for ease of user experience.
A new metric is added for OSD blocklist count
With this release, to configure a corresponding alert, a new metric ceph_cluster_osd_blocklist_count is added on the Ceph Dashboard.
Support for Ceph client authorization on the dashboard
With this release, you can create, edit, list, import, export, and, delete users with different capabilities on the IBM Storage Ceph Dashboard.
Introduction of ceph-exporter daemon
With this release, ceph-exporter daemon is introduced to collect and expose performance counters of all Ceph daemons as Prometheus metrics. It is deployed
on each node of the cluster to be performant at large scale clusters.
Support force promote for RBD mirroring through Dashboard
Previously, although RBD mirror promote/demote was implemented on the Ceph Dashboard, there was no option to force promote.
With this release, support for force promoting RBD mirroring through Ceph Dashboard is added. If the promotion fails on the Ceph Dashboard, the user is given the
option to force the promotion.
Support for collecting and exposing the labeled performance counters
With this release, support for collecting and exposing the labeled performance counters of Ceph daemons as Prometheus metrics with labels is introduced.
Ceph File System
Switch the unfair Mutex lock to fair mutex
Previously, the implementations of the Mutex, for example, std::mutex in C++, would not guarantee fairness and would not guarantee that the lock would be
acquired by threads in the order called lock(). In most cases, this worked well but in an overloaded case, the client requests handling thread and submit thread
would always successfully acquire the submit_mutex in a long time, causing MDLog::trim() to get stuck. That meant the MDS daemons would fill journal logs into
the metadata pool, but could not trim the expired segments in time.
With this enhancement, the unfair Mutex lock is switched to fair mutex and all the submit_mutex waiters are woken up one by one in FIFO mode.
cephfs-top limitation is increased for more client loading
Previously, due to a limitation in the cephfs-top utility, only less than 100 clients were loaded at a time, which could not be scrolled and also hung if more clients
were loaded.
With this release, cephfs-top users can scroll vertically, as well as, horizontally. This enables cephfs-top to load nearly 10,000 clients. The users can scroll the
loaded clients, and view them on the screen.
Users now have the option to sort clients based on the fields of their choice in cephfs-top
With this release, users have the option to sort the clients based on the fields of their choice in cephfs-top and also, limit the number of clients to be displayed.
This enables the user to analyze the metrics based on the order of fields as per requirement.
Non-head omap entries are now included in the omap entries
Previously, a directory fragment would not split if non-head snapshotted entries were not taken into account when deciding to merge or split a fragment. Due to
this, the number of omap entries in a directory object would exceed a certain limit, and result in cluster warnings.
With this release, non-head omap entries are included in the number of omap entries when deciding to merge or split a directory fragment to never exceed the limit.
RADOS
Usability and design improvements are implemented to the mClock scheduler
With this release, significant usability and design improvements are implemented to the mClock scheduler to address the slow backfill issues.
The balanced profile is set as the default mClock profile because it represents a compromise between prioritizing client IO or recovery IO. Users can then
choose either the high_client_ops profile to prioritize client IO or the high_recovery_ops profile to prioritize recovery IO.
QoS parameters like reservation and limit are now specified in terms of a fraction (range: 0.0 to 1.0) of the OSD’s IOPS capacity.
The cost parameters - osd_mclock_cost_per_io_usec_* and osd_mclock_cost_per_byte_usec_* have been removed. The cost of an operation is
now determined using the random IOPS and maximum sequential bandwidth capability of the OSD’s underlying device.
Degraded object recovery is given higher priority when compared to misplaced object recovery because degraded objects present a data safety issue not
present with objects that are merely misplaced. Therefore, backfilling operations with the balanced and high_client_ops mClock profiles may progress
slower than what was seen with the WeightedPriorityQueue (WPQ) scheduler.
The QoS allocations in all the mClock profiles are optimized based on the above fixes and enhancements.
RADOS Block Devices (RBD)
IBM Storage Ceph 5
Object map for the snapshot accurately reflects the contents of the snapshot
Previously, due to an implementation defect, a stale snapshot context would be used when handling a write-like operation. Due to this, the object map for the
snapshot was not guaranteed to accurately reflect the contents of the snapshot in case the snapshot was taken without quiescing the workload. In differential
backup and snapshot-based mirroring, use cases with object-map and/or fast-diff features enabled, the destination image could get corrupted.
With this enhancement, the implementation defect is fixed and everything works as expected.
RBD Mirroring
More fields added to rbd mirror status command
Previously, the rbd mirror status command lacked important information about the time taken to sync the bytes in the last snapshot, due to which monitoring
and troubleshooting the mirroring process effectively was challenging.
With this release, additional fields, such as last_snapshot_bytes and last_snapshot_seconds are added in the rbd mirror status command providing
administrators with valuable information about the performance of the mirroring process. This enables efficient monitoring and troubleshooting snapshot-based
mirroring operations.
Renaming and adding performance counters is supported
With this release, you can rename and add performance counters in rbd-mirror daemon to improve clarity and provide detailed information about snapshot
synchronization between source and destination clusters. The renamed existing snapshot and journal-based performance counters added new performance
counters and switched to using labels instead of image specification in counter names.
Ceph Object Gateway
Users can now enable data transition to Azure
With this enhancement, users can enable data transition to a remote cloud service, such as Azure, as part of the lifecycle configuration. See Transitioning data to
Azure cloud service for more details.
D3N is implemented in Ceph Object Gateway
With this release, Data-Delivery Network (D3N) is implemented, which uses high-speed storage such as NVMe flash to cache datasets on the access side, allowing
big data jobs to use the compute and fast storage resources available on each Ceph Object Gateway node.
See D3N Data Cache in the IBM Storage Ceph Object Gateway guide for more details.
Multi-site Ceph Object Gateway
Users can archive older data to an AWS bucket
With this release, users can enable data transition to a remote cloud service, such as Amazon Web Services (AWS), as part of the lifecycle configuration.
See Transitioning data to Amazon S3 cloud service for more details.
Bucket granular multi-site sync policies is now supported
IBM now supports bucket granular multi-site sync policies.
Important: With the agreed limited scope for the GA release of the bucket granular sync policy in 6.1, the following features are now supported:
Greenfield deployment: This release only supports new multi-site deployments. To set up bucket granular sync replication, a new zonegroup/zone must be
configured at a minimum. Please note that this release does not allow migrating deployed/running RGW multi-site replication configurations to the newly
featured RGW bucket sync policy replication.
Data flow - symmetrical: Although both unidirectional and bi-directional/symmetrical replication can be configured, only symmetrical replication flows are
supported in this release.
1-to-1 bucket replication: Currently, only replication between buckets with identical names is supported. This means that, if the bucket on site A is named
bucket1, it can only be replicated to bucket1 on site B. Replicating from bucket1 to bucket2 in a different zone is not currently supported.
The following features are not supported in this release:
Source filters
Storage class
Destination owner translation
User mode
See Using multi-site sync policies in the IBM Storage Ceph Object Gateway guide for more details.
Bucket notifications are sent when an object is synced to a zone
Previously, bucket notifications would not be sent when an object got synced to a zone.
With this release, bucket notifications are sent when an object is synced to a zone, to allow external systems to receive information into the zone syncing status at
object-level. The following bucket notification event types are added - s3:ObjectSynced:* and s3:ObjectSynced:Created. When configured with the bucket
notification mechanism, a notification event is sent from the synced Ceph Object Gateway upon successful sync of an object.
Note: Both topics and the notification configuration should be done separately in each zone from which you would like to see the notification events being sent.
Disable per-bucket replication when zones replicate by default
With this release, the ability to disable per-bucket replication when the zones replicate by default, using multisite sync policy, is introduced to ensure that selective
buckets can opt out.
Server-Side encryption is now supported
With this release, Server-Side Encryption S3 (SSE-S3) is supported. This enables S3 users to protect data at rest with a unique key through Server-Side encryption
with Amazon S3-managed encryption keys (SSE-S3).
Users can use the PutBucketEncryption S3 feature to enforce object encryption
Previously, to enforce object encryption in order to protect data, users were required to add a header to each request which was not possible in all cases.
With this release, Ceph Object Gateway is updated to support PutBucketEncryption
S3 action. Users can use the PutBucketEncryption S3 feature with the Ceph Object Gateway without adding headers to each request. This is handled by the
Ceph Object Gateway.
Deploy Ceph Object Gateway multi-site with rgw module
6 IBM Storage Ceph
With this release, you can use the Ceph Manager rgw module to deploy realms, zonegroups, and different related entities. You can also configure the secondary
zone with the generated tokens.
Bug fixes
This section describes bugs with significant user impact, which were fixed in this release of IBM Storage Ceph. In addition, the section includes descriptions of fixed
known issues found in previous versions.
Ceph Manager plug ins
Cephadm utility
Ceph Dashboard
Ceph Metrics
Ceph File System
Ceph Volume utility
Ceph Object Gateway
Multi-site Ceph Object Gateway
RADOS
RADOS Block Devices (RBD)
RBD Mirroring
Ceph Manager plug ins
A new Ceph manager module option, exclude_perf_counters, is introduced
Previously, after the introduction of the new Ceph exporter, metrics coming from Ceph daemon performance counters were exposed by Ceph exporter and
Prometheus Manager module. Due to this, metrics based on Ceph daemon performance counters were duplicated.
With this fix, a new Ceph Manager module option, exclude_perf_counters, is introduced. By default, it is set to True, preventing the Prometheus Manager
module from exporting the metrics coming from the performance counters.
(2186549)
Python tasks no longer wait for the GIL
Previously, the Ceph manager daemons held the Python Global Interpreter Lock (GIL) during some RPCs. Due to this, other python tasks in the queue starved
waiting for the GIL.
With this fix, the GIL is released during all libcephfs or librbd calls and other Python tasks can acquire the GIL normally.
(BZ#2219440)
Emails generated are not flagged as spam
Previously, the email header created by the Ceph alert manager module would not include the message-id and date fields and the mails would get flagged as spam.
With this release, the email header is modified to include these two fields and the emails generated by the module are not flagged as spam.
(BZ#2210906)
Cephadm utility
The config or keyring files are no longer temporarily removed from the nodes
Previously, in cephadm, the calculation of where client config/keyrings files should be written would occasionally end up empty. Due to this, config/keyring files are
temporarily removed from the nodes where they should be left alone until the next checks and the correct calculation.
With this fix, the timing of the calculation is altered and the keyring/conf files are no longer randomly removed from the nodes temporarily.
(BZ#2161545)
Message about the limit policy is lowered to debug-level
Previously, cephadm would log about hitting the limit policy of an OSD specification at info-level. Due to this, every time the limit policy was hit, cephadm would
check if any new OSDs should be deployed to match the specification, resulting in the logs to be spammed with messages about hitting the limit.
With this fix, the message about the limit policy is lowered to debug-level and the logs are no longer spammed with messages about limit policy unless they set the
log level to debug.
(BZ#2207480)
Error-handling is updated to work with the python version used in IBM Storage Ceph 6
Previously, a new timeout on cephadm commands added in Red Hat Ceph Storage 6.1 would not work correctly due to a python version difference affecting the
type of error generated. Due to this, whenever one of these timeouts actually happened, the generated error was not handled and the cephadm module would
crash.
With this fix, the error-handling is updated to work with the python version used in Red Hat Ceph Storage 6.1. When one of these timeouts occurs, it results in a
health warning about what timed out, and the cephadm module no longer crashes.
(BZ#2209493)
Setting or getting the global configuration works when using the ceph config
get command
Previously, running the ceph config get command would not retrieve any value when the entity was set as global. Due to this, the module would fail because
the module first tried to get the current value of the option for idempotency concern.
IBM Storage Ceph 7
With this fix, checks are made by running the ceph config dump command and setting or getting the global configuration works.
(BZ#2190187)
Special lines are no longer included in the host’s /etc/hosts file when mounting into the container
Previously, podman versions added a special line to the /etc/hosts file inside the container which messed with the host name resolution within the container,
causing it to think that the FQDN of the current host is host.containers.internal.
With this fix, the host’s /etc/hosts file is mounted into the container without the special lines and users can use /etc/hosts for hostname resolution. Users will
no longer encounter errors related to being unable to find the IP for host.container.internal while accessing Grafana graphs in the Ceph dashboard.
(BZ#2216295)
Scheduled daemon actions no longer get dropped on mgr failover
Previously, scheduled daemon actions would get dropped on Ceph Manager failover. Due to this, if users sent a ceph orch daemon redeploy
MANAGER_DAEMON_NAME, where the Ceph Manager daemon specified is active, it would just failover, and never actually redeploy. Users would have to run the
command again, once that manager daemon became a standby.
With this fix, scheduled daemon actions are saved somewhere persistent when they are scheduled. When you run ceph
orch daemon redeploy MANAGER_DAEMON_NAME command specifying the active manager, It will first failover control to another manager, and then the other
manager redeploys the previously active manager specified in the command.
(BZ#2097187)
The manager daemons correctly identify they have been upgraded and no longer failover
Previously, a check for whether the current active manager was running the upgraded version of cephadm would not work, and sometimes, managers that were
upgraded would think that they were on an older version. Due to this, the upgrade would get into a state where all the mgr daemons were upgraded, but they were
still in line to be redeployed by a manager using the upgraded version of cephadm, causing it to repeatedly just failover between the manager daemons.
With this fix, the check is rectified and now, the manager daemons are aware of their correct version, and can redeploy other daemons using the upgraded version
of cephadm.
(BZ#2108489)
Bootstrap no longer fails if a comma-separated list of quoted IPs are passed in as the public network in the initial Ceph configuration
Previously, Cephadm bootstrap would improperly parse comma-delimited lists of IP addresses, if the list was quoted. Due to this, the bootstrap would fail if a
comma-separated list of quoted IP addresses, for example, 172.120.3.0/24, 172.117.3.0/24, 172.118.3.0/24 ,172.119.3.0/24, was provided as the
public_network in the initial Ceph configuration passed to bootstrap with the --config parameter.
(BZ#2111680)
Cephadm no longer attempts to parse the provided yaml files more than necessary
Previously, Cephadm bootstrap would attempt to manually parse the provided yaml files more than necessary. Due to this, sometimes, even if the user had provided
a valid yaml file to Cephadm bootstrap, the manual parsing would fail, depending on the individual specification, causing the entire specification to be discarded.
With this fix, Cephadm no longer attempts to parse the yaml more than necessary. The host specification is searched only for the purpose of spreading SSH keys.
Otherwise, the specification is just passed up to the manager module. The cephadm bootstrap --apply-spec command now works as expected with any valid
specification.
(BZ#2112309)
host.containers.internal entry is no longer added to the /etc/hosts file of deployed containers
Previously, certain podman versions would, by default, add a host.containers.internal entry to the /etc/hosts file of deployed containers. Due to this,
issues arose in some services with respect to this entry, as it was misunderstood to represent FQDN of a real node.
With this fix, Cephadm mounts the host’s /etc/hosts file when deploying containers. The host.containers.internal entry in the /etc/hosts file in the
containers, is no longer present, avoiding all bugs related to the entry, although users can still see the host’s /etc/hosts for name resolution within the container.
(BZ#2133549)
Cephadm now logs device information only when an actual change occurs
Previously, Cephadm would compare all fields reported for OSDs, to check for new or changed devices. But one of these fields included a timestamp that would
differ every time. Due to this, Cephadm would log Detected new or changed devices every time it refreshed a host’s devices, regardless of whether anything
actually changed or not.
With this fix, the comparison of device information against previous information no longer takes the timestamp fields into account that are expected to constantly
change. Cephadm now logs only when there is an actual change in the devices.
(BZ#2136336)
The generated Prometheus URL is now accessible
Previously, if a host did not have an FQDN, the Prometheus URL generated would be http://host-shortname:9095, and it would be inaccessible.
With this fix, if no FQDN is available, the host IP is used over the shortname. The URL generated for Prometheus is now in a format that is accessible, even if the host
Prometheus is deployed on a service that has no FQDN available.
(BZ#2153726)
Cephadm no longer has permission issues while writing files to the host
Previously, cephadm would first create files within the /tmp directory, and then move them to their final location. Due to this, in certain setups, a permission issue
would arise when writing files, making cephadm effectively unable to operate until permissions were modified.
With this fix, Cephadm uses a subdirectory within /tmp to write files to the host that do not have the same permission issues.
(BZ#2182035)
8 IBM Storage Ceph
Ceph Dashboard
No CephPGImbalance alerts are fired on the dashboard
Previously, multiple Ceph PG imbalance alerts were fired as the query was not suitable.
With this fix, the query is fixed and no multiple CephPGImbalance alerts are fired.
(BZ#2110290)
No CephPoolGrowthWarning alerts are fired on the dashboard
Previously, incorrect query for CephPoolGrowthWarning alert caused “Evaluating rule failed” errors to repeat indefinitely in Prometheus logs of a stretch cluster.
With this release, the query is fixed and no errors are observed.
(BZ#2114835)
The default option in the OSD creation step of Expand Cluster wizard works as expected
Previously, the default option in the OSD creation step of Expand Cluster wizard was not working on the dashboard, causing the user to be misled by showing the
option as “selected”.
With this fix, the default option works as expected. Additionally, a “Skip” button is added if the user decides to skip the step.
(BZ#2111751)
Users can create normal or mirror snapshots
Previously, even though the users were allowed to create a normal image snapshot and mirror image snapshot, it was not possible to create a normal image
snapshot.
With this fix, the user can choose from two options to select either normal or mirror image snapshot modes.
(BZ#2145104)
Flicker no longer occurs on the Host page
Previously, the host page would flicker after 5 seconds if there were more than 1 hosts, causing a bad user experience.
With this fix, the API is optimized to load the page normally and the flicker no longer occurs.
(BZ#2164327)
Ceph Metrics
The metrics names produced by Ceph exporter and prometheus manager module are the same
Previously, the metrics coming from the Ceph daemons (performance counters) were produced by the Prometheus manager module. The new Ceph exporter would
replace the Prometheus manager module, and the metrics name produced would not follow the same rules applied in the Prometheus manager module. Due to
this, the name of the metrics for the same performance counters were different depending on the provider of the metric (Prometheus manager module or Ceph
exporter).
With this fix, the Ceph exporter uses the same rules as the ones in the Prometheus manager module to generate metric names from Ceph performance counters.
The metrics produced by Ceph exporter and Prometheus manager module are exactly the same.
(BZ#2186557)
Ceph File System
Directory permissions are correctly mirrored from source clusters to remote clusters
Previously, cephfs-mirror would not update top-level directory permissions after the initial directory creation. Due to this, top-level directory permissions on remote
clusters were not mirrored correctly when permissions were updated on source clusters.
With this fix, top-level directory permissions are synced on every snapshot sync and directory permissions are perfectly mirrored.
(BZ#2160542)
The specified path is validated when creating a Ceph File System export over NFS
Previously, anything could be specified as a path, such as a file, symlink, or a non-existent path and it would not be validated while creating a Ceph File System
export over NFS and the export would be successfully created. Due to this, mounting the export would fail with an error message.
With this fix, the specified path is validated when creating the export and the operator is informed with a proper error message.
(BZ#2170739)
The fallocate path clears the suid/sgid if an unprivileged user changes the file
Previously, the fallocate path would not clear the suid/sgid if an unprivileged user changed the file. There is no Posix item that requires clearing the suid/sgid in
fallocate path but this is the default behaviour for most of the filesystems and the VFS layer. So, the user space libcephfs client would not comply with most
filesystems in the kernel and this could be easily hacked.
With this fix, the fallocate path clears the suid/sgid if an unprivileged user changes the file, making the user space libcephfs client comply with most other
filesystems and fix the attack hole.
(BZ#2185710)
The recovered files under the lost+found directory can now be deleted in Ceph File System
With this fix, after recovering a Ceph File System post the disaster recovery the recovered files under the lost+found directory can be deleted.
(BZ#2222231)
IBM Storage Ceph 9
snap-schedule module no longer throws any traceback
Previously, due to an incorrect subvolume path resolution, traceback was seen on command-line whenever a bad path was provided along with a subvolume
argument.
With this fix, users are recommended to ignore the user-specified path for subvolumes and snap-schedule module does not throw any traceback on the commandline interface.
(BZ#2196748)
MDS no longer crashes when allocating CInode
Previously, when replaying the journals, if the inodetable or the sessionmap versions did not match, the CInode would be added to the inode_map. But the ino
may still be in the inodetable or sessions' `prealloc inos list. Due to this, when allocating a new ino number, if the corresponding CInode was already in
the inode_map, the MDS would crash.
With this fix, allocating ino# is skipped when allocating the new CInode and the MDS does not crash.
(BZ#2189132)
Creation of pool-level snaps for pools actively associated with a filesystem is disallowed
Previously, the ceph osd pool mksnap command allowed the creation of pool-level snaps for pools actively associated with a filesystem. Due to this, there
would be possible data loss when snapshots were deleted from either the filesystem or the pool due to pool ID collision.
With this fix, creation of pool-level snaps for pools actively associated with a filesystem is disallowed and no data loss occurs.
(BZ#2189787)
Link requests no longer fail with -EXDEV
Previously, if an inode had more than one link and after one of its dentries was unlinked, it would be moved to a stray directory. Before the link merge/migrate
finished, if a link request came, it would fail with -EXDEV error. While in non-multiple link cases, it was possible that the clients could pass one invalidate ino, which
is still under unlinking. Due to this, some link requests would fail directly.
With this fix, if users wait for the link merge, migrate or purge to finish, no link requests fail with -EXDEV.
(BZ#2196405)
The minimum compatible python version for cephfs-top is 3.6.0
With this fix, the minimum compatible python version for cephfs-top is 3.6.0. It checks if the current python version is greater than or equal to 3.6.0 during the
build, and thus ensures that the cephfs-top curses display launches successfully.
(BZ#2203165)
Structure variables are no longer stale or unsafe when accessed after session reconnection
Previously, the Ceph File System user-space clients could access stale/unsafe structure variables when rebuilding a request and this would lead to the clients
misbehaving sometimes after reconnecting to the Ceph Manager daemons while re-issuing requests.
With this fix, the structure variables are no longer stale or unsafe when accessed after session reconnection. This is ensured by deep-copying them instead of
shallow-copying and the Ceph File System (CephFS) user-space clients work as expected.
(BZ#2203906)
Calls to Ceph Manager daemons and volumes no longer return -ESHUTDOWN
Previously, calls to Ceph Manager daemons and volumes would return -ESHUTDOWN when the ceph-mgr process was shutting down, which was not necessary.
With this fix, Ceph Manager plugin handles the shutdown without returning a special error code and -ESHUTDOWN is never returned.
(BZ#2150306)
Support for snap-schedule module
With this fix, Ceph continues to provide the subvolume interface to the snap-schedule module of Ceph Manager.
(BZ#2187659)
Deadlock no longer occurs between unlink requests and stray dentry reintegration
Previously, there was a race between unlink requests and stray dentry reintegration when trying to manipulate the same inodes. This caused a deadlock between
unlink requests and stray dentry reintegration.
With this fix, for any unlink request, checks are done to determine whether the corresponding inodes are under reintegration. If inodes are under reintegration, the
unlink requests need to be added to the waiter list and wait for the reintegration to finish and then continue. Deadlocks no longer occur.
(BZ#2188460)
MDLog replay thread does not deadlock anymore and log events get replayed as expected
Previously, a bug in the MDS journaler machinery would cause the MDLog replay thread to infinitely block. Due to this, MDLog replayer thread would infinitely block
causing log events from being replayed.
With this fix, the race condition that causes the thread to get deadlock is never hit. This is done by refactoring the logic that acquires locks. MDLog replay thread
does not deadlock anymore and log events get replayed as expected.
(BZ#2161483)
mtime and change_attr are now updated for snapshot directory when snapshots are created
Previously, libcephfs clients would not update mtime, and would change the attribute when snaps were created or deleted. Due to this, NFS clients could not list
CephFS snapshots within a CephFS NFS-Ganesha export correctly.
With this fix, mtime and change_attr are updated for the snapshot directory, .snap, when snapshots are created, deleted, and renamed. Correct mtime and
change_attr ensure that listing snapshots do not return stale snapshot entries.
10 IBM Storage Ceph
(BZ#1975689)
cephfs-top -d [--delay] option accepts only integer values ranging between 1 to 25
Previously, cephfs-top -d [--delay] option would not work properly, due to the addition of a few new curses methods. The new curses method would accept
only integer values, due to which an exception was thrown on getting the float values from a helper function.
With this fix, cephfs-top -d [--delay] option accepts only integer values ranging between 1 and 25, and cephfs-top utility works as expected.
(BZ#2136031)
Creating same dentries after the unlink finishes does not crash the MDS daemons
Previously, there was a racy condition between unlink and creating operations. Due to this, if the previous unlink request was delayed due to any reasons, and
creating same dentries was attempted during this time, it would fail by crashing the MDS daemons or new creation would succeed but the written content would be
lost.
With this fix, users need to ensure to wait until the unlink finishes, to avoid conflict when creating the same dentries.
(BZ#2140784)
Non-existing cluster no longer shows up when running the ceph nfs cluster info
CLUSTER_ID command.
Previously, existence of a cluster would not be checked when ceph nfs cluster info
CLUSTER_ID command was run, due to which, information of the non-existing cluster would be shown, such as virtual_ip and backend, null, and empty
respectively.
With this fix, the ceph nfs cluster info CLUSTER_ID command checks the cluster existence and an Error ENOENT: cluster does not exist is thrown in case a
non-existing cluster is queried.
(BZ#2149415)
The snap-schedule module no longer incorrectly refers to the volumes module
Previously, the snap-schedule module would incorrectly refer to the volumes module when attempting to fetch the subvolume path. Due to using the incorrect
name of the volumes module and remote method name, the ImportError traceback would be seen.
With this fix, the untested and incorrect code is rectified, and the method is implemented and correctly invoked from the snap-schedule CLI interface methods. The
snap-schedule module now correctly resolves the subvolume path when trying to add a subvolume level schedule.
(BZ#2153196)
Integer overflow and ops_in_flight value overflow no longer happens
Previously, _calculate_ops would rely on a configuration option filer_max_purge_ops, which could be modified on the fly too. Due to this, if the value of
ops_in_flight is set to more than uint64’s capability, then there would be an integer overflow, and this would make ops_in_flight far more greater than
max_purge_ops and it would not be able to go back to a reasonable value.
With this fix, the usage of filer_max_purge_ops in ops_in_flight is ignored, since it is already used in Filer::_do_purge_range(). Integer overflow and
ops_in_flight value overflow no longer happens.
(BZ#2159307)
Invalid OSD requests are no longer submitted to RADOS
Previously, when the first dentry had enough metadata and the size was larger than max_write_size, an invalid OSD request would be submitted to RADOS. Due
to this, RADOS would fail the invalid request, causing CephFS to be read-only.
With this fix, all the OSD requests are filled with validated information before sending it to RADOS and no invalid OSD requests cause the CephFS to be read-only.
(BZ#2160598)
MDS now processes all stray directory entries.
Previously, a bug in the MDS stray directory processing logic caused the MDS to skip processing a few stray directory entries. Due to this, the MDS would not
process all stray directory entries, causing deleted files to not free up space.
With this fix, the stray index pointer is corrected, so that the MDS processes all stray directories.
(BZ#2161479)
Pool-level snaps for pools attached to a Ceph File System are disabled
Previously, the pool-level snaps and fs-level snaps had their own snap ID namespace and this caused a clash between the IDs, and the Ceph Monitor was unable to
uniquely identify a snap as to whether it is a pool-level snap or a fs-level snap. Due to this, there were chances for the wrong snap to get deleted when referring to
an ID, which is present in the set of pool-level snaps and fs-level snaps.
With this fix, the pool-level snaps for the pools attached to a Ceph File System are disabled and no clash of pool IDs occurs. Hence, no unintentional data loss
happens when a CephFS snap is removed.
Note: It was possible to still create pool-level snapshots for the pools using the below two commands:
rados mksnap snap_name
ceph osd mksnap pool
snap
With the current fix, creating a snapshot at pool-level is blocked when the rados mksnap command is used. However, the ceph osd mksnap pool
snap command still creates the pool-level snapshots, which may result in a known issue.
(BZ#2168541)
Client requests no longer bounce indefinitely between MDS and clients
Previously, there was a mismatch between the Ceph protocols for client requests between CephFS client and MDS. Due to this, the corresponding information
would be truncated or lost when communicating between CephFS clients and MDS, and the client requests would indefinitely bounce between MDS and clients.
IBM Storage Ceph 11
With this fix, the type of the corresponding members in the protocol for the client requests is corrected by making them the same type and the new code is made to
be compatible with the old Cephs. The client request does not bounce between MDS and clients indefinitely, and stops after being well retried.
(BZ#2172791)
Support for snap-schedule module
With this fix, Ceph continues to provide the subvolume interface to the snap-schedule module of Ceph Manager.
(BZ#2187659)
A code assert is added to the Ceph Manager daemon service to detect metadata corruption
Previously, a type of snapshot-related metadata corruption would be introduced by the manager daemon service for workloads running Postgres, and possibly
others.
With this fix, a code assert is added to the manager daemon service which is triggered if a new corruption is detected. This reduces the proliferation of the damage,
and allows the collection of logs to ascertain the cause.
Note: If daemons crash after the cluster is upgraded to IBM Storage Ceph 6.1, contact IBM support for analysis and corrective action.
(BZ#2175307)
MDS daemons no longer crash due to sessionmap version mismatch issue
Previously, MDS sessionmap journal log would not correctly persist when MDS failover occurred. Due to this, when a new MDS was trying to replay the journal logs,
the sessionmap journal logs would mismatch with the information in the MDCache or the information from other journal logs, causing the MDS daemons to trigger
an assert to crash themselves.
With this fix, trying to force replay the sessionmap version instead of crashing the MDS daemons results in no MDS daemon crashes due to sessionmap version
mismatch issue.
(BZ#2182564)
MDS no longer gets indefinitely stuck while waiting for the cap revocation acknowledgement
Previously, if __setattrx() failed, the _write() would retain the CEPH_CAP_FILE_WR caps reference, the MDS would be indefinitely stuck waiting for the cap
revocation acknowledgment. It would also cause other clients' requests to be stuck indefinitely.
With this fix, the CEPH_CAP_FILE_WR caps reference is released if the __setattrx() fails and MDS' caps revoke request is not stuck.
(BZ#2182613)
Ceph Volume utility
Devices already used by Ceph are filtered out in ceph-volume
Previously, due to a bug, ceph-volume would not filter out devices already used by Ceph. Due to this, adding new OSDs with ceph-volume failed when using precreated LVs.
With this fix, devices already used by Ceph are filtered out in ceph-volume as expected and new OSDs with pre-created LVs can now be added.
(BZ#2188246)
The correct size is calculated for each database device in ceph-volume
Previously, as of RHCS 4.3, ceph-volume would not make a single VG with all database devices inside, since each database device had its own VG. Due to this, the
database size was calculated differently for each LV.
With this release, the logic is updated to take into account the new database devices with LVM layout. The correct size is calculated for each database device.
(BZ#2185588)
Ceph Object Gateway
The bucket listing feature enables the rgw-restore-bucket-index tool to complete reindexing
Previously, the rgw-restore-bucket-index tool would restore the bucket’s index partially until the next user listed out the bucket. Due to this, the bucket’s statistics
would report incorrectly until the reindexing completed.
With this enhancement, the bucket listing feature is added which enables the tool to complete the reindexing and the bucket statistics are reported correctly.
Additionally, a small change to the build process is added that would not affect end-users.
(BZ#2182456)
Lifecycle transition no longer fails for objects with modified metadata
Previously, setting an ACL on an existing object would change its mtime due to which lifecycle transition failed for such objects.
With this fix, unless it is a copy operation, the object’s mtime remains unchanged while modifying just the object metadata, such as setting ACL or any other
attributes.
(BZ#2213801)
Support for AWS PublicAccessBlock
Previously, Ceph Object Storage did not support the AWS public access block S3 APIs.
With this release, Ceph Object Storage supports the AWS public access block S3 APIs such as PutPublicAccessBlock.
(BZ#2064260)
Swift object storage dialect now includes support for SHA-256 and SHA-512 digest algorithms
Previously, support for digest algorithms was added by OpenStack Swift in 2022, but Ceph Object Gateway had not implemented them.
12 IBM Storage Ceph
With this release, Ceph Object Gateway’s Swift object storage dialect now includes support for SHA-256 and SHA-512 digest methods in tempurl operations.
Ceph Object Gateway can now correctly handle tempurl operations by recent OpenStack Swift clients.
(BZ#2105950)
Topic creation is now allowed with or without trailing slash
Previously, http endpoints with one trailing slash in the push-endpoint URL, failed to create a topic.
With this fix, topic creation is allowed with or without trailing slash and it creates successfully.
(BZ#2082666)
Blocksize is changed to 4K
Previously, Ceph Object Gateway GC processing would consume excessive time due to the use of a 1K blocksize that would consume the GC queue. This caused
slower processing of large GC queues.
With this fix, blocksize is changed to 4K, which has accelerated the processing of large GC queues.
(BZ#2142167)
Timestamp is sent in the multipart upload bucket notification event to the receiver
Previously, no timestamp was sent on the multipart upload bucket notification event. Due to this, the receiver of the event would not know when the multipart
upload ended.
With this fix, the timestamp when the multipart upload ends is sent in the notification event to the receiver.
(BZ#2149259)
Object size and etag values are no longer sent as 0/empty
Previously, some object metadata would not be decoded before dispatching bucket notifications from the lifecycle. Due to this, object size and etag values were
sent as 0/empty in notifications from lifecycle events.
With this fix, object metadata is fetched and values are now correctly sent with notifications.
(BZ#2153533)
Ceph Object Gateway recovers from kafka broker disconnections
Previously, if the kafka broker was down for more than 30 seconds, there would be no reconnect after the broker was up again. Due to this, bucket notifications
would not be sent, and eventually, after queue fill up, S3 operations that require notifications would be rejected.
With this fix, the broker reconnect happens regardless of the time duration the broker is down and the Ceph Object Gateway is able to recover from kafka broker
disconnects.
(BZ#2184268)
S3 PUT requests with chunked Transfer-Encoding does not require content-length
Previously, S3 clients that PUT objects with Transfer-Encoding:chunked, without providing the x-amz-decoded-content-length field, would fail. As a
result, the S3 PUT requests would fail with 411 Length Required http status code.
With this fix, S3 PUT requests with chunked Transfer-Encoding need not specify a content-length, and S3 clients can perform S3 PUT requests as expected.
(BZ#2186760)
Users can now configure the remote S3 service with the right credentials
Previously, while configuring remote cloud S3 object store service to transition objects, access keys starting with digit were incorrectly parsed. Due to this, there
were chances for the object transition to fail.
With this fix, the keys are parsed correctly. Users cannot configure the remote S3 service with the right credentials for transition.
(BZ#2187394)
Multi-site Ceph Object Gateway
Objects replicated from another zone now returns the header
Previously, in a multi-site configuration, objects that were replicated from another zone, did not return the header x-amz-replication-status=REPLICA.
With this release, in a multi-site configuration, objects that have replicated from another zone, return the header x-amz-replication-status=REPLICA, to
allow multi-site users to identify if the object was replicated locally or not.
(BZ#1467648)
Bucket attributes are no longer overwritten in the archive sync module
Previously, bucket attributes were overwritten in the archive sync module. Due to this, bucket policy or any other attributes would be reset when archive zone
sync_object() was executed.
With this fix, ensure to not reset bucket attributes. Any bucket attribute set on source replicates to the archive zone without being reset.
(BZ#1937618)
Zonegroup is added to the bucket ARN in the notification event
Previously, zonegroup was missing from bucket ARN in the notification event. Due to this, while the notification events handler received events from multiple zone
groups, it would cause confusion in the identification of the source bucket of the event.
With this fix, zonegroup is added to the bucket ARN and the notification events handler receiving events from multiple zone groups has all the required information.
(BZ#2004175)
IBM Storage Ceph 13
bucket read_sync_status() command no longer returns a negative ret value
Previously, bucket read_sync_status() would always return a negative ret value. Due to this, the bucket sync marker command would fail with : ERROR:
sync.read_sync_status()
returned error=0.
With this fix, the actual ret value from the bucket read_sync_status() operation is returned and the bucket sync marker command runs successfully.
(BZ#2127926)
New bucket instance information are stored on the newly created bucket
Previously, in the archive zone, a new bucket would be created when a source bucket was deleted, in order to preserve the archived versions of objects. The new
bucket instance information would be stored in the old instance rendering the new bucket on the archived zone to be in accessible.
With this fix, the bucket instance information is stored in the newly created bucket. Deleted buckets on source are still accessible in the archive zone.
Note: This fix is primarily for newly created buckets and does not work on older buckets.
(BZ#2186774)
Segmentation fault no longer occurs when bucket has a num_shards value of 0
Previously, multi-site sync would result in segmentation faults when a bucket had num_shards value of 0. This resulted in inconsistent sync behavior and
segmentation fault.
With this fix, num_shards=0 is properly represented in data sync and buckets with shard value 0 does not have any issues with syncing.
(BZ#2187617)
RADOS
QoS parameters are restored to built-in default values whenever an attempt is made to modify them
Previously, the modifications to the QoS parameters would not reset to the profile defaults because the OSD would not have permissions to remove the modified
entry. Due to this, the QoS parameters of the built-in were shown as “modified” by the configuration subsystem even though the intended change would not come
into effect. This gave erroneous information to the user.
With this fix, the required permissions for the OSD to remove the modified QoS parameters from the configuration store are enabled. The QoS parameters, such as
reservation, weight, and limit for a built-in profile are restored to the built-in default values whenever an attempt is made to modify them.
(BZ#2124137)
Users can check for accurate number of placement groups with the crush rule
Previously, the check_pg_num() function would not take into account the root OSDs used by the crush rule. This resulted in an inaccurate placement group
number per OSD count.
With this fix, check_pg_num() counts the projected placement group number which are part of the pools affected by the crush rule. The same applies to the
number of OSDs as well; instead of dividing the projected placement group number total by all the osdmap 's OSDs, it is divided only by the OSDs used by the crush
rule.
(BZ#2155766)
Manager continues to send beacons in the event of an error during authentication check
Previously, if an error was encountered when performing an authentication check with a monitor, the manager would get into a state where it would no longer have
an active connection. Due to this, the manager could no longer send beacons and the monitor would mark it as lost.
With this fix, a session (active connection) is reopened in the event of an error and the manager is able to continue to send beacons and is no longer marked as lost.
(BZ#2171847)
MonClient no longer fails to authenticate with EAGAIN
Previously, if MonClient failed to authenticate with EAGAIN, there was a possibility that it would reach a prohibited state in which it would not have an active
connection to ceph-mon and would not try further to acquire one. Due to this, even though the Ceph Manager daemon was technically alive, it became invisible to
monitors in the cluster.
With this fix, the authentication with EAGAIN is handled properly and it works as expected.
(BZ#2187258)
Upon querying the IOPS capacity for an OSD, only the configuration option that matches the underlying device type shows the measured/default value
Previously, the osd_mclock_max_capacity_iops_[ssd|hdd] values were set depending on the OSD’s underlying device type. The configuration options also
had default values that were displayed when queried. For example, if the underlying device type for an OSD was SSD, the default value for the HDD option,
osd_mclock_max_capacity_iops_hdd, was also displayed with a non-zero value. Due to this, displaying values for both HDD and SSD options of an OSD when
queried, caused confusion regarding the correct option to interpret.
With this fix, the IOPS capacity-related configuration option of the OSD that matches the underlying device type is set and the alternate/inactive configuration
option is set to 0. When a user queries the IOPS capacity for an OSD, only the configuration option that matches the underlying device type shows the
measured/default value. The alternative/inactive option is set to 0 to clearly indicate that it is disabled.
(BZ#2111282)
RADOS Block Devices (RBD)
profile rbd-read-only OSD capability allows opening an image in read-only mode
Previously, due to an implementation defect, the profile rbd-read-only OSD capability would disallow opening an image even in read-only mode. Due to this,
a bogus "Operation not permitted" error was returned for images in custom namespaces, although images in a default namespace were not affected.
With this fix, the implementation defect is fixed and the profile rbd-read-only OSD capability allows opening an image in read-only mode regardless of the
images in the namespaces.
14 IBM Storage Ceph
(BZ#2209652)
RBD Mirroring
Detect the blocklisted client
Previously, if the client requested an exclusive lock while blocklisted, the delayed request would not continue and the call that requested the lock would never
complete.
With this fix, the blocklisted client is detected and the stuck condition completes with an appropriate error code.
(BZ#2153673)
Snapshot removal step is moved to the primary site
Previously, a remote cluster could not grab lock quickly enough to remove synced snapshot while I/O was underway due to latency between sites. This would cause
the mirror image snapshot sync process to be stuck during snapshot removal and would not continue with syncing further snapshots.
With this fix, the snapshot removal step is moved to the primary site which grabs the lock fast enough to remove the snapshot and the mirror image snapshot sync
does not get stuck and works as expected.
(BZ#2181055)
Error message when enabling image mirroring within a namespace now provides more insight
Previously, attempting to enable image mirroring within a namespace would fail with a "cannot enable mirroring in current pool mirroring mode" error. The error
would neither provide insight into the problem nor provide any solution.
With this fix, to provide more insight, the error handling is improved and the error now states "cannot enable mirroring: mirroring is not enabled on a namespace".
(BZ#2024444)
Snapshot mirroring no longer halts permanently
Previously, if a primary snapshot creation request was forwarded to rbd-mirror daemon when the rbd-mirror daemon was axed for some practical reason before
marking the snapshot as complete, the primary snapshot would be permanently incomplete. This is because, upon retrying that primary snapshot creation request,
librbd would notice that such a snapshot already existed. It would not check whether this "pre-existing" snapshot was complete or not. Due to this, the mirroring
of snapshots was permanently halted.
With this fix, as part of the next mirror snapshot creation, including being triggered by a scheduler, checks are made to ensure that any incomplete snapshots are
deleted accordingly to resume the mirroring.
(BZ#2120624)
Set the syncing_percent value to 0
Previously, when an image snapshot copy was interrupted and restarted, the rbd-mirror would not have details of the last object copied for that image. The
syncing_percent calculation, which relies on this information, returned an invalid percentage value.
With this fix, if the information about the last object is unavailable, you can set the syncing_percent to 0 until the last object copied is updated.
(BZ#2106421)
Detect the block-listed client
Previously, if the client was block-listed and requested an exclusive lock, the delayed request would not continue and the call that requested the lock would never
complete.
With this fix, the block-listed client is detected and the stuck condition completes with an appropriate error code.
(BZ#2153673)
Known issues
This section documents known issues found in this release of IBM Storage Ceph.
Ceph Object Gateway
Ceph Object Gateway
Bucket lifecycle processing might get delayed if the Ceph Object Gateway instances are killed due to a crash or kill
Presently, if a Ceph Object Gateway instance is killed (ungraceful shutdown) due to a crash or kill -9 while lifecycle processing for one or more buckets is taking
place, processing might not continue on those buckets until two scheduling periods have elapsed, for example, two days. At this point, the buckets are marked stale
and reinitialized. There is no workaround for this issue.
As a workaround, if the user needs to find the version, the daemons' container names include the version.
(BZ#2072680)
Technology previews
This section provides an overview of Technology Preview features introduced or updated in this release of IBM Storage Ceph
IBM Storage Ceph 15
Important: Technology Preview features are not supported with IBM production service level agreements (SLAs), might not be functionally complete, and IBM does not
recommend using them for production. These features provide early access to upcoming product features, enabling customers to test functionality and provide feedback
during the development process.
NFS-CephFS and NFS on Ceph Object Gateway are supported in IBM Storage Ceph 6.1
With this release, NFS-CephFS and NFS on Ceph Object Gateway are supported in IBM Storage Ceph 6.1. For more information, see Managing the NFS-Ganesha
gateway (Technology Preview).
Sources
The updated IBM Storage Ceph source code packages are available at the following location:
For Red Hat Enterprise Linux 8: http://ftp.redhat.com/redhat/linux/enterprise/8Base/en/RHCEPH/SRPMS/
For Red Hat Enterprise Linux 9: http://ftp.redhat.com/redhat/linux/enterprise/9Base/en/RHCEPH/SRPMS/
For IBM Storage Ceph tools: https://public.dhe.ibm.com/ibmdl/export/pub/storage/ceph/6
Asynchronous updates
This section describes the bug fixes, known issues, and enhancements of the z-stream releases.
Release notes for 6.1z7
Release notes for 6.1z6
Release notes for 6.1z5
Release notes for 6.1z4
Release notes for 6.1z3
Release notes for 6.1z2
Release notes for 6.1z7
Enhancements
This section lists all the major updates, enhancements, and new features introduced in this release of IBM Storage Ceph.
Bug fixes
This section describes bugs with significant user impact, which were fixed in this release of IBM Storage Ceph.
Enhancements
This section lists all the major updates, enhancements, and new features introduced in this release of IBM Storage Ceph.
Ceph File System
Ceph Object Gateway
RADOS
Ceph File System
New clone creation no longer slows down due to parallel clone limit
Previously, upon reaching the limit of parallel clones, the rest of the clones would queue up, slowing down the cloning.
With this enhancement, upon reaching the limit of parallel clones at a time, the new clone creation requests are rejected. This feature is enabled by default but can
be disabled.
(BZ#2196829)
The Python librados supports iterating object omap key/values
Previously, the iteration would break whenever a binary/unicode key was encountered.
With this release, the Python librados supports iterating object omap key/values with unicode or binary keys and the iteration continues as expected.
(BZ#2232161)
Ceph Object Gateway
Improved temporary file placement and error messages for the with the /usr/bin/rgw-restore-bucket-index tool
Previously, the /usr/bin/rgw-restore-bucket-index tool only placed temporary files into the /tmp directory. This could lead to issues if the directory ran out
of space, resulting in the following error messing being emitted: ln: failed to access '/tmp/rgwrbi-object-list.XXX': No such file or
directory.
With this enhancement, users can now specify a specific directory to place temporary files, by using the -t command-line option. Additionally, if the specified
directory is full, users now receive an error message specifying the problem: ERROR: the temporary directory's partition is full, preventing
continuation.
16 IBM Storage Ceph
(BZ#2270322)
S3 requests are no longer cut off in the middle of transmission during shutdown
Previously, a few clients faced issues with the S3 request being cut off in the middle of transmission during shutdown without waiting.
With this enhancement, the S3 requests can be configured to wait for the duration defined in the rgw_exit_timeout_secs parameter for all outstanding requests to
complete before exiting the Ceph Object Gateway process unconditionally. Ceph Object Gateway will wait for up to 120 seconds (configurable) for all on-going S3
requests to complete before exiting unconditionally. During this time, new S3 requests will not be accepted. This configuration is off by default.
Note: In containerized deployments, an additional extra_container_args parameter configuration of --stop-timeout=120 (or the value of rgw_exit_timeout_secs
parameter, if not default) is also necessary.
(BZ#2298708)
RADOS
New mon_cluster_log_level command option to control the cluster log level verbosity for external entities
Previously, debug verbosity logs were sent to all external logging systems regardless of their level settings. As a result, the /var/ filesystem would rapidly fill up.
With this enhancement, mon_cluster_log_file_level and mon_cluster_log_to_syslog_level command options have been removed. From this release, use only the
new generic mon_cluster_log_level command option to control the cluster log level verbosity for the cluster log file and all external entities.
(BZ#2053021)
Bug fixes
This section describes bugs with significant user impact, which were fixed in this release of IBM Storage Ceph.
Ceph File System
Ceph Block Device (RBD)
Ceph Block Device (RBD) Mirroring
Ceph Dashboard
RADOS
Ceph Object Gateway
Multi-site Ceph Object Gateway
Ceph Metrics
Ceph Manager plugins
Ceph File System
Missing fields added to the cephfs-top--dump output
Previously, the date, client_count, and filters fields were missing from the cephfs-top --dump command output.
With this release, the missing fields date, client_count, and filters are added to the cephfs-top--dump command output.
(BZ#2184991)
FSPerfStats:re_register_queries() is no longer called unexpectedly when creating new file systems
Previously, whenever the stats plugin was enabled and a new File System was created, the fsmap observers were notified. As a result,
FSPerfStats:re_register_queries() would be unexpectedly called, causing errors in mgr logs.
With this fix, self.mx_last_updated in FSPerfStats::init is initialized, and errors no longer occur in the mgr logs.
(BZ#2280743)
Checks are now performed on m_instance_watcher and m_mirror_watcher with a new patch
Previously, FSMirror::is_failed() function used null m_instance_watcher and m_mirror_watcher pointers to call the member function
InstanceWatcher::is_failed() and MirrorWatcher::is_failed(), respectively.
With this fix, the patch checks m_instance_watcher and m_mirror_watcher before usage.
(BZ#2280636)
The target mon_host details are removed from the peer list and mirror daemon status
Previously, the snapshot mirror peer-list command showed more information than just the peer list. This output caused confusion if there should be only one MON
IP address or all the MON host IP addresses should be displayed.
With this fix, mon_host is removed from the fs snapshot mirror peer_list command and the target mon_host details are removed from the peer list and mirror
daemon status.
(BZ#2277144)
Daemon commands now recognize Ceph implemented flags
Previously, daemon commands were run and implemented while ignoring Ceph flags. Ignored flags included --help and --status flags.
As a result, commands being run that were meant to only issue help or status information. For example, running the ceph --admin-daemon $socketfile client evict
–help command would stop all active connections, instead of running the --help flag on the command.
With this release, client connections are not stopped when running commands with --help and --status flags and a valid error message is emitted.
(BZ#2160855)
IBM Storage Ceph 17
MDS no longer fails with the assert function during recovery
Previously, some metadata was marked as corrupted because the state tests for the file system snap systems were incorrect. Due to this, the MDS wrongly
indicated some metadata is damaged.
With this fix, the state tests for the file system snap systems have been corrected and the MDS no longer indicates the correct metadata as corrupt.
(BZ#2272983)
The erroneous patch from the kernel driver is processed appropriately and MDS no longer enters an infinite loop
Previously, an erroneous patch to the kernel driver caused the MDS to enter an infinite loop processing an operation due to which MDS would become largely
unavailable.
With this fix, the erroneous message from the kernel driver is processed appropriately and the MDS no longer enters an infinite loop.
(BZ#2294415)
The client process no longer causes loss of access to Ceph File System due to incorrect lock API usage
Previously, an incorrect lock API usage caused the client process to crash, causing loss of access to the Ceph File System.
With this fix, the correct lock API is being used and the Ceph File System works as expected.
(BZ#2249038)
MDS_CLIENT_OLDEST_TID clients now does not fail to advance oldest client/flush tid
Previously, the MDS_CLIENT_OLDEST_TID client would fail to advance the oldest client/flush tid. Due to this the MDS would pile up the completed request list in
large size resulting in the MDS going read-only.
With this fix, the oldest_client_tid is updated through the session renew caps message to make sure that MDS would not pile up the completed request and
now MDS_CLIENT_OLDEST_TID does not fail to advance the oldest client/flush tid.
(BZ#2257730)
Ceph Block Device (RBD)
rbd du command no longer crashes if a 0-sized block device image is encountered
Previously, due to an implementation defect, rbd du command would crash if it encountered a zero-sized Ceph Block Device image.
With this fix, the implementation defect is fixed and the rbd du command no longer crashes if a 0-sized Ceph Block Device image is encountered.
(BZ#2293935)
rbd_diff_iterate2() API returns correct results for block device images with LUKS encryption loaded
Previously, due to an implementation defect, rbd_diff_iterate2() API returned incorrect results for Ceph Block Device images with LUKS encryption loaded.
With this fix, rbd_diff_iterate2() API returns correct results for Ceph Block Device images with LUKS encryption loaded.
(BZ#2295023)
Ceph Block Device (RBD) Mirroring
rbd-mirror daemon now properly disposes of outdated PoolReplayer instances
Previously, due to an implementation defect, rbd-mirror daemon did not properly dispose of outdated PoolReplayer instances in particular when refreshing
the mirror peer configuration.
Due to this there was unnecessary resource consumption and number of PoolReplayer instances competed against each other causing rbd-mirror daemon
health to be reported as ERROR and replication would hang in some cases. To resume replication the administrator had to restart the rbd-mirror daemon.
With this fix, the implementation defect is corrected and rbd-mirror daemon now properly disposes of outdated PoolReplayer instances.
(BZ#2279531)
Ceph Dashboard
Grafana dashboard now shows RGW Sync Overview page graphs
Previously, due to regex for the Ceph Object Gateway sync metrics in ceph-exporter, the Grafana dashboards did not display the RGW Sync Overview page graphs.
In place of the graphs, the No Data message was displayed.
With this fix, the ceph-exporter provides the desired metrics value with the expected labels and the RGW sync overview page shows all data correctly.
(BZ#2270223)
RADOS
MGR module no longer takes up one CPU core every minute and CPU usage is normal
Previously, expensive calls from the placement group auto-scaler module to get OSDMap from the Monitor resulted in the MGR module taking up one CPU core
every minute. Due to this, the CPU usage was high in the MGR daemon.
With this fix, the number of OSD map calls made from the placement group auto-scaler module is reduced. The CPU usage is now normal.
(BZ#2241037)
osd pool rmsnap command is now working properly
Previously, the osd pool rmsnap command was faulty and as a result removing snaps using this command would leave clone objects behind.
18 IBM Storage Ceph
With this fix, the osd pool rmsnap command is corrected and a new command force-remove-snap is introduced to remove the leaked clone objects if the faulty
command was used.
(BZ#2272362)
Conflicting Ceph Monitor weights no longer cause Ceph network entity std::discrete_distribution assertion failures
Previously, Ceph network entities, such as Ceph MGR, would run into an assertion failure in std::discrete_distribution due to different Ceph Monitors
having 0 and non-zero weights. As a result, the affected Ceph network fails.
With this fix, Ceph Monitors with 0 weight are ignored and the Ceph network entities no longer receive the std::discrete_distribution assertion failure.
(BZ#2269542)
Ceph Monitor now completes shutdown, as expected
Previously, the Ceph Monitor would stop during a Ceph Monitor shutdown. This occurred due to older versions of RocksDB treating the ETIMEDOUT and EBUSY
statuses differently when freeing resources.
With this fix, the ETIMEDOUT and EBUSY statuses are unified, and the Ceph Monitor does not stop unexpectedly during a shutdown process.
(BZ#2294221)
Autoscaler no longer runs while the norecover flag is set
Previously, the autoscaler would run while the norecover flag was set leading to creation of new placement groups (PGs) and these PGs requiring to be backfilled.
Running of autoscaler while the norecover flag is set allowed in cases where I/O is blocked on missing or degraded objects to avoid client I/O hanging indefinitely.
With this fix, the autoscaler does not run while the norecover flag is set.
(BZ#2293908)
Ceph Object Gateway
Multipart uploads now complete successfully even with failed multipart object parts with the _multipart prefix
Previously, failed multipart object parts with the _multipart prefix were not automatically deleted after a multipart upload. The upload did not complete
successfully due to the failed object parts not being deleted, even with multiple retries. As a result, failed multipart upload parts were still present and listed under
a bucket list, leading to their size being considered in the bucket statistics.
With this fix, the failed parts are automatically deleted after a multipart upload, and the upload completes successfully.
(BZ#2266680)
Batch object deleting is now allowed, with IAM policy permissions
Previously, during a batch delete process, also known as multi object delete, due to the incorrect evaluation of IAM policies returned AccessDenied output if no
explicit or implicit deny were present. The AccessDenied occurred even if there were Allow privileges. As a result, batch deleting fails with the AccessDenied
error.
With this fix, the policies are evaluated as expected and batch deleting succeeds, when IAM policies are enabled.
(BZ#2298712)
Multi-site Ceph Object Gateway
Running copy_object no longer fails to sync the copied object to the original source zone
Previously, if a copy_object was run on a replicated object, it would fail to sync the copied object to the original source zone where the source object was first
written and replicated from. This is because copy_object retains the source attributes by default.
As a result, whenever get_obj() was issued from a fetch_remote_obj() call during sync, a check for RGW_ATTR_OBJ_REPLICATION_TRACE would be made
and if that destination zone was already present in the trace, a NOT_MODIFIED error was returned, thus failing to replicate the copied object.
With this fix, the conflicting source attributes are removed during replication, thus resolving the issue.
(BZ#2297360)
The authentication of the forwarded request on the primary site no longer fails
Previously, an S3 request issued to secondary failed if temporary credentials returned by STS were used to sign the request. The failure occurred because the
request would be forwarded to the primary and signed using a system user's credentials which do not match the temporary credentials in the session token of the
forwarded request. Due to unmatched credentials, the authentication of the forwarded request on the primary site fails, resulting in the failure of the S3 operation.
With this fix, the authentication is by-passed by using temporary credentials in the session token in case a request is forwarded from secondary to primary. The
system user's credentials are used to complete the authentication successfully.
(BZ#2271595)
Ceph Metrics
ceph-exporter is now able to handle exceptions during metric collection and maintains stable operation
Previously, the ceph-exporter thread in ceph-client.admin was not properly handling exceptions, particularly when processing JSON data related to cluster
metrics. As a result, the ceph-exporter crashed with an invalid_argument error while collecting metrics as the cluster approached its near-full ratio. The error
caused the Ceph health status to enter a warning state.
With this fix, exception handling is being handled properly in the ceph-exporter thread. In addition, error logging was added to capture details of any exceptions.
The ceph-exporter now gracefully handles exceptions during metric collection, preventing crashes and maintaining stable operation even when the cluster
approaches capacity limits.
(BZ#2266538)
IBM Storage Ceph 19
Ceph Manager plugins
Prometheus manager module no longer crashes during startup
Previously, a race condition between loading Prometheus and the Orchestrator module caused the Prometheus module to crash during startup. Due to this, the
cluster metrics being produced by the Prometheus module would not be displayed in several monitoring tools. For example, metrics displayed in the Ceph
Dashboard. A Ceph health error would be thrown at the user due to the error while loading the Prometheus Manager module.
With this fix, the initialization of the modify_instance_id is removed from the startup and moved to the get_metadata_and_osd_status which is called periodically by
the metrics collection thread or when some clients performs a get on/metrics endpoint. Loading errors no longer occur in the Prometheus Manager module.
(BZ#2303151)
Release notes for 6.1z6
An update is now available for IBM Storage Ceph 6.1 which provides security updates.
Release notes for 6.1z5
Enhancements
This section lists all the major updates, enhancements, and new features introduced in this release of IBM Storage Ceph.
Bug fixes
This section describes bugs with significant user impact, which were fixed in this release of IBM Storage Ceph.
Enhancements
This section lists all the major updates, enhancements, and new features introduced in this release of IBM Storage Ceph.
The Ceph Ansible utility
RBD Mirroring
Ceph File System
The Ceph Ansible utility
All bootstrap CLI parameters are now made available for usage in the cephadm-ansible module
Previously, only a subset of the bootstrap CLI parameters were available and it was limiting the module usage.
With this enhancement, all bootstrap CLI parameters are made available for usage in the cephadm-ansible module.
(BZ#2269514)
RBD Mirroring
RBD diff-iterate now executes locally if exclusive lock is available.
Previously, when diffing against the beginning of time (fromsnapname == NULL) in fast-diff mode (whole_object == true with fast-diff image feature
enabled and valid), RBD diff-iterate wasn’t guaranteed to execute locally.
With this enhancement, rbd_diff_iterate2() API performance improvement is implemented and RBD diff-iterate is now guaranteed to execute locally if
exclusive lock is available. This brings a dramatic performance improvement for QEMU live disk synchronization and backup use cases assuming fast-diff image
feature is enabled.
(BZ#2259053)
Ceph File System
Snapshot scheduling support is now provided for subvolumes
With this enhancement, snapshot scheduling support is provided for subvolumes. All snapshot scheduling commands accept `--subvol` and `--group` arguments
to refer to appropriate subvolumes and subvolume groups. If a subvolume is specified without a subvolume group argument, then the default subvolume group is
considered. Also, a valid path need not be specified when referring to subvolumes and just a placeholder string is sufficient due to the nature of argument parsing
employed.
Example
# ceph fs snap-schedule add - 15m --subvol sv1 --group g1
# ceph fs snap-schedule status - --subvol sv1 --group g1
(BZ#2243783)
Bug fixes
20 IBM Storage Ceph
This section describes bugs with significant user impact, which were fixed in this release of IBM Storage Ceph.
Ceph File System
RADOS
Ceph Object Gateway
Multi-site Ceph Object Gateway
Ceph File System
snap-schedule repeat and retention specification for monthly snapshots is changed from m to M
Previously, the snap-schedule repeat specification and retention specification for monthly snapshots was not consistent with other Ceph components.
With this fix, the specifications are changed from m to M and it is now consistent with other Ceph components. For example, to retain 5 monthly snapshots, you need
to issue the following command:
# ceph fs snap-schedule retention add /some/path M 5
(BZ#2270220)
ceph-mds no longer crashes during inode replication in a multi-mds cluster
Previously, due to incorrect lock assertion in ceph-mds, ceph-mds would crash when some inodes were replicated in a multi-mds cluster.
With this fix, the lock state in the assertion is validated and no crash is observed.
(BZ#2265505)
RADOS
The OSD daemon no longer asserts due to allocation failure
Previously, the OSD daemon would assert if the allocator gets 4K requests when it is configured with a different allocation unit. Due to this, the OSD daemon would
fail to boot, thereby creating a data unavailability or data loss situation.
With this fix, the LBA alignment requirement in allocators is eliminated. The OSD daemon no longer asserts due to allocation failure.
(BZ#2264053)
Ceph Object Gateway
Multipart object is no longer damaged during multipart upload
Previously, during a multipart upload, when restarting an upload of a failed part, the cleanup of the first attempt sometimes removed tail objects from the
secondary attempt. A restart of a failed part upload could occur, for example, due to a time-out. As a result of this failure, the Ceph Object Gateway multipart object
would respond to a HEAD request but fail during a GET request, as some tail objects were missing.
With this fix, the code now cleans up properly during the first attempt and the multipart object is no longer damaged.
(BZ#2265071)
Ceph Object Gateway now passes requests with well-formed payloads of the new stream encoding forms
Previously, Ceph Object Gateway did not recognize STREAMING-AWS4-HMAC-SHA256-PAYLOAD and STREAMING-UNSIGNED-PAYLOAD-TRAILER encoding
forms, resulting in failed requests.
With this fix, a new logic is introduced. This logic recognizes, parses, and verifies new trailing request signatures provided, when applicable, for new encoding forms.
This allows Ceph Object Gateway to pass requests with well-formed payloads of the new stream.
(BZ#2260354)
Ceph Object Gateway roles are no longer illegally accessed and no crash happens
Previously, a variable representing a Ceph Object Gateway role was being accessed before it was initialized resulting in an illegal access which induced a segfault.
With this fix, operations are reordered and there is no illegal access and roles are enforced as required.
(BZ#2260859)
Multi-site Ceph Object Gateway
CURL path normalization is now disabled at startup
Previously, by default, "path normalization" would be performed by CURL as part of the Ceph Object Gateway replication stack. Due to this, object names were
illegally reformatted during replication and objects whose names contained embedded "." and ".." were not replicated.
With this fix, CURL path normalization is disabled at startup and affected objects are now replicated as expected.
(BZ#2265153)
Release notes for 6.1z4
Enhancements
This section lists all the major updates, enhancements, and new features introduced in this release of IBM Storage Ceph.
Bug fixes
This section describes bugs with significant user impact, which were fixed in this release of IBM Storage Ceph.
IBM Storage Ceph 21
Enhancements
This section lists all the major updates, enhancements, and new features introduced in this release of IBM Storage Ceph.
Ceph File System
Ceph Object Gateway
Ceph File System
MDS dynamic metadata balancer is off by default
With this enhancement, MDS dynamic metadata balancer is off by default to improve the poor balancer behavior that could fragment trees in undesirable or
unintended ways simply by increasing the max_mds file system setting.
Operators must turn on the balancer explicitly to use it.
(BZ#2255435)
The resident segment size perf counter in the MDS is tracked with a higher priority
With this enhancement, the MDS resident segment size (or RSS) perf counter is tracked with a higher priority to allow callers to consume its value to generate
useful warnings which is useful for rook to ascertain the MDS RSS size and act accordingly.
(BZ#2256731)
ceph auth commands give a message when permissions in MDS are incorrect
With this enhancement, permissions in MDS capability now either start with r, rw, * or all. This results in theceph
auth commands like ceph auth add, ceph auth caps, ceph auth get-or-create and ceph auth get-or-create-key to generate a clear message
when the permissions in the MDS caps are incorrect.
(BZ#2218189)
Ceph Object Gateway
The radosgw-admin bucket stats command prints bucket versioning
With this enhancement, the radosgw-admin bucket stats command prints the versioning status for buckets as enabled or off since versioning can be
enabled or disabled after creation.
(BZ#2256364)
Bug fixes
This section describes bugs with significant user impact, which were fixed in this release of IBM Storage Ceph.
Ceph File System
Cephadm utility
Ceph Dashboard
RADOS
Ceph Object Gateway
Ceph File System
MDS triggers only one reintegration for each case
Previously, when unlinking the CInode which had multiple links, the MDS would trigger the same reintegration multiple times resulting in slowing down of normal
requests and thereby MDS performance.
With this fix, only one reintegration is triggered for each case and no redundant reintegration is triggered.
(BZ#2249566)
The fs volume info now correctly shows the number of subvolumes queued for deletion
Previously, the os.listdir utility used for listing the trash directory containing subvolumes for deletion did not have subvolume base directory context. Hence, it
always failed with ENOENT. Due to this the pending_subvolume_deletions field in fs
volume infooutput was always zero.
With this fix, the listdir utility which has the context of the subvolume base directory is used and hence the pending_subvolume_deletions field in fs
volume info correctly represents the remaining subvolumes queued for deletion.
(BZ#2240585)
The loner member is now set to true
Previously, for a file lock in the LOCK_EXCL_XSYN state, the non-loner clients would be issued empty caps. However, since the loner of this state was set to false,
it could make the locker to issue the Fcb caps, which is incorrect. Due to this, some client requests incorrectly revoked some caps and infinitely waited and caused
slow requests.
With this fix, the loner member is set to true and as a result the corresponding request is not blocked.
(BZ#2251767)
Snap-schedule manager module now alerts about missing --fs argument
Previously, the snap-schedule Manager module would not prompt for --fs argument even in multi-fs setup causing the first filesystem to be selected for all
operations. This led to unintentional changes in the system.
22 IBM Storage Ceph
With this fix, in multi-fs setup, the snap-schedule Manager module prompts the user for the --fs argument and refuses to act on the command.
The users now get alerts about the missing --fs argument without which the snap-schedule Manager module could have used the first file system name for the
command as default.
(BZ#2224060)
The MDS no longer crashes when the journal logs are flushed
Previously, when the journal logs were successfully flushed, you could set the lockers’ state to LOCK_SYNC or LOCK_PREXLOCK when the xclock count was nonzero. However, the MDS would not allow that and would crash.
With this fix, MDS allows the lockers’ state to LOCK_SYNC or LOCK_PREXLOCK when the xclock count is non-zero and the MDS does not crash.
(BZ#2248998)
The modules capture all exceptions to stay operational
Previously, the modules would crash when unknown or unexpected exceptions were not captured at the time of development, leading to module crash and loss of
functionality.
With this fix, the module captures all exceptions. The resulting traceback is also dumped to the console or the log file to report unexpected events. As a result, the
module continues to stay operational, providing a better user experience.
(BZ#2227812)
Incremental snapshot sync are now fixed
Previously, after creating a second snapshot, it would not sync incrementally, as prev would always be none. Due to this, prevsnapshot would be none even if it
existed resulting in a performance impact on snapshot synchronization.
With this fix, the prev snapshot has its address and this fixes the incremental snapshot sync, which improves the overall performance of cephfs-mirroring.
(BZ#2248637)
MDS now properly trims the caches
Previously, the standby-replay MDS daemons would not trim their caches causing MDS to run out of memory.
With this fix, MDS properly trims its cache when in standby-replay and no longer runs out of memory.
(BZ#2252126)
MDS now proceeds with failover recovery normally
Previously, MDS would not queue the next client request for replay in the up:client-replay state causing the MDS to hang in the up:client-replay state.
With this fix, the next client replay request is queued automatically as part of request cleanup and now MDS proceeds with failover recovery normally.
(BZ#2244866)
The per dump command now works correctly
Previously, the diagnostic counters showing journal replay progress were not updated in the up:replay state causing the perf dump command to not evaluate
progress.
With this fix, the counters are updated during replay and the perf dump command works properly.
(BZ#2259301)
Cephadm utility
Cephadm now maintains list of potential upgrade failure issues correctly
Previously, cephadm missed including the last two items in the list of known upgrade failure reasons due a bug in the code. Due to this, if the upgrade were to fail
for one of those last two reasons, most notably in the case of an offline host, the cephadm
mgr module would crash.
With this fix, the list of potential upgrade failure issues is repaired and the cephadm
mgr module no longer crashes due to upgrade failures.
(BZ#2256453)
Ceph Dashboard
OSD panel now shows correct panel colors when OSDs are down or out
Previously, there was no threshold set in the Ceph cluster dashboard's OSD panel to show OSDs which are out or down in a danger state.
With this fix, the thresholds to the OSD panel are added now the OSD panel shows correct panel colors when OSD's are down or out.
(BZ#2240132)
RADOS
The ceph config dump command with pretty-print option now gives correct output
Previously, the ceph config dump command without the pretty-print formatted output showed the localized option name and its value. Due to this, the ceph
config dump command result with and without the pretty-print option was inconsistent.
With this fix, the ceph config dump command now shows localized option name with and without the pretty-print option and the output is now always
consistent.
ceph config dump --format <type>
IBM Storage Ceph 23
command where type is the pretty-print type.
An example of a normalized vs localized option is as below::
Normalized: mgr/dashboard/ssl_server_port
Localized: mgr/dashboard/x/ssl_server_port
(BZ#2250161)
Slow requests are not reported for the down and out OSD daemon
Previously, the cluster would report slow requests on OSD which was already out resulting in an inconsistent behavior as the OSD which is already out should not
have any operations registered for it.
With this fix, the daemon_state records in the mgr daemon are longer stored for the down and out OSD daemon. As a result, slow requests are not reported now
for the down and out OSD daemon.
(BZ#2189920)
Disjointed blobs are now handled properly and without assert
Previously, the Elastic Shared Blob feature mishandled cloning blobs that consist of disjointed allocation units. Due to this the Immediate assert()
function corrupted data which was not persisted to disk.
With this fix, now disjointed blobs are handled properly and without assert.
(BZ#2254497)
Database blocklisting is no longer a terminal failure for a module
Previously, libcephsqlite would blocklist active connections when RADOS access was lost causing some ceph-mgr modules, like devicehealth to become
unavailable.
With this fix, reopen database connections when blocklist errors occur and now database blocklisting is no longer a terminal failure for a module.
(BZ#2248719)
rados cppool now asks to confirm the operation with --yes-i-really-mean-it for pools with self managed snapshots
Previously, rados cppool did not ask for confirmation for operations with pools having self managed snapshots unlike the other Ceph commands.
The rados cppool does not preserve self managed snapshots when copying a pool and it should be used with care and awareness on pools with self managed
snapshots like the RBD pools. Also, the obligatory switch is not enforced for RBD pools.
With this fix, rados cpppool ceases the operation and exits with a warning even if the user misses the switch and asks to confirm the operation with --yes-ireally-mean-it for pools with self managed snapshots.
(BZ#2252780)
Ceph Object Gateway
The modification time for the bucket is now displayed correctly
Previously, the output for radosgw-admin bucket stats ... command would not load the value of bucket modification time (mtime) correctly, resulting in
displaying the output for mtime as 0.000000.
With this fix, the information loads and the mtime for the bucket is displayed correctly.
(BZ#2257805)
Custom work times can now be scheduled
Previously, the lifecycle should_work() predicate would not take the next date notion into account. As a result, any custom work time XY:TW-AB:CD would
break the LC processing when AB < XY. Due to this, lifecycles with custom work times caused lifecycle processing to never be run.
With this fix, the should_work() function now considers next date and the custom work times.
(BZ#2255939)
Bucket stat now generates correct output
Previously, due to a code change, the radosgw-admin bucket check stat calculation and bucket reshard stat calculation would generate incorrect values
when there were objects that had transitioned from unversioned to versioned. Due to this, the stat calculations would be incorrect for some buckets.
With this fix, the stat calculations are correct and incorrect bucket stat outputs are no longer generated.
(BZ#2257373)
Release notes for 6.1z3
Enhancements
This section lists all the major updates, enhancements, and new features introduced in this release of IBM Storage Ceph.
Bug fixes
This section describes bugs with significant user impact, which were fixed in this release of IBM Storage Ceph.
Enhancements
This section lists all the major updates, enhancements, and new features introduced in this release of IBM Storage Ceph.
Ceph File System
Ceph Object Gateway
24 IBM Storage Ceph
NFS Ganesha
RADOS
Ceph File System
Laggy clients are now evicted only if there are no laggy OSDs
Previously, monitoring performance dumps from the MDS would sometimes show that the OSDs were laggy, objecter.op_laggy and objecter.osd_laggy,
causing laggy clients (cannot flush dirty data for cap revokes).
With this enhancement, if defer_client_eviction_on_laggy_osds is set to true and a client gets laggy because of a laggy OSD then client eviction will not
take place until OSDs are no longer laggy.
(BZ#2228065)
The snap schedule module now supports a new retention specification
With this enhancement, users can define a new retention specification to retain the number of snapshots.
For example, if a user defines to retain 50 snapshots irrespective of the snapshot creation cadence, the snapshots retained is 1 less than the maximum specified as
the pruning happens after a new snapshot is created. In this case, 49 snapshots are retained so that there is a margin of 1 snapshot to be created on the file system
on the next iteration to avoid breaching the system configured limit of mds_max_snaps_per_dir.
Note: Configure the mds_max_snaps_per_dir and snapshot scheduling carefully to avoid unintentional deactivation of snapshot schedules due to file system
returning a Too many links error if the mds_max_snaps_per_dir is breached.
(BZ#2227807)
Ceph Object Gateway
rgw-restore-bucket-index tool can now restore the bucket indices for versioned buckets
With this enhancement, to enable the rgw-restore-bucket-index tool to work as broadly as possible, the ability to restore the bucket indices for versioned
buckets, in addition to its existing ability working with un-versioned buckets, is added.
(BZ#2182385)
NFS Ganesha
NFS Ganesha version updated to V5.6
With this enhancement of updated version of NFS Ganesha, following issues have been fixed:
FSAL's state_free function called by free_state did not actually free.
CEPH: Fixed the cmount_path.
CEPH: Currently client_oc true was broken, forced it to false.
(BZ#2249958)
RADOS
New reports available for sub-events for delayed operations.
Previously, slow operations were marked as delayed but without a detailed description.
With this enhancement, you can view the detailed descriptions of delayed sub-events for operations.
(BZ#2240838)
Turning the noautoscale flag on/off now retains each pool's original autoscale mode configuration.
Previously, the pg_autoscaler did not persist in each pool's autoscale
mode configuration when the noautoscale flag was set. Due to this, after turning the noautoscale flag on/off, the user would have to go back and set the
autoscale mode for each pool again.
With this enhancement, the pg_autoscaler module persists individual pool configuration for the autoscale mode after thenoautoscale flag is set.
(BZ#2241201)
BlueStore instance cannot be opened twice
Previously, when using containers, it was possible to create unrelated inodes that targeted the same block device mknod b causing multiple containers to think that
they have exclusive access.
With this enhancement, the reinforced advisory locking with O_EXCL open flag dedicated for block devices is now implemented,thereby improving the protection
against running OSD twice at the same time on one block device.
(BZ#2239449)
Bug fixes
This section describes bugs with significant user impact, which were fixed in this release of IBM Storage Ceph.
Ceph-ansible
Ceph File System
Ceph Object Gateway
RADOS
IBM Storage Ceph 25
RBD Mirroring
Ceph-ansible
The Ceph packages are now installed without stopping any of the running Ceph services
Previously, during the upgrade, all the Ceph services stopped running as the Ceph 4 packages got uninstalled instead of updating.
With this fix, the new Ceph 5 packages get installed during upgrades and do not impact the running Ceph processes.
(BZ#2211324)
Repackaged the correct version of cephadm-ansible matching Ceph 6
Previously, the wrong version of cephadm-ansible was released with Ceph 6. The released version was for Ceph 7 and hence when running the preflight playbook it
tried to connect to Ceph 7 repositories which were not available.
With this fix, the correct version of cephadm-ansible matching Ceph 6 is repackaged and now the preflight playbook connects to the correct repositories.
(BZ#2244978)
Ceph File System
Client always sends a caps revocation acknowledgment to the MDS daemon
Previously, whenever an MDS daemon sent a caps revocation request to a client and during this time, if the client released the caps and removed the inode, then the
client would drop the request directly, but the MDS daemon would need to wait for a caps revoking acknowledgment from the client. Due to this, even when there
was no need for caps revocation, the MDS daemon would continue waiting for an acknowledgment from the client, causing a warning in MDS daemon health status.
With this fix, the client always sends a caps revocation acknowledgment to the MDS daemon, even when there is no inode existing and the MDS daemon no longer
stays stuck.
(BZ#2227999)
User-space Ceph File System (CephFS) works as expected post upgrade
Previously, the user-space CephFS client would sometimes crash during a cluster upgrade. This would occur due to stale feature bits on the MDS side that were
held on the user-space side.
With this fix, ensure that the user-space CephFS client has updated MDS feature bits that allow the clients to work as expected after a cluster upgrade.
(BZ#2249814)
Blocklist and evict client for large session metadata
Previously, large client metadata buildup in the MDS would sometimes cause the MDS to switch to read-only mode.
With this fix, the client that is causing the buildup is blocklisted and evicted, allowing the MDS to work as expected.
(BZ#2238666)
Ceph Object Gateway
Testing for reshardable bucket layouts is added to prevent crashes
Previously, with the added bucket layout code to enable dynamic bucket resharding with multisite, there was no check to verify if the bucket layout supported
resharding during dynamic, immediate, or rescheduled resharding. Due to this, the Ceph Object gateway daemon would crash in case of dynamic bucket resharding
and the `radosgw-admin` command would crash in case of immediate or scheduled resharding.
With this fix, a test for reshardable bucket layouts is added and the crashes no longer occur. When immediate and scheduled resharding occurs, an error message is
displayed. When dynamic bucket resharding occurs, the bucket is skipped.
(BZ#2245147)
The RetainUntilDate encoding is now fixed for new PutObjectRetention requests
Previously, for S3 Object Lock with PutObjectRetention requests that specify a RetainUntilDate after the year 2106 was truncated to 32 bits when stored,
so a much earlier date was used for object lock enforcement.
With this fix, the RetainUntilDate encoding is resolved for new PutObjectRetention requests.
Note: This fix cannot repair the dates of existing object locks. Such objects can be identified with a HeadObject request based on the x-amz-object-lockretain-until-dateresponse header.
(BZ#2252337)
The S3 bucket now shows consistent stats
Previously, a race condition in S3 CompleteMultipartUpload, where multiple updates complete around the same time, caused cleanup of the uploaded part
objects to be skipped. Due to this, the S3 bucket stats showed an unexpected value for numObjects
With this fix, the race condition is resolved and now the S3 bucket now shows consistent stats.
(BZ#2181424)
Theuser modify –placement-id command can now be used with an empty --storage-class argument
Previously, if the --storage-class argument was not used when running the 'user modify --placement-id' command, the command would fail.
With this fix, the --storage-class argument can be left empty without causing the command to fail.
(BZ#2245697)
The segfault case is now resolved
26 IBM Storage Ceph
Previously, a segfault occurred depending on the timing due to a race condition that affected async notification endpoint completion handlers.
With this fix, the race condition is now resolved and now a segfault does not occur.
(BZ#2252256)
Multipart uploaded objects can not bypass governance
Previously, the requested object-lock governance information provided with the new multipart uploads was not persisted with the uploaded part information,
permitting the governance to be bypassed after the upload completed. Due to this, there was a possibility of a violation of data retention policy for multipart
uploaded objects.
With this fix, object lock parameters are persisted with part upload information and checked at completion, such that there is no window to bypass governance.
(BZ#2252792)
RADOS
Ceph reports a POOL_APP_NOT_ENABLED warning if the pool has 0 objects stored in it
Previously, Ceph status failed to report pool application warning if the pool was empty resulting in RGW bucket creation failure if the application tag was enabled for
RGW pools.
With this fix, Ceph reports a POOL_APP_NOT_ENABLED warning even if the pool has 0 objects stored in it.
(BZ#2232663)
Spillover now appears properly
Previously, a refactoring removed the spillover detection code resulting in the spillover from dedicated Block.DB to main Block device never being detected.
With this fix, the removed code is added back and now spillover is properly detected as before.
(BZ#2237881)
Monitors now do not get stuck in elections during crash/shutdown tests
Previously, the disallowed_leaders attribute of the MonitorMap was conditionally filled only when the monitors entered stretch_mode causing the monitors
that just for revived to be unable to enter stretch_mode right away because of being in a probing state. This led to a mismatch in thedisallowed_leaders set
between the monitors across the cluster. Due to this, the monitors would fail to elect a leader and the election would get stuck resulting in Ceph becoming
unresponsive.
With this fix, monitors now do not have to be in stretch_mode to fill thedisallowed_leaders attribute and hence, monitors do not get stuck in elections during
crash/shutdown tests.
(BZ#2243741)
All OSDs hosted by the machine get notified whenever the auto-tuner applies a new osd_memory_target
Previously, when enabling the osd_memory_target_autotuneoption, the memory target was applied at the host-level. But the code that applies the memory
target would not determine the correct crush location of the parent host for the change to be propagated to the OSD(s) of the host. Hence, none of the OSDs hosted
by the machine would get notified via the config observer. This resulted in theosd_memory_target to remain unchanged for those sets of OSDs.
With this fix, the correct crush location of the OSDs parent (host) is determined based on the host mask. Hence, all the OSDs hosted by the machine get notified
whenever the auto-tuner applies a new osd_memory_target and the change is reflected accordingly.
(BZ#2213873)
RBD Mirroring
rbd_support module no longer fails to recover from repeated blocklisting of its client
Previously, it was observed that the rbd_support module failed to recover from repeated blocklisting of its client due to a recursive deadlock in the rbd_support
module, a race condition in the rbd_support module's librbd client, and a bug in the librbd cython bindings that sometimes crashed the ceph-mgr.
With this release, all these 3 issues are fixed and rbd_support module no longer fails to recover from repeated blocklisting of its client.
(BZ#2247543)
Release notes for 6.1z2
Enhancements
This section lists all the major updates, enhancements, and new features introduced in this release of IBM Storage Ceph.
Bug fixes
This section describes bugs with significant user impact, which were fixed in this release of IBM Storage Ceph.
Enhancements
This section lists all the major updates, enhancements, and new features introduced in this release of IBM Storage Ceph.
Ceph Object Gateway
Multi-site Ceph Object gateway
Ceph Dashboard
IBM Storage Ceph 27
Ceph Object Gateway
rgw-restore-bucket-index tool can now restore the bucket indices for versioned buckets
With this enhancement, to enable the rgw-restore-bucket-index tool to work as broadly as possible, the ability to restore the bucket indices for versioned
buckets, in addition to its existing ability working with unversioned buckets, is added.
Realm, zone group, and/or zone can be specified when running the rgw-restore-bucket-index command
Previously, the tool could only work with the default realm, zone group, and zone.
With this enhancement, realm, zone group, and/or zone can be specified when running the rgw-restore-bucket-index command. Three additional commandline options are added:
"-r <realm>"
"-g <zone group>"
"-z <zone>"
(BZ#2183926)
Additional features and enhancements are added to rgw-gap-list and
rgw-orphan-list scripts to enhance end-users' experience
With this enhancement, to improve the end-users' experience with rgw-gap-list and rgw-orphan-list scripts, a number of features and enhancements have
been added, including internal checks, more command-line options, and enhanced output.
(BZ#2228242)
Multi-site Ceph Object gateway
Original multipart uploads can now be identified in multi-site configurations
Previously, a data corruption bug was fixed in 6.1z1 that effected multi-part uploads with server-side encryption in multi-site configurations.
With this enhancement, a new tool, radosgw-admin bucket resync encrypted
multipart, can be used to identify these original multipart uploads. The LastModified timestamp of any identified object is incremented by 1ns to cause peer
zones to replicate it again. For multi-site deployments that make any use of Server-Side encryption, users are recommended to run this command against every
bucket in every zone after all zones have upgraded.
(BZ#2227842)
Ceph Dashboard
Dashboard host loading speed is improved and pages now load faster
Previously, large clusters of five or more hosts had a linear increase in load time on hosts page and main page.
With this enhancement, the dashboard host loading speed is improved and pages now load faster.
(BZ#2220922)
Bug fixes
This section describes bugs with significant user impact, which were fixed in this release of IBM Storage Ceph.
Ceph Dashboard
Ceph File System
RBD Mirroring
RADOS
Ceph File System
Errors are handled gracefully in MDLog::_recovery_thread
Previously, a write would fail if the MDS was already block-listed due to the fs
fail issued by the QA tests. For instance, the QA test test_rebuild_moved_file (tasks/data-scan) would fail due to this reason.
With this fix, the write failures are gracefully handled in MDLog::_recovery_thread.
(BZ#2228357)
ceph mds metadata command now functions as expected across upgrades
Previously, monitors could lose track of MDS metadata during upgrades and cancel the PAXOS transactions causing the MDS metadata to be unavailable.
With this fix, MDS metadata is added in batches with FSMap changes to ensure consistency. The ceph mds metadata command now functions as expected
across upgrades.
(BZ#2236385)
MDS locks are obtained in the correct order
Previously, MDS would acquire metadata tree locks in the wrong order, resulting in a create and getattr RPC request to deadlock.
With this fix, locks are obtained in the correct order in MDS and the requests no longer deadlock.
(BZ#2236188)
28 IBM Storage Ceph
Sending split_realms information is skipped from CephFS MDS
Previously, the split_realms information would be incorrectly sent from the CephFS MDS which could not be correctly decoded by kclient. Due to this, the
clients would not care about the split_realms and treat it as a corrupted snaptrace.
With this fix, split_realms are not sent to kclient and no crashes take place.
(BZ#2228004)
Revocation requests no longer get stuck
Previously, before the revoke request was sent out, which would increase the seq, if the clients released the corresponding caps and sent out the cap update
request with the old seq, the MDS would miss checking the seq(s) and cap calculation. Due to this, the revocation requests would be stuck infinitely and would
throw warnings about the revocation requests not responding from clients.
With this fix, an acknowledgement is always sent for revocation requests and they no longer get stuck.
(BZ#2224243)
Snapshot data is no longer lost after setting writing flags
Previously, in clients, if the writing flag was set to 1 when the Fb caps were used, it would be skipped in case of any dirty caps and reuse the existing capsnap,
which is incorrect. Due to this, two consecutive snapshots would be overwritten and lose data.
With this fix, the writing flags are correctly set and no snapshot data is lost.
(BZ#2224239)
Thread renaming no longer fails
Previously, in a few rare cases, during renaming, if another thread tried to lookup the dst dentry, there were chances for it to get inconsistent result, wherein both
the src dentry and dst dentry would link to the same inode simultaneously. Due to this,the rename request would fail as two different dentries were being linked to
the same inode.
With this fix, the thread waits for the renaming action to finish and everything works as expected.
(BZ#2224230)
RBD Mirroring
Non-primary images are now deleted when the primary image is deleted
Previously, a race condition in the rbd-mirror daemon image replayer prevented a non-primary image from being deleted when the primary was deleted. Due to this,
the non-primary image would not be deleted and the storage space was used.
With this fix, the rbd-mirror image replayer is modified to eliminate the race condition. Non-primary images are now deleted when the primary image is deleted.
(BZ#2215392)
Demoted mirror snapshot is removed following the promotion of the image
Previously, due to an implementation defect, the demoted mirror snapshots would not be removed following the promotion of the image, whether on the secondary
image or on the primary image. Due to this, demoted mirror snapshots would pile up and consume storage space.
With this fix, the implementation defect is fixed and the appropriate demoted mirror snapshot is removed following the promotion of the image.
(BZ#2214278)
The librbd client correctly propagates the block-listing error to the caller
Previously, when the rbd_support module's RADOS client was block-listed, the module's mirror_snapshot_schedule handler would not always shut down
correctly. The handler's librbd client would not propagate the block-list error, thereby stalling the handler's shutdown. This lead to the failures of the
mirror_snapshot_schedule handler and the rbd_support module to automatically recover from repeated client block-listing. The rbd_support module
stopped scheduling mirror snapshots after its client was repeatedly block-listed.
With this fix, the race in the librbd client between its exclusive lock acquisition and handling of block-listing is fixed. This allows the librbd client to propagate
the block-listing error correctly to the caller, for example, the mirror_snapshot_schedule handler, while waiting to acquire an exclusive lock. The
mirror_snapshot_schedule handler and the rbd_support_module automatically recover from repeated client block-listing.
(BZ#2211290)
The mirror_snapshot_schedule handler and the rbd_support module automatically recovers from repeated client blocklisting
Previously, when the rbd_support module’s RADOS client was blocklisted, the module’s mirror_snapshot_schedule would not always shut down correctly.
The handler's librbd client would not propagate the blocklist error stalling the handler's shutdown. This lead to the failures of the mirror_snapshot_schedule
handler and the rbd_support module to automatically recover from repeated client blocklisting. The rbd_support module stopped scheduling mirror snapshots
after its client was repeatedly blocklisted.
With this release, the race in the librbd client between its exclusive lock acquisition and handling of blocklisting is fixed. This allows the librbd client to propagate
the blocklisting error correctly to the caller, for example, the mirror_snapshot_schedule handler, while waiting to acquire an exclusive lock. The
mirror_snapshot_schedule handler and the rbd_support module automatically recovers from repeated client blocklisting.
(BZ#2211290)
Ceph Dashboard
Size and provisioned column names are changed in RBD images table to avoid mismatch
Previously, size and provisioned columns in RBD images table were labeled incorrectly, causing the column's name to mismatch with the data displayed.
With this fix, the columns' names are changed to Used and Total Used. Column Used displays the disk usage for the image and Total Used displays the total disk
usage.
IBM Storage Ceph 29
(BZ#2210944)
Dashboard now relies on the rgw_dns_name configuration for Object Gateway requests
Previously, the dashboard relied on the hostname instead of IP address to do the Object Gateway requests. Due to this, when the hostname was unresolvable, the
Object Gateway pages would break.
With this fix, the dashboard relies on the rgw_dns_name configuration which the user can set manually. If set, the value of rgw_dns_name is used. Otherwise, it
resolves to the IP address. Errors like SignatureDoesNotMatch and hostname
not resolvable in container environments are now resolved.
(BZ#2229179)
ceph-exporter daemons no longer crash during upgrade from 6.1z1 to 6.1z2
Previously, there was a format change for the output of counter dump and counter schema commands, which was delivered to 6.1z2. Due to this, during the
upgrade from 6.1z1 to 6.1z2, some exporter daemons crashed which were still in queue to be upgraded and using the old format.
With this fix, looping through the output of counter dump and counter
schema is avoided if the the new format is unsupported. ceph-exporter daemons no longer crash.
(BZ#2233762)
JSON entries emitted by counter dump and counter schema no longer repeat, can be parsed, and are valid
Previously, the JSON emitted by the counter dump and counter
schema commands would have repeated entries with the same name, due to which the commands became invalid.
With this fix, each entry in the JSON emitted by the counter dump and counter schema has an array of labeled counters or schemas. The entries are no longer
repeated, can be parsed, and are valid JSON.
(BZ#2229267)
Grafana panels for performance of daemons in the Ceph Dashboard now show correct data
Previously, the labels exporter weren't compatible with the queries used in the grafana dashboard. Due to this, the grafana panels were empty for Ceph daemons
performance in the Ceph Dashboard.
With this fix, the labels names are made compatible with the grafana dashboard queries and the grafana panels for performance of daemons show correct data.
(BZ#2217817)
RADOS
Ceph Monitor no longer crashes with a ceph_assert(osdmon()->is_writeable()) error
Previously, there was a problem in the stretch mode part of the Ceph Monitor when performing failover. Users could not exit the wait_for_writeable_ctx
function properly while waiting for a Paxos service called osdmon to be writeable. Due to this, the monitor would crash with a ceph_assert(osdmon()>is_writeable()) error.
With this fix, the function is exited by returning nothing after going into wait_for_writeable_ctx error and crashes no longer occur when we failover in stretch
cluster.
(BZ#2138216)
Introduction to IBM Storage Ceph
IBM Storage Ceph cluster is a distributed data object store designed to provide excellent performance, reliability and scalability.
IBM Storage Ceph is a scalable, open, software-defined storage platform that combines an enterprise-hardened version of the Ceph storage system, with a Ceph
management platform, deployment utilities, and support services. IBM Storage Ceph is designed for cloud infrastructure and web-scale object storage.
Distributed object stores are the future of storage, because they accommodate unstructured data, and because clients can use modern object interfaces and legacy
interfaces simultaneously.
For example,
APIs in many languages (C/C++, Java, Python)
RESTful interfaces (S3/Swift)
Block device interface
Filesystem interface
The power of IBM Storage Ceph cluster can transform your organization’s IT infrastructure and your ability to manage vast amounts of data, especially for cloud computing
platforms like Red Hat Enterprise Linux OSP. The cluster delivers extraordinary scalability–thousands of clients accessing petabytes to exabytes of data and beyond.
At the heart of every Ceph deployment is the IBM Storage Ceph cluster. It consists of three types of daemons:
Ceph OSD Daemon
Ceph OSDs store data on behalf of Ceph clients. Additionally, Ceph OSDs utilize the CPU, memory and networking of Ceph nodes to perform data replication,
erasure coding, rebalancing, recovery, monitoring and reporting functions.
Ceph Monitor
A Ceph Monitor maintains a master copy of the IBM Ceph Storage cluster map with the current state of the cluster. Monitors require high consistency, and use Paxos
to ensure agreement about the state of the cluster.
30 IBM Storage Ceph
Ceph Manager
The Ceph Manager maintains detailed information about placement groups, process metadata and host metadata in lieu of the Ceph Monitor—significantly
improving performance at scale. The Ceph Manager handles execution of many of the read-only Ceph CLI queries, such as placement group statistics. The Ceph
Manager also provides the RESTful monitoring APIs.
Figure 1. Daemons
Ceph client interfaces read data from and write data to the IBM Ceph Storage cluster. Clients need the following data to communicate with the IBM Ceph Storage cluster:
The Ceph configuration file, or the cluster name (usually ceph) and the monitor address.
The pool name.
The user name and the path to the secret key.
Ceph clients maintain object IDs and the pool names where they store the objects. However, they do not need to maintain an object-to-OSD index or communicate with a
centralized object index to look up object locations. Then, Ceph clients provide an object name and pool name to librados, which computes an object’s placement group
and the primary OSD for storing and retrieving data using the CRUSH (Controlled Replication Under Scalable Hashing) algorithm. The Ceph client connects to the primary
OSD where it may perform read and write operations. There is no intermediary server, broker or bus between the client and the OSD.
When an OSD stores data, it receives data from a Ceph client—whether the client is a Ceph Block Device, a Ceph Object Gateway, a Ceph Filesystem or another interface
and it stores the data as an object.
Note: An object ID is unique across the entire cluster, not just an OSD’s storage media.
Ceph OSDs store all data as objects in a flat namespace. There are no hierarchies of directories. An object has a cluster-wide unique identifier, binary data, and metadata
consisting of a set of name/value pairs.
Figure 2. Object
Ceph clients define the semantics for the client’s data format. For example, the Ceph block device maps a block device image to a series of objects stored across the
cluster.
Note: Objects consisting of a unique ID, data, and name/value paired metadata can represent both structured and unstructured data, as well as legacy and leading edge
data storage interfaces.
IBM Storage Ceph clusters consist of the following types of nodes:
Ceph Monitor
Each Ceph Monitor node runs the ceph-mon daemon, which maintains a primary copy of the storage cluster map. The storage cluster map includes the storage
cluster topology. A client connecting to the Ceph storage cluster retrieves the current copy of the storage cluster map from the Ceph Monitor, enabling the client to
read from and write data to the storage cluster.
Important: The storage cluster can run with just one Ceph Monitor; however, to ensure high availability in a production storage cluster, IBM supports deployments
with at least three Ceph Monitor nodes. Deploy a total of 5 Ceph Monitors for storage clusters exceeding 750 Ceph OSDs.
Ceph Manager
The Ceph Manager daemon, ceph-mgr, co-exists with the Ceph Monitor daemons running on Ceph Monitor nodes to provide extra services. The Ceph Manager
provides an interface for other monitoring and management systems using Ceph Manager modules. Running the Ceph Manager daemons is a requirement for
normal storage cluster operations.
Ceph OSD
Each Ceph Object Storage Device (OSD) node runs the ceph-osd daemon, which interacts with logical disks that are attached to the node. The storage cluster
stores data on these Ceph OSD nodes.
Ceph can run with few OSD nodes, of which the default is three, but production storage clusters realize better performance beginning at modest scales. For
example, 50 Ceph OSDs in a storage cluster. Ideally, a Ceph storage cluster has multiple OSD nodes, allowing for the possibility to isolate failure domains by
configuring the CRUSH map.
Ceph MDS
Each Ceph Metadata Server (MDS) node runs the ceph-mds daemon, which manages metadata related to files stored on the Ceph File System (CephFS). The Ceph
MDS daemon also coordinates access to the shared storage cluster.
Ceph Object Gateway
Ceph Object Gateway node runs the ceph-radosgw daemon, and is an object storage interface built on top of librados to provide applications with a RESTful
access point to the Ceph storage cluster. The Ceph Object Gateway supports two interfaces:
S3
Provides object storage functionality with an interface that is compatible with a large subset of the Amazon S3 RESTful API.
IBM Storage Ceph 31
Swift
Provides object storage functionality with an interface that is compatible with a large subset of the OpenStack Swift API.
IBM Storage Ceph is also available within a unified offering in IBM Storage Ready Nodes for IBM Storage Ceph. This offering provides Ready Nodes as IBM supported
servers validated for IBM Storage Ceph 6. For more information, see IBM Storage Ready Nodes for IBM Storage Ceph.
Note:
For more information about Ceph architecture, see Architecture.
For the minimum hardware recommendations, see Hardware.
Overview
Learn about the architecture, data security, and hardening concepts for IBM Storage Ceph.
Architecture
Data security and hardening
Use this information to learn about data security and hardening information for IBM Storage Ceph Clusters and their clients. The information here also provides
advice and good practices information for hardening the security of IBM Storage Ceph, with a focus on the Ceph Orchestrator using cephadm for IBM Storage Ceph
deployments.
Architecture
Know the architecture information for Ceph Storage Clusters and their clients.
Ceph architecture
Ceph client components
Core Ceph components
Ceph architecture
IBM Storage Ceph cluster is a distributed data object store designed to provide excellent performance, reliability and scalability. Distributed object stores are the future of
storage, because they accommodate unstructured data, and because clients can use modern object interfaces and legacy interfaces simultaneously.
For example:
APIs in many languages (C/C++, Java, Python)
RESTful interfaces (S3/Swift)
Block device interface
Filesystem interface
The power of IBM Storage Ceph cluster can transform your organization’s IT infrastructure and your ability to manage vast amounts of data, especially for cloud computing
platforms like Red Hat Enterprise Linux OSP. The cluster delivers extraordinary scalability–thousands of clients accessing petabytes to exabytes of data and beyond.
At the heart of every Ceph deployment is the IBM Storage Ceph cluster. It consists of three types of daemons:
Ceph OSD Daemon
Ceph OSDs store data on behalf of Ceph clients. Additionally, Ceph OSDs utilize the CPU, memory and networking of Ceph nodes to perform data replication,
erasure coding, rebalancing, recovery, monitoring and reporting functions.
Ceph Monitor
A Ceph Monitor maintains a master copy of the IBM Ceph Storage cluster map with the current state of the cluster. Monitors require high consistency, and use Paxos
to ensure agreement about the state of the cluster.
Ceph Manager
The Ceph Manager maintains detailed information about placement groups, process metadata and host metadata in lieu of the Ceph Monitor—significantly
improving performance at scale. The Ceph Manager handles execution of many of the read-only Ceph CLI queries, such as placement group statistics. The Ceph
Manager also provides the RESTful monitoring APIs.
Figure 1. Daemons
Ceph client interfaces read data from and write data to the IBM Ceph Storage cluster. Clients need the following data to communicate with the IBM Ceph Storage cluster:
32 IBM Storage Ceph
The Ceph configuration file, or the cluster name (usually ceph) and the monitor address.
The pool name.
The user name and the path to the secret key.
Ceph clients maintain object IDs and the pool names where they store the objects. However, they do not need to maintain an object-to-OSD index or communicate with a
centralized object index to look up object locations. Then, Ceph clients provide an object name and pool name to librados, which computes an object’s placement group
and the primary OSD for storing and retrieving data using the CRUSH (Controlled Replication Under Scalable Hashing) algorithm. The Ceph client connects to the primary
OSD where it may perform read and write operations. There is no intermediary server, broker or bus between the client and the OSD.
When an OSD stores data, it receives data from a Ceph client—whether the client is a Ceph Block Device, a Ceph Object Gateway, a Ceph Filesystem or another interface
and it stores the data as an object.
Note: An object ID is unique across the entire cluster, not just an OSD’s storage media.
Ceph OSDs store all data as objects in a flat namespace. There are no hierarchies of directories. An object has a cluster-wide unique identifier, binary data, and metadata
consisting of a set of name/value pairs.
Figure 2. Object
Ceph clients define the semantics for the client’s data format. For example, the Ceph block device maps a block device image to a series of objects stored across the
cluster.
Note: Objects consisting of a unique ID, data, and name/value paired metadata can represent both structured and unstructured data, as well as legacy and leading edge
data storage interfaces.
Ceph client components
Ceph clients differ materially in how they present data storage interfaces. A Ceph Block Device presents block storage that mounts just like a physical storage drive. A
Ceph gateway presents an object storage service with S3-compliant and Swift-compliant RESTful interfaces with its own user management. However, all Ceph clients use
the Reliable Autonomic Distributed Object Store (RADOS) protocol to interact with the IBM Storage Ceph cluster.
They all have the same basic needs:
The Ceph configuration file, and the Ceph monitor address.
The pool name.
The user name and the path to the secret key.
Ceph clients tend to follow some similar patterns, such as object-watch-notify and striping. The following sections describe a little bit more about RADOS, librados and
common patterns used in Ceph clients.
Prerequisites
Ceph client native protocol
Ceph client object watch and notify
Ceph client Mandatory Exclusive Locks
Ceph client object map
Ceph client data striping
Ceph on-wire encryption
Prerequisites
A basic understanding of distributed storage systems.
Ceph client native protocol
Modern applications need a simple object storage interface with asynchronous communication capability. The IBM Storage Ceph Cluster provides a simple object storage
interface with asynchronous communication capability. The interface provides direct, parallel access to objects throughout the cluster.
Pool Operations
Snapshots
Read/Write Objects
IBM Storage Ceph 33
Create or Remove
Entire Object or Byte Range
Append or Truncate
Create/Set/Get/Remove XATTRs
Create/Set/Get/Remove Key/Value Pairs
Compound operations and dual-ack semantics
Ceph client object watch and notify
A Ceph client can register a persistent interest with an object and keep a session to the primary OSD open. The client can send a notification message and payload to all
watchers and receive notification when the watchers receive the notification. This enables a client to use any object as a synchronization/communication channel.
Figure 1. Ceph client object watch and notify
Ceph client Mandatory Exclusive Locks
Mandatory Exclusive Locks is a feature that locks an RBD to a single client, if multiple mounts are in place. This helps address the write conflict situation when multiple
mounted clients try to write to the same object. This feature is built on object-watch-notify explained in the previous section. So, when writing, if one client first
establishes an exclusive lock on an object, another mounted client will first check to see if a peer has placed a lock on the object before writing.
With this feature enabled, only one client can modify an RBD device at a time, especially when changing internal RBD structures during operations like snapshot
create/delete. It also provides some protection for failed clients. For instance, if a virtual machine seems to be unresponsive and you start a copy of it with the same
disk elsewhere, the first one will be blacklisted in Ceph and unable to corrupt the new one.
Mandatory Exclusive Locks are not enabled by default. You have to explicitly enable it with --image-feature parameter when creating an image.
Example
[root@mon ~]# rbd create --size 102400 mypool/myimage --image-feature 13
Here, the numeral 13 is a summation of 1, 4 and 8 where 1 enables layering support, 4 enables exclusive locking support and 8 enables object map support. So, the above
command creates a 100 GB RBD image, enable layering, exclusive lock and object map.
Mandatory Exclusive Locks is also a prerequisite for object map. Without enabling exclusive locking support, object map support cannot be enabled.
34 IBM Storage Ceph
Mandatory Exclusive Locks also does some ground work for mirroring.
Ceph client object map
Object map is a feature that tracks the presence of backing RADOS objects when a client writes to an rbd image. When a write occurs, that write is translated to an offset
within a backing RADOS object. When the object map feature is enabled, the presence of these RADOS objects is tracked. So, we can know if the objects actually exist.
Object map is kept in-memory on the librbd client so it can avoid querying the OSDs for objects that it knows don’t exist. In other words, object map is an index of the
objects that actually exist.
Object map is beneficial for certain operations, viz:
Resize
Export
Copy
Flatten
Delete
Read
A shrink resize operation is like a partial delete where the trailing objects are deleted.
An export operation knows which objects are to be requested from RADOS.
A copy operation knows which objects exist and need to be copied. It does not have to iterate over potentially hundreds and thousands of possible objects.
A flatten operation performs a copy-up for all parent objects to the clone so that the clone can be detached from the parent i.e, the reference from the child clone to the
parent snapshot can be removed. So, instead of all potential objects, copy-up is done only for the objects that exist.
A delete operation deletes only the objects that exist in the image.
A read operation skips the read for objects it knows doesn’t exist.
So, for operations like resize, shrinking only, exporting, copying, flattening, and deleting, these operations would need to issue an operation for all potentially affected
RADOS objects, whether they exist or not. With object map enabled, if the object doesn’t exist, the operation need not be issued.
For example, if we have a 1 TB sparse RBD image, it can have hundreds and thousands of backing RADOS objects. A delete operation without object map enabled would
need to issue a remove
object operation for each potential object in the image. But if object map is enabled, it only needs to issue remove object operations for the objects that exist.
Object map is valuable against clones that don’t have actual objects but get objects from parents. When there is a cloned image, the clone initially has no objects and all
reads are redirected to the parent. So, object map can improve reads as without the object map, first it needs to issue a read operation to the OSD for the clone, when that
fails, it issues another read to the parent — with object map enabled. It skips the read for objects it knows doesn’t exist.
Object map is not enabled by default. You have to explicitly enable it with --image-features parameter when creating an image. Also, Mandatory
Exclusive Locks is a prerequisite for object map. Without enabling exclusive locking support, object map support cannot be enabled. To enable object map support
when creating a image, execute:
Example
[root@mon ~]# rbd create --size 102400 mypool/myimage --image-feature 13
Here, the numeral 13 is a summation of 1, 4 and 8 where 1 enables layering support, 4 enables exclusive locking support and 8 enables object map support. So, the above
command creates a 100 GB RBD image, enable layering, exclusive lock and object map.
Ceph client data striping
Storage devices have throughput limitations, which impact performance and scalability. So storage systems often support striping—storing sequential pieces of
information across multiple storage devices—to increase throughput and performance. The most common form of data striping comes from RAID. The RAID type most
similar to Ceph’s striping is RAID 0, or a striped volume. Ceph’s striping offers the throughput of RAID 0 striping, the reliability of n-way RAID mirroring and faster recovery.
Ceph provides three types of clients: Ceph Block Device, Ceph Filesystem, and Ceph Object Storage. A Ceph Client converts its data from the representation format it
provides to its users, such as a block device image, RESTful objects, CephFS filesystem directories, into objects for storage in the Ceph Storage Cluster.
Tip: The objects Ceph stores in the Ceph Storage Cluster are not striped. Ceph Object Storage, Ceph Block Device, and the Ceph Filesystem stripe their data over multiple
Ceph Storage Cluster objects. Ceph Clients that write directly to the Ceph storage cluster using librados must perform the striping, and parallel I/O for themselves to
obtain these benefits.
The simplest Ceph striping format involves a stripe count of 1 object. Ceph Clients write stripe units to a Ceph Storage Cluster object until the object is at its maximum
capacity, and then create another object for additional stripes of data. The simplest form of striping may be sufficient for small block device images, S3 or Swift objects.
However, this simple form doesn’t take maximum advantage of Ceph’s ability to distribute data across placement groups, and consequently doesn’t improve performance
very much. The following diagram depicts the simplest form of striping:
Figure 1. Data Striping
IBM Storage Ceph 35
If you anticipate large images sizes, large S3 or Swift objects for example, video, you may see considerable read/write performance improvements by striping client data
over multiple objects within an object set. Significant write performance occurs when the client writes the stripe units to their corresponding objects in parallel. Since
objects get mapped to different placement groups and further mapped to different OSDs, each write occurs in parallel at the maximum write speed. A write to a single disk
would be limited by the head movement for example, 6ms per seek and bandwidth of that one device for example, 100MB/s. By spreading that write over multiple objects,
which map to different placement groups and OSDs, Ceph can reduce the number of seeks per drive and combine the throughput of multiple drives to achieve much faster
write or read speeds.
Note: Striping is independent of object replicas. Since CRUSH replicates objects across OSDs, stripes get replicated automatically.
In the following diagram, client data gets striped across an object set (object set
1 in the following diagram) consisting of 4 objects, where the first stripe unit is stripe unit 0 in object 0, and the fourth stripe unit is stripe unit 3 in object 3.
After writing the fourth stripe, the client determines if the object set is full. If the object set is not full, the client begins writing a stripe to the first object again, see object
0 in the following diagram. If the object set is full, the client creates a new object set, see object set 2 in the following diagram, and begins writing to the first stripe,
with a stripe unit of 16, in the first object in the new object set, see object 4 in the diagram below.
Figure 2. Client Data Stripping
36 IBM Storage Ceph
Three important variables determine how Ceph stripes data:
Object Size
Objects in the Ceph Storage Cluster have a maximum configurable size, such as 2 MB, or 4 MB. The object size should be large enough to accommodate many stripe
units, and should be a multiple of the stripe unit.
Important: IBM recommends a safe maximum value of 16 MB.
Stripe Width
Stripes have a configurable unit size, for example 64 KB. The Ceph Client divides the data it will write to objects into equally sized stripe units, except for the last
stripe unit. A stripe width should be a fraction of the Object Size so that an object may contain many stripe units.
Stripe Count
The Ceph Client writes a sequence of stripe units over a series of objects determined by the stripe count. The series of objects is called an object set. After the Ceph
Client writes to the last object in the object set, it returns to the first object in the object set.
Important: Test the performance of your striping configuration before putting your cluster into production. You CANNOT change these striping parameters after you stripe
the data and write it to objects.
Once the Ceph Client has striped data to stripe units and mapped the stripe units to objects, Ceph’s CRUSH algorithm maps the objects to placement groups, and the
placement groups to Ceph OSD Daemons before the objects are stored as files on a storage disk.
Note: Since a client writes to a single pool, all data striped into objects get mapped to placement groups in the same pool. So they use the same CRUSH map and the same
access controls.
IBM Storage Ceph 37
Ceph on-wire encryption
You can enable encryption for all Ceph traffic over the network with the introduction of the messenger version 2 protocol. The secure mode setting for messenger v2
encrypts communication between Ceph daemons and Ceph clients, giving you end-to-end encryption.
The second version of Ceph’s on-wire protocol, msgr2, includes several new features:
A secure mode encrypting all data moving through the network.
Encapsulation improvement of authentication payloads.
Improvements to feature advertisement and negotiation.
The Ceph daemons bind to multiple ports allowing both the legacy, v1-compatible, and the new, v2-compatible, Ceph clients to connect to the same storage cluster. Ceph
clients or other Ceph daemons connecting to the Ceph Monitor daemon will try to use the v2 protocol first, if possible, but if not, then the legacy v1 protocol will be used.
By default, both messenger protocols, v1 and v2, are enabled. The new v2 port is 3300, and the legacy v1 port is 6789, by default.
The messenger v2 protocol has two configuration options that control whether the v1 or the v2 protocol is used:
ms_bind_msgr1
This option controls whether a daemon binds to a port speaking the v1 protocol; it is true by default.
ms_bind_msgr2
This option controls whether a daemon binds to a port speaking the v2 protocol; it is true by default.
Similarly, two options control based on IPv4 and IPv6 addresses used:
ms_bind_ipv4
This option controls whether a daemon binds to an IPv4 address; it is true by default.
ms_bind_ipv6
This option controls whether a daemon binds to an IPv6 address; it is true by default.
The msgr2 protocol supports two connection modes:
crc
Provides strong initial authentication when a connection is established with cephx.
Provides a crc32c integrity check to protect against bit flips.
Does not provide protection against a malicious man-in-the-middle attack.
Does not prevent an eavesdropper from seeing all post-authentication traffic.
secure
Provides strong initial authentication when a connection is established with cephx.
Provides full encryption of all post-authentication traffic.
Provides a cryptographic integrity check.
The default mode is crc.
Ensure that you consider cluster CPU requirements when you plan the IBM Storage Ceph deployment to include encryption overhead.
Important: Using secure mode is supported by Ceph clients using librbd, such as OpenStack Nova, Glance, and Cinder.
Address Changes
For both versions of the messenger protocol to co-exist in the same storage cluster, the address formatting has changed:
Syntax for old address format
IP_ADDR:PORT/CLIENT_ID
For example, 1.2.3.4:5678/91011
Syntax for new address format
PROTOCOL_VERSION:IP_ADDR:PORT/CLIENT_ID
For example, v2:1.2.3.4:5678/91011, where PROTOCOL_VERSION can be either v1 or v2.
Because the Ceph daemons now bind to multiple ports, the daemons display multiple addresses instead of a single address. Here is an example from a dump of the
monitor map:
epoch 1
fsid 50fcf227-be32-4bcb-8b41-34ca8370bd17
last_changed 2021-12-12 11:10:46.700821
created 2021-12-12 11:10:46.700821
min_mon_release 14 (nautilus)
0: [v2:10.0.0.10:3300/0,v1:10.0.0.10:6789/0] mon.a
1: [v2:10.0.0.11:3300/0,v1:10.0.0.11:6789/0] mon.b
2: [v2:10.0.0.12:3300/0,v1:10.0.0.12:6789/0] mon.c
Also, the mon_host configuration option and specifying addresses on the command line using -m supports the new address format.
Connection Phases
There are four phases for making an encrypted connection:
Banner
On connection, both the client and the server send a banner. Currently, the Ceph banner is ceph 0 0n.
38 IBM Storage Ceph
Authentication Exchange
All data, sent or received, is contained in a frame for the duration of the connection. The server decides if authentication has completed, and what the connection
mode will be. The frame format is fixed, and can be in three different forms depending on the authentication flags being used.
Message Flow Handshake Exchange
The peers identify each other and establish a session. The client sends the first message, and the server will reply with the same message. The server can close
connections if the client talks to the wrong daemon. For new sessions, the client and server proceed to exchanging messages. Client cookies are used to identify a
session, and can reconnect to an existing session.
Message Exchange
The client and server start exchanging messages, until the connection is closed.
Reference
See the IBM Ceph Storage Data Security and Hardening for details on enabling the msgr2 protocol.
Core Ceph components
An IBM Storage Ceph cluster can have a large number of Ceph nodes for limitless scalability, high availability and performance. Each node leverages non-proprietary
hardware and intelligent Ceph daemons that communicate with each other to:
Write and read data
Compress data
Ensure durability by replicating or erasure coding data
Monitor and report on cluster health--also called 'heartbeating'
Redistribute data dynamically--also called 'backfilling'
Ensure data integrity; and,
Recover from failures.
To the Ceph client interface that reads and writes data, an IBM Storage Ceph cluster looks like a simple pool where it stores data. However, librados and the storage
cluster perform many complex operations in a manner that is completely transparent to the client interface. Ceph clients and Ceph OSDs both use the CRUSH (Controlled
Replication Under Scalable Hashing) algorithm. The following sections provide details on how CRUSH enables Ceph to perform these operations seamlessly.
Prerequisites
Pools
The Ceph storage cluster stores data objects in logical partitions called pools. Ceph administrators can create pools for particular types of data, such as for block
devices, object gateways, or simply just to separate one group of users from another. Pools play an important role in data durability, performance, and high
availability to Ceph.
Ceph authentication
To identify users and protect against man-in-the-middle attacks, Ceph provides its cephx authentication system, which authenticates users and daemons.
Placement groups
CRUSH ruleset
Input/output operations
Replication
Erasure coding
ObjectStore
BlueStore
Self management operations
Heartbeat
Peering
Ceph stores copies of placement groups on multiple OSDs. Each copy of a placement group has a status. These OSDs peer check each other to ensure that they
agree on the status of each copy of the placement group. Peering issues usually resolve themselves.
Rebalancing and recovery
Data integrity
High availability
Clustering the Ceph Monitor
Prerequisites
A basic understanding of distributed storage systems.
Pools
The Ceph storage cluster stores data objects in logical partitions called pools. Ceph administrators can create pools for particular types of data, such as for block devices,
object gateways, or simply just to separate one group of users from another. Pools play an important role in data durability, performance, and high availability to Ceph.
From the perspective of a Ceph client, the storage cluster is very simple. When a Ceph client reads or writes data using an I/O context, it always connects to a storage pool
in the Ceph storage cluster. The client specifies the pool name, a user and a secret key, so the pool appears to act as a logical partition with access controls to its data
IBM Storage Ceph 39
objects.
In fact, a Ceph pool is not only a logical partition for storing object data. A pool plays a critical role in how the Ceph storage cluster distributes and stores data. However,
these complex operations are transparent to the Ceph client.
Ceph pools define:
Pool type
In early versions of Ceph, a pool simply maintained multiple deep copies of an object. Today, Ceph can maintain multiple copies of an object, or it can use erasure
coding to ensure durability. The data durability method is pool-wide, and does not change after creating the pool. The pool type defines the data durability method
when creating the pool. Pool types are completely transparent to the client.
Placement groups
In an exabyte scale storage cluster, a Ceph pool might store millions of data objects or more. Ceph must handle many types of operations, including data durability
via replicas or erasure code chunks, data integrity by scrubbing or CRC checks, replication, rebalancing and recovery. Consequently, managing data on a per-object
basis presents a scalability and performance bottleneck. Ceph addresses this bottleneck by sharding a pool into placement groups. The CRUSH algorithm computes
the placement group for storing an object and computes the Acting Set of OSDs for the placement group. CRUSH puts each object into a placement group. Then,
CRUSH stores each placement group in a set of OSDs. System administrators set the placement group count when creating or modifying a pool.
CRUSH ruleset
CRUSH can detect failure domains and performance domains. CRUSH can identify OSDs by storage media type and organize OSDs hierarchically into nodes, racks,
and rows. CRUSH enables Ceph OSDs to store object copies across failure domains. For example, copies of an object may get stored in different server rooms,
aisles, racks and nodes. If a large part of a cluster fails, such as a rack, the cluster can still operate in a degraded state until the cluster recovers.
Additionally, CRUSH enables clients to write data to particular types of hardware, such as SSDs, hard drives with SSD journals, or hard drives with journals on the
same drive as the data. The CRUSH ruleset determines failure domains and performance domains for the pool. Administrators set the CRUSH ruleset when creating
a pool.
Note: An administrator CANNOT change a pool’s ruleset after creating the pool.
Durability
In exabyte scale storage clusters, hardware failure is an expectation and not an exception. When using data objects to represent larger grained storage interfaces
such as a block device, losing one or more data objects for that larger grained interface can compromise the integrity of the larger grained storage entity potentially
rendering it useless. So data loss is intolerable. Ceph provides high data durability in two ways:
Replica pools store multiple deep copies of an object using the CRUSH failure domain to physically separate one data object copy from another. That is,
copies get distributed to separate physical hardware. This increases durability during hardware failures.
Erasure coded pools store each object as K+M chunks, where K represents data chunks and M represents coding chunks. The sum represents the number of
OSDs used to store the object and the M value represents the number of OSDs that can fail and still restore data should the M number of OSDs fail.
Ceph authentication
To identify users and protect against man-in-the-middle attacks, Ceph provides its cephx authentication system, which authenticates users and daemons.
Note: The cephx protocol does not address data encryption for data transported over the network or data stored in OSDs.
Cephx uses shared secret keys for authentication, meaning both the client and the monitor cluster have a copy of the client’s secret key. The authentication protocol
enables both parties to prove to each other that they have a copy of the key without actually revealing it. This provides mutual authentication, which means the cluster is
sure the user possesses the secret key, and the user is sure that the cluster has a copy of the secret key.
Cephx
The cephx authentication protocol operates in a manner similar to Kerberos.
A user/actor invokes a Ceph client to contact a monitor. Unlike Kerberos, each monitor can authenticate users and distribute keys, so there is no single point of failure or
bottleneck when using cephx. The monitor returns an authentication data structure similar to a Kerberos ticket that contains a session key for use in obtaining Ceph
services. This session key is itself encrypted with the user’s permanent secret key, so that only the user can request services from the Ceph monitors. The client then uses
the session key to request its desired services from the monitor, and the monitor provides the client with a ticket that will authenticate the client to the OSDs that actually
handle data. Ceph monitors and OSDs share a secret, so the client can use the ticket provided by the monitor with any OSD or metadata server in the cluster. Like
Kerberos, cephx tickets expire, so an attacker cannot use an expired ticket or session key obtained surreptitiously. This form of authentication will prevent attackers with
access to the communications medium from either creating bogus messages under another user’s identity or altering another user’s legitimate messages, as long as the
user’s secret key is not divulged before it expires.
To use cephx, an administrator must set up users first. In the following diagram, the client.admin user invokes ceph auth get-or-create-key from the command
line to generate a username and secret key. Ceph’s auth subsystem generates the username and key, stores a copy with the monitor(s) and transmits the user’s secret
back to the client.admin user. This means that the client and the monitor share a secret key.
Note: The client.admin user must provide the user ID and secret key to the user in a secure manner.
Figure 1. CephX
40 IBM Storage Ceph
Placement groups
Storing millions of objects in a cluster and managing them individually is resource intensive. So Ceph uses placement groups (PGs) to make managing a huge number of
objects more efficient.
A PG is a subset of a pool that serves to contain a collection of objects. Ceph shards a pool into a series of PGs. Then, the CRUSH algorithm takes the cluster map and the
status of the cluster into account and distributes the PGs evenly and pseudo-randomly to OSDs in the cluster.
Here is how it works:
When a system administrator creates a pool, CRUSH creates a user-defined number of PGs for the pool. Generally, the number of PGs should be a reasonably fine-grained
subset of the data. For example, 100 PGs per OSD per pool would mean that each PG contains approximately 1% of the pool’s data.
The number of PGs has a performance impact when Ceph needs to move a PG from one OSD to another OSD. If the pool has too few PGs, Ceph will move a large
percentage of the data simultaneously and the network load will adversely impact the cluster’s performance. If the pool has too many PGs, Ceph will use too much CPU
and RAM when moving tiny percentages of the data and thereby adversely impact the cluster’s performance. For details on calculating the number of PGs to achieve
optimal performance, see PG Count
Ceph ensures against data loss by storing replicas of an object or by storing erasure code chunks of an object. Since Ceph stores objects or erasure code chunks of an
object within PGs, Ceph replicates each PG in a set of OSDs called the Acting Set for each copy of an object or each erasure code chunk of an object. A system
administrator can determine the number of PGs in a pool and the number of replicas or erasure code chunks. However, the CRUSH algorithm calculates which OSDs are in
the acting set for a particular PG.
The CRUSH algorithm and PGs make Ceph dynamic. Changes in the cluster map or the cluster state may result in Ceph moving PGs from one OSD to another automatically.
Here are a few examples:
Expanding the Cluster: When adding a new host and its OSDs to the cluster, the cluster map changes. Since CRUSH evenly and pseudo-randomly distributes PGs
to OSDs throughout the cluster, adding a new host and its OSDs means that CRUSH will reassign some of the pool’s placement groups to those new OSDs. That
means that system administrators do not have to rebalance the cluster manually. Also, it means that the new OSDs contain approximately the same amount of data
as the other OSDs. This also means that new OSDs do not contain newly written OSDs, preventing hot spots in the cluster.
An OSD Fails: When an OSD fails, the state of the cluster changes. Ceph temporarily loses one of the replicas or erasure code chunks, and needs to make another
copy. If the primary OSD in the acting set fails, the next OSD in the acting set becomes the primary and CRUSH calculates a new OSD to store the additional copy or
erasure code chunk.
By managing millions of objects within the context of hundreds to thousands of PGs, the Ceph storage cluster can grow, shrink and recover from failure efficiently.
For Ceph clients, the CRUSH algorithm via librados makes the process of reading and writing objects very simple. A Ceph client simply writes an object to a pool or
reads an object from a pool. The primary OSD in the acting set can write replicas of the object or erasure code chunks of the object to the secondary OSDs in the acting set
on behalf of the Ceph client.
If the cluster map or cluster state changes, the CRUSH computation for which OSDs store the PG will change too. For example, a Ceph client may write object foo to the
pool bar. CRUSH will assign the object to PG 1.a, and store it on OSD 5, which makes replicas on OSD 10 and OSD 15 respectively. If OSD 5 fails, the cluster state
changes. When the Ceph client reads object foo from pool bar, the client via librados will automatically retrieve it from OSD 10 as the new primary OSD dynamically.
The Ceph client via librados connects directly to the primary OSD within an acting set when writing and reading objects. Since I/O operations do not use a centralized
broker, network oversubscription is typically NOT an issue with Ceph.
The following diagram depicts how CRUSH assigns objects to PGs, and PGs to OSDs. The CRUSH algorithm assigns the PGs to OSDs such that each OSD in the acting set is
in a separate failure domain, which typically means the OSDs will always be on separate server hosts and sometimes in separate racks.
Figure 1. Placement Groups
IBM Storage Ceph 41
CRUSH ruleset
Ceph assigns a CRUSH ruleset to a pool. When a Ceph client stores or retrieves data in a pool, Ceph identifies the CRUSH ruleset, a rule within the rule set, and the toplevel bucket in the rule for storing and retrieving data. As Ceph processes the CRUSH rule, it identifies the primary OSD that contains the placement group for an object.
That enables the client to connect directly to the OSD, access the placement group and read or write object data.
To map placement groups to OSDs, a CRUSH map defines a hierarchical list of bucket types. The list of bucket types are located under types in the generated CRUSH
map. The purpose of creating a bucket hierarchy is to segregate the leaf nodes by their failure domains and/or performance domains, such as drive type, hosts, chassis,
racks, power distribution units, pods, rows, rooms, and data centers.
With the exception of the leaf nodes representing OSDs, the rest of the hierarchy is arbitrary. Administrators may define it according to their own needs if the default types
don’t suit their requirements. CRUSH supports a directed acyclic graph that models the Ceph OSD nodes, typically in a hierarchy. So Ceph administrators can support
multiple hierarchies with multiple root nodes in a single CRUSH map. For example, an administrator can create a hierarchy representing higher cost SSDs for high
performance, and a separate hierarchy of lower cost hard drives with SSD journals for moderate performance.
Input/output operations
Ceph clients retrieve a Cluster Map from a Ceph monitor, bind to a pool, and perform input/output (I/O) on objects within placement groups in the pool. The pool’s CRUSH
ruleset and the number of placement groups are the main factors that determine how Ceph will place the data. With the latest version of the cluster map, the client knows
about all of the monitors and OSDs in the cluster and their current state. However, the client doesn’t know anything about object locations.
The only inputs required by the client are the object ID and the pool name. It is simple: Ceph stores data in named pools. When a client wants to store a named object in a
pool it takes the object name, a hash code, the number of PGs in the pool and the pool name as inputs; then, CRUSH (Controlled Replication Under Scalable Hashing)
calculates the ID of the placement group and the primary OSD for the placement group.
Ceph clients use the following steps to compute PG IDs.
1. The client inputs the pool ID and the object ID. For example, pool = liverpool and object-id = john.
2. CRUSH takes the object ID and hashes it.
3. CRUSH calculates the hash modulo of the number of PGs to get a PG ID. For example, 58.
4. CRUSH calculates the primary OSD corresponding to the PG ID.
5. The client gets the pool ID given the pool name. For example, the pool liverpool is pool number 4.
6. The client prepends the pool ID to the PG ID. For example, 4.58.
7. The client performs an object operation such as write, read, or delete by communicating directly with the Primary OSD in the Acting Set.
The topology and state of the Ceph storage cluster are relatively stable during a session. Empowering a Ceph client via librados to compute object locations is much
faster than requiring the client to make a query to the storage cluster over a chatty session for each read/write operation. The CRUSH algorithm allows a client to compute
where objects should be stored, and enables the client to contact the primary OSD in the acting set directly to store or retrieve data in the objects. Since a cluster at the
exabyte scale has thousands of OSDs, network oversubscription between a client and a Ceph OSD is not a significant problem. If the cluster state changes, the client can
simply request an update to the cluster map from the Ceph monitor.
Replication
Like Ceph clients, Ceph OSDs can contact Ceph monitors to retrieve the latest copy of the cluster map. Ceph OSDs also use the CRUSH algorithm, but they use it to
compute where to store replicas of objects. In a typical write scenario, a Ceph client uses the CRUSH algorithm to compute the placement group ID and the primary OSD in
the Acting Set for an object. When the client writes the object to the primary OSD, the primary OSD finds the number of replicas that it should store. The value is found in
the osd_pool_default_size setting. Then, the primary OSD takes the object ID, pool name and the cluster map and uses the CRUSH algorithm to calculate the IDs of
secondary OSDs for the acting set. The primary OSD writes the object to the secondary OSDs. When the primary OSD receives an acknowledgment from the secondary
OSDs and the primary OSD itself completes its write operation, it acknowledges a successful write operation to the Ceph client.
42 IBM Storage Ceph
Figure 1. Replicated IO
With the ability to perform data replication on behalf of Ceph clients, Ceph OSD Daemons relieve Ceph clients from that duty, while ensuring high data availability and data
safety.
Note: The primary OSD and the secondary OSDs are typically configured to be in separate failure domains. CRUSH computes the IDs of the secondary OSDs with
consideration for the failure domains.
Data copies
In a replicated storage pool, Ceph needs multiple copies of an object to operate in a degraded state. Ideally, a Ceph storage cluster enables a client to read and write data
even if one of the OSDs in an acting set fails. For this reason, Ceph defaults to making three copies of an object with a minimum of two copies clean for write operations.
Ceph will still preserve data even if two OSDs fail. However, it will interrupt write operations.
In an erasure-coded pool, Ceph needs to store chunks of an object across multiple OSDs so that it can operate in a degraded state. Similar to replicated pools, ideally an
erasure-coded pool enables a Ceph client to read and write in a degraded state.
Important: The following jerasure coding values for k, and m are supported:
k=8 m=3
k=8 m=4
k=4 m=2
Erasure coding
Ceph can load one of many erasure code algorithms. The earliest and most commonly used is the Reed-Solomon algorithm. An erasure code is actually a forward error
correction (FEC) code. FEC code transforms a message of K chunks into a longer message called a code word of N chunks, such that Ceph can recover the original message
from a subset of the N chunks.
More specifically, N = K+M where the variable K is the original amount of data chunks. The variable M stands for the extra or redundant chunks that the erasure code
algorithm adds to provide protection from failures. The variable N is the total number of chunks created after the erasure coding process. The value of M is simply N-K
which means that the algorithm computes N-K redundant chunks from K original data chunks. This approach guarantees that Ceph can access all the original data. The
system is resilient to arbitrary N-K failures. For instance, in a 10 K of 16 N configuration, or erasure coding 10/16, the erasure code algorithm adds six extra chunks to the
10 base chunks K. For example, in a M = K-N or 16-10 = 6 configuration, Ceph will spread the 16 chunks N across 16 OSDs. The original file could be reconstructed
from the 10 verified N chunks even if 6 OSDs fail—ensuring that the IBM Storage Ceph cluster will not lose data, and thereby ensures a very high level of fault tolerance.
Like replicated pools, in an erasure-coded pool the primary OSD in the up set receives all write operations. In replicated pools, Ceph makes a deep copy of each object in
the placement group on the secondary OSDs in the set. For erasure coding, the process is a bit different. An erasure coded pool stores each object as K+M chunks. It is
divided into K data chunks and M coding chunks. The pool is configured to have a size of K+M so that Ceph stores each chunk in an OSD in the acting set. Ceph stores the
rank of the chunk as an attribute of the object. The primary OSD is responsible for encoding the payload into K+M chunks and sends them to the other OSDs. The primary
OSD is also responsible for maintaining an authoritative version of the placement group logs.
For example, in a typical configuration a system administrator creates an erasure coded pool to use six OSDs and sustain the loss of two of them. That is, (K+M = 6) such
that (M = 2).
When Ceph writes the object NYAN containing ABCDEFGHIJKL to the pool, the erasure encoding algorithm splits the content into four data chunks by simply dividing the
content into four parts: ABC, DEF, GHI, and JKL. The algorithm will pad the content if the content length is not a multiple of K. The function also creates two coding
chunks: the fifth with YXY and the sixth with QGC. Ceph stores each chunk on an OSD in the acting set, where it stores the chunks in objects that have the same name,
NYAN, but reside on different OSDs. The algorithm must preserve the order in which it created the chunks as an attribute of the object shard_t, in addition to its name.
For example, Chunk 1 contains ABC and Ceph stores it on OSD5 while chunk 5 contains YXY and Ceph stores it on OSD4.
Figure 1. Ceph client object watch and notify
IBM Storage Ceph 43
In a recovery scenario, the client attempts to read the object NYAN from the erasure-coded pool by reading chunks 1 through 6. The OSD informs the algorithm that
chunks 2 and 6 are missing. These missing chunks are called erasures. For example, the primary OSD could not read chunk 6 because the OSD6 is out, and could not read
chunk 2, because OSD2 was the slowest and its chunk was not taken into account. However, as soon as the algorithm has four chunks, it reads the four chunks: chunk 1
containing ABC, chunk 3 containing GHI, chunk 4 containing JKL, and chunk 5 containing YXY. Then, it rebuilds the original content of the object ABCDEFGHIJKL, and
original content of chunk 6, which contained QGC.
Splitting data into chunks is independent from object placement. The CRUSH ruleset along with the erasure-coded pool profile determines the placement of chunks on the
OSDs. For instance, using the Locally Repairable Code (lrc) plugin in the erasure code profile creates additional chunks and requires fewer OSDs to recover from. For
example, in an lrc profile configuration K=4 M=2 L=3, the algorithm creates six chunks (K+M), just as the jerasure plugin would, but the locality value (L=3) requires
that the algorithm create 2 more chunks locally. The algorithm creates the additional chunks as such, (K+M)/L. If the OSD containing chunk 0 fails, this chunk can be
recovered by using chunks 1, 2 and the first local chunk. In this case, the algorithm only requires 3 chunks for recovery instead of 5.
Note: Using erasure-coded pools disables Object Map.
Reference
For more information about CRUSH, the erasure-coding profiles, and plug-ins, see Storage strategies.
For more details on Object Map, see Ceph client object map.
ObjectStore
ObjectStore provides a low-level interface to an OSD’s raw block device. When a client reads or writes data, it interacts with the ObjectStore interface. Ceph write
operations are essentially ACID transactions: that is, they provide Atomicity, Consistency, Isolation and Durability. ObjectStore ensures that a Transaction is all-ornothing to provide Atomicity. The ObjectStore also handles object semantics. An object stored in the storage cluster has a unique identifier, object data and metadata.
So ObjectStore provides Consistency by ensuring that Ceph object semantics are correct. ObjectStore also provides the Isolation portion of an ACID transaction
by invoking a Sequencer on write operations to ensure that Ceph write operations occur sequentially. In contrast, an OSDs replication or erasure coding functionality
provides the Durability component of the ACID transaction. Since ObjectStore is a low-level interface to storage media, it also provides performance statistics.
Ceph implements several concrete methods for storing data:
BlueStore
A production grade implementation using a raw block device to store object data.
Memstore
A developer implementation for testing read/write operations directly in RAM.
K/V Store
An internal implementation for Ceph’s use of key/value databases.
Since administrators will generally only address BlueStore, the following sections will only describe those implementations in greater detail.
BlueStore
BlueStore is the next generation storage implementation for Ceph. As the market for storage devices now includes solid state drives or SSDs and non-volatile memory
over PCI Express or NVMe, their use in Ceph reveals some of the limitations of the FileStore storage implementation. While FileStore has many improvements to
facilitate SSD and NVMe storage, other limitations remain. Among them, increasing placement groups remains computationally expensive, and the double write penalty
44 IBM Storage Ceph
remains. Whereas, FileStore interacts with a file system on a block device, BlueStore eliminates that layer of indirection and directly consumes a raw block device for
object storage. BlueStore uses the very light weight BlueFS file system on a small partition for its k/v databases. BlueStore eliminates the paradigm of a directory
representing a placement group, a file representing an object and file XATTRs representing metadata. BlueStore also eliminates the double write penalty of FileStore,
so write operations are nearly twice as fast with BlueStore under most workloads.
BlueStore stores data as:
Object Data
In BlueStore, Ceph stores objects as blocks directly on a raw block device. The portion of the raw block device that stores object data does NOT contain a
filesystem. The omission of the filesystem eliminates a layer of indirection and thereby improves performance. However, much of the BlueStore performance
improvement comes from the block database and write-ahead log.
Block Database
In BlueStore, the block database handles the object semantics to guarantee Consistency. An object’s unique identifier is a key in the block database. The values
in the block database consist of a series of block addresses that refer to the stored object data, the object’s placement group, and object metadata. The block
database may reside on a BlueFS partition on the same raw block device that stores the object data, or it may reside on a separate block device, usually when the
primary block device is a hard disk drive and an SSD or NVMe will improve performance. The block database provides a number of improvements over FileStore;
namely, the key/value semantics of BlueStore do not suffer from the limitations of filesystem XATTRs. BlueStore may assign objects to other placement groups
quickly within the block database without the overhead of moving files from one directory to another, as is the case in FileStore. BlueStore also introduces new
features. The block database can store the checksum of the stored object data and its metadata, allowing full data checksum operations for each read, which is
more efficient than periodic scrubbing to detect bit rot. BlueStore can compress an object and the block database can store the algorithm used to compress an
object—ensuring that read operations select the appropriate algorithm for decompression.
Write-ahead Log
In BlueStore, the write-ahead log ensures Atomicity, similar to the journaling functionality of FileStore. Like FileStore, BlueStore logs all aspects of each
transaction. However, the BlueStore write-ahead log or WAL can perform this function simultaneously, which eliminates the double write penalty of FileStore.
Consequently, BlueStore is nearly twice as fast as FileStore on write operations for most workloads. BlueStore can deploy the WAL on the same device for
storing object data, or it may deploy the WAL on another device, usually when the primary block device is a hard disk drive and an SSD or NVMe will improve
performance.
Note: It is only helpful to store a block database or a write-ahead log on a separate block device if the separate device is faster than the primary storage device. For
example, SSD and NVMe devices are generally faster than HDDs. Placing the block database and the WAL on separate devices may also have performance benefits due to
differences in their workloads.
Self management operations
Ceph clusters perform a lot of self monitoring and management operations automatically. For example, Ceph OSDs can check the cluster health and report back to the
Ceph monitors. By using CRUSH to assign objects to placement groups and placement groups to a set of OSDs, Ceph OSDs can use the CRUSH algorithm to rebalance the
cluster or recover from OSD failures dynamically.
Heartbeat
Ceph OSDs join a cluster and report to Ceph Monitors on their status. At the lowest level, the Ceph OSD status is up or down reflecting whether or not it is running and able
to service Ceph client requests. If a Ceph OSD is down and in the Ceph storage cluster, this status may indicate the failure of the Ceph OSD. If a Ceph OSD is not running
for example, it crashes the Ceph OSD cannot notify the Ceph Monitor that it is down. The Ceph Monitor can ping a Ceph OSD daemon periodically to ensure that it is
running. However, heartbeating also empowers Ceph OSDs to determine if a neighboring OSD is down, to update the cluster map and to report it to the Ceph Monitors. This
means that Ceph Monitors can remain light weight processes.
Peering
Ceph stores copies of placement groups on multiple OSDs. Each copy of a placement group has a status. These OSDs peer check each other to ensure that they agree on
the status of each copy of the placement group. Peering issues usually resolve themselves.
Note: When Ceph monitors agree on the state of the OSDs storing a placement group, that does not mean that the placement group has the latest contents.
When Ceph stores a placement group in an acting set of OSDs, refer to them as Primary, Secondary, and so forth. By convention, the Primary is the first OSD in the Acting
Set. The Primary that stores the first copy of a placement group is responsible for coordinating the peering process for that placement group. The Primary is the ONLY OSD
that will accept client-initiated writes to objects for a given placement group where it acts as the Primary.
An Acting Set is a series of OSDs that are responsible for storing a placement group. An Acting Set may refer to the Ceph OSD Daemons that are currently responsible for
the placement group, or the Ceph OSD Daemons that were responsible for a particular placement group as of some epoch.
The Ceph OSD daemons that are part of an Acting Set may not always be up. When an OSD in the Acting Set is up, it is part of the Up Set. The Up Set is an important
distinction, because Ceph can remap PGs to other Ceph OSDs when an OSD fails.
Note: In an Acting Set for a PG containing osd.25, osd.32 and osd.61, the first OSD, osd.25, is the Primary. If that OSD fails, the Secondary, osd.32, becomes the
Primary, and Ceph will remove osd.25 from the Up Set.
Rebalancing and recovery
When an administrator adds a Ceph OSD to a Ceph storage cluster, Ceph updates the cluster map. This change to the cluster map also changes object placement, because
the modified cluster map changes an input for the CRUSH calculations. CRUSH places data evenly, but pseudo randomly. So only a small amount of data moves when an
IBM Storage Ceph 45
administrator adds a new OSD. The amount of data is usually the number of new OSDs divided by the total amount of data in the cluster. For example, in a cluster with 50
OSDs, 1/50th or 2% of the data might move when adding an OSD.
The following diagram depicts the rebalancing process where some, but not all of the PGs migrate from existing OSDs, OSD 1 and 2 in the diagram, to the new OSD, OSD 3,
in the diagram. Even when rebalancing, CRUSH is stable. Many of the placement groups remain in their original configuration, and each OSD gets some added capacity, so
there are no load spikes on the new OSD after the cluster rebalances.
Figure 1. Rebalancing and recovery
Data integrity
As part of maintaining data integrity, Ceph provides numerous mechanisms to guard against bad disk sectors and bit rot.
Scrubbing
Ceph OSD Daemons can scrub objects within placement groups. That is, Ceph OSD Daemons can compare object metadata in one placement group with its replicas
in placement groups stored on other OSDs. Scrubbing usually performed daily catches bugs or storage errors. Ceph OSD Daemons also perform deeper scrubbing
by comparing data in objects bit-for-bit. Deep scrubbing usually performed weekly finds bad sectors on a drive that weren’t apparent in a light scrub.
CRC Checks
When using BlueStore, Ceph can ensure data integrity by conducting a cyclical redundancy check (CRC) on write operations; then, store the CRC value in the
block database. On read operations, Ceph can retrieve the CRC value from the block database and compare it with the generated CRC of the retrieved data to
ensure data integrity instantly.
High availability
In addition to the high scalability enabled by the CRUSH algorithm, Ceph must also maintain high availability. This means that Ceph clients must be able to read and write
data even when the cluster is in a degraded state, or when a monitor fails.
Clustering the Ceph Monitor
Before Ceph clients can read or write data, they must contact a Ceph Monitor to obtain the most recent copy of the cluster map. An IBM Storage Ceph cluster can operate
with a single monitor; however, this introduces a single point of failure. That is, if the monitor goes down, Ceph clients cannot read or write data.
For added reliability and fault tolerance, Ceph supports a cluster of monitors. In a cluster of Ceph Monitors, latency and other faults can cause one or more monitors to fall
behind the current state of the cluster. For this reason, Ceph must have agreement among various monitor instances regarding the state of the storage cluster. Ceph
always uses a majority of monitors and the Paxos algorithm to establish a consensus among the monitors about the current state of the storage cluster. Ceph Monitors
nodes require NTP to prevent clock drift.
Storage administrators usually deploy Ceph with an odd number of monitors so determining a majority is efficient. For example, a majority may be 1, 2:3, 3:5, 4:6, and so
forth.
Data security and hardening
46 IBM Storage Ceph
Use this information to learn about data security and hardening information for IBM Storage Ceph Clusters and their clients. The information here also provides advice and
good practices information for hardening the security of IBM Storage Ceph, with a focus on the Ceph Orchestrator using cephadm for IBM Storage Ceph deployments.
Important: While following these instructions helps harden the security of your environment, security and compliance is not guaranteed from following these
recommendations.
Security is an important concern and should be a strong focus of any IBM Storage Ceph deployment. Data breaches and downtime are costly and difficult to manage, laws
may require passing audits and compliance processes, and projects have an expectation of a certain level of data privacy and security. The information in this section
provides a general introduction to security for IBM Storage Ceph, as well as the role of IBM in supporting your system’s security.
Introduction to IBM Storage Ceph
Supporting Software
Threat and Vulnerability Management
Encryption and Key Management
Identity and Access Management
Infrastructure Security
Data Retention
Federal Information Processing Standard (FIPS)
Summary
Introduction to IBM Storage Ceph
IBM Storage Ceph is a highly scalable and reliable object storage solution, which is typically deployed in conjunction with cloud computing solutions like OpenStack, as a
standalone storage service, or as network attached storage using interfaces.
All IBM Storage Ceph deployments consist of a storage cluster commonly referred to as the Ceph Storage Cluster or RADOS (Reliable Autonomous Distributed Object
Store), which consists of three types of daemons:
Ceph Monitors (ceph-mon): Ceph monitors provide a few critical functions such as establishing an agreement about the state of the cluster, maintaining a history
of the state of the cluster such as whether an OSD is up and running and in the cluster, providing a list of pools through which clients write and read data, and
providing authentication for clients and the Ceph Storage Cluster daemons.
Ceph Managers (ceph-mgr): Ceph manager daemons track the status of peering between copies of placement groups distributed across Ceph OSDs, a history of
the placement group states, and metrics about the Ceph cluster. They also provide interfaces for external monitoring and management systems.
Ceph OSDs (ceph-osd): Ceph Object Storage Daemons (OSDs) store and serve client data, replicate client data to secondary Ceph OSD daemons, track and report
to Ceph Monitors on their health and on the health of neighboring OSDs, dynamically recover from failures, and backfill data when the cluster size changes, among
other functions.
All IBM Storage Ceph deployments store end-user data in the Ceph Storage Cluster or RADOS (Reliable Autonomous Distributed Object Store). Generally, users DO NOT
interact with the Ceph Storage Cluster directly; rather, they interact with a Ceph client.
There are three primary Ceph Storage Cluster clients:
Ceph Object Gateway (radosgw): The Ceph Object Gateway, also known as RADOS Gateway, radosgw or rgw provides an object storage service with RESTful
APIs. Ceph Object Gateway stores data on behalf of its clients in the Ceph Storage Cluster or RADOS.
Ceph Block Device (rbd): The Ceph Block Device provides copy-on-write, thin-provisioned, and cloneable virtual block devices to a Linux kernel via Kernel RBD
(krbd) or to cloud computing solutions like OpenStack via librbd.
Ceph File System (cephfs): The Ceph File System consists of one or more Metadata Servers (mds), which store the inode portion of a file system as objects on the
Ceph Storage Cluster. Ceph file systems can be mounted via a kernel client, a FUSE client, or via the libcephfs library for cloud computing solutions like
OpenStack.
Additional clients include librados, which enables developers to create custom applications to interact with the Ceph Storage cluster and command line interface
clients for administrative purposes.
Supporting Software
An important aspect of IBM Storage Ceph security is to deliver solutions that have security built-in upfront, that IBM supports over time. Specific steps which IBM takes
with IBM Storage Ceph include:
Maintaining upstream relationships and community involvement to help focus on security from the start.
Selecting and configuring packages based on their security and performance track records.
Building binaries from associated source code (instead of simply accepting upstream builds).
Applying a suite of inspection and quality assurance tools to prevent an extensive array of potential security issues and regressions.
Digitally signing all released packages and distributing them through cryptographically authenticated distribution channels.
Providing a single, unified mechanism for distributing patches and updates.
In addition, IBM maintains a dedicated security team that analyzes threats and vulnerabilities against our products, and provides relevant advice and updates through the
Customer Portal. This team determines which issues are important, as opposed to those that are mostly theoretical problems. The IBM Product Security team maintains
expertise in, and makes extensive contributions to the upstream communities associated with our subscription products. A key part of the process, IBM Security
Advisories, deliver proactive notification of security flaws affecting IBM solutions, along with patches that are frequently distributed on the same day the vulnerability is
first published.
IBM Storage Ceph 47
Threat and Vulnerability Management
IBM Storage Ceph is typically deployed in conjunction with cloud computing solutions, so it can be helpful to think about an IBM Storage Ceph deployment abstractly as
one of many series of components in a larger deployment. These deployments typically have shared security concerns, which this guide refers to as Security Zones. Threat
actors and vectors are classified based on their motivation and access to resources. The intention is to provide you with a sense of the security concerns for each zone,
depending on your objectives.
Threat Actors
Security Zones
Connecting Security Zones
Security-Optimized Architecture
Threat Actors
A threat actor is an abstract way to refer to a class of adversary that you might attempt to defend against. The more capable the actor, the more rigorous the security
controls that are required for successful attack mitigation and prevention. Security is a matter of balancing convenience, defense, and cost, based on requirements.
In some cases, it’s impossible to secure an IBM Storage Ceph deployment against all threat actors described here. When deploying IBM Storage Ceph, you must decide
where the balance lies for your deployment and usage.
As part of your risk assessment, you must also consider the type of data you store and any accessible resources, as this will also influence certain actors. However, even if
your data is not appealing to threat actors, they could simply be attracted to your computing resources.
Nation-State Actors: This is the most capable adversary. Nation-state actors can bring tremendous resources against a target. They have capabilities beyond that
of any other actor. It’s difficult to defend against these actors without stringent controls in place, both human and technical.
Serious Organized Crime: This class describes highly capable and financially driven groups of attackers. They are able to fund in-house exploit development and
target research. In recent years, the rise of organizations such as the Russian Business Network, a massive cyber-criminal enterprise, has demonstrated how cyber
attacks have become a commodity. Industrial espionage falls within the serious organized crime group.
Highly Capable Groups: This refers to ‘Hacktivist’ type organizations who are not typically commercially funded, but can pose a serious threat to service providers
and cloud operators.
Motivated Individuals Acting Alone: These attackers come in many guises, such as rogue or malicious employees, disaffected customers, or small-scale industrial
espionage.
Script Kiddies: These attackers don’t target a specific organization, but run automated vulnerability scanning and exploitation. They are often a nuisance; however,
compromise by one of these actors is a major risk to an organization’s reputation.
The following practices can help mitigate some of the risks identified above:
Security Updates: You must consider the end-to-end security posture of your underlying physical infrastructure, including networking, storage, and server
hardware. These systems will require their own security hardening practices. For your IBM Storage Ceph deployment, you should have a plan to regularly test and
deploy security updates.
Product Updates: IBM recommends running product updates as they become available. Updates are typically released every six weeks (and occasionally more
frequently). IBM endeavors to make point releases and z-stream releases fully compatible within a major release in order to not require additional integration
testing.
Access Management: Access management includes authentication, authorization, and accounting. Authentication is the process of verifying the user’s identity.
Authorization is the process of granting permissions to an authenticated user. Accounting is the process of tracking which user performed an action. When granting
system access to users, apply the principle of least privilege, and only grant users the granular system privileges they actually need. This approach can also help
mitigate the risks of both malicious actors and typographical errors from system administrators.
Manage Insiders: You can help mitigate the threat of malicious insiders by applying careful assignment of role-based access control (minimum required access),
using encryption on internal interfaces, and using authentication/authorization security (such as centralized identity management). You can also consider additional
non-technical options, such as separation of duties and irregular job role rotation.
Security Zones
A security zone comprises users, applications, servers, or networks that share common trust requirements and expectations within a system. Typically they share the
same authentication, authorization requirements, and users. Although you might refine these zone definitions further, this section refers to four distinct security zones,
three of which form the bare minimum that is required to deploy a security-hardened IBM Storage Ceph cluster. These security zones are listed below from least to most
trusted:
Public Security Zone: The public security zone is an entirely untrusted area of the cloud infrastructure. It can refer to the Internet as a whole or simply to networks
that are external to your OpenStack deployment over which you have no authority. Any data with confidentiality or integrity requirements that traverse this zone
should be protected using compensating controls such as encryption. The public security zone SHOULD NOT be confused with the Ceph Storage Cluster’s front- or
client-side network, which is referred to as the public_network in IBM Storage Ceph and is usually NOT part of the public security zone or the Ceph client
security zone.
Ceph Client Security Zone: With IBM Storage Ceph, the Ceph client security zone refers to networks accessing Ceph clients such as Ceph Object Gateway, Ceph
Block Device, Ceph Filesystem, or librados. The Ceph client security zone is typically behind a firewall separating itself from the public security zone. However,
Ceph clients are not always protected from the public security zone. It is possible to expose the Ceph Object Gateway’s S3 and Swift APIs in the public security
zone.
48 IBM Storage Ceph
Storage Access Security Zone: The storage access security zone refers to internal networks providing Ceph clients with access to the Ceph Storage Cluster. We use
the phrase storage access security zone so that this document is consistent with the terminology used in the OpenStack Platform Security and Hardening Guide. The
storage access security zone includes the Ceph Storage Cluster’s front- or client-side network, which is referred to as the public_network in IBM Storage Ceph.
Ceph Cluster Security Zone: The Ceph cluster security zone refers to the internal networks providing the Ceph Storage Cluster’s OSD daemons with network
communications for replication, heartbeating, backfilling, and recovery. The Ceph cluster security zone includes the Ceph Storage Cluster’s backside network,
which is referred to as the cluster_network in IBM Storage Ceph.
These security zones can be mapped separately, or combined to represent the majority of the possible areas of trust within a given IBM Storage Ceph deployment.
Security zones should be mapped out against your specific deployment topology. The zones and their trust requirements will vary depending upon whether the
storage cluster is operating in a standalone capacity or is serving a public, private, or hybrid cloud.
For a visual representation of these security zones, see Security-Optimized Architecture.
Reference
For more information, see the following:
Network Communication
Connecting Security Zones
Any component that spans across multiple security zones with different trust levels or authentication requirements must be carefully configured. These connections are
often the weak points in network architecture, and should always be configured to meet the security requirements of the highest trust level of any of the zones being
connected. In many cases, the security controls of the connected zones should be a primary concern due to the likelihood of attack. The points where zones meet do
present an opportunity for attackers to migrate or target their attack to more sensitive parts of the deployment.
In some cases, IBM Storage Ceph administrators might want to consider securing integration points at a higher standard than any of the zones in which the integration
point resides. For example, the Ceph Cluster Security Zone can be isolated from other security zones easily, because there is no reason for it to connect to other security
zones. By contrast, the Storage Access Security Zone must provide access to port 6789 on Ceph monitor nodes, and ports 6800-7300 on Ceph OSD nodes. However, port
3000 should be exclusive to the Storage Access Security Zone, because it provides access to Ceph Grafana monitoring information that should be exposed to Ceph
administrators only. A Ceph Object Gateway in the Ceph Client Security Zone will need to access the Ceph Cluster Security Zone’s monitors (port 6789) and OSDs (ports
6800-7300), and may expose its S3 and Swift APIs to the Public Security Zone such as over HTTP port 80 or HTTPS port 443; yet, it may still need to restrict access to the
admin API.
The design of IBM Storage Ceph is such that the separation of security zones is difficult. As core services usually span at least two zones, special consideration must be
given when applying security controls to them.
Security-Optimized Architecture
An IBM Storage Ceph cluster’s daemons typically run on nodes that are subnet isolated and behind a firewall, which makes it relatively simple to secure a cluster.
By contrast, IBM Storage Ceph clients such as Ceph Block Device (rbd), Ceph Filesystem (cephfs), and Ceph Object Gateway (rgw) access the IBM storage cluster, but
expose their services to other cloud computing platforms.
Figure 1. Security-optimized architecture
IBM Storage Ceph 49
Encryption and Key Management
The IBM Storage Ceph cluster typically resides in its own network security zone, especially when using a private storage cluster network.
Note: Security zone separation might be insufficient for protection if an attacker gains access to Ceph clients on the public network.
There are situations where there is a security requirement to assure the confidentiality or integrity of network traffic, and where IBM Storage Ceph uses encryption and
key management, including:
SSH
SSL Termination
Messenger v2 protocol
Encryption in Transit
Compression modes of messenger v2 protocol
Encryption at Rest
Key rotation
SSH
SSL Termination
Messenger v2 protocol
Encryption in transit
Compression modes of messenger v2 protocol
Encryption at Rest
Enabling key rotation
50 IBM Storage Ceph
SSH
All nodes in the IBM Storage Ceph cluster use SSH as part of deploying the cluster. This means that on each node:
A cephadm user exists with password-less root privileges.
The SSH service is enabled and by extension port 22 is open.
A copy of the cephadm user’s public SSH key is available.
Important: Any person with access to the cephadm user by extension has permission to run commands as root on any node in the IBM Storage Ceph cluster.
For more information, see the following:
Reference
For more information, see the following:
How cephadm works
SSL Termination
The Ceph Object Gateway may be deployed in conjunction with HAProxy and keepalived for load balancing and failover. Earlier versions of Civetweb do not support SSL
and later versions support SSL with some performance limitations.
You can configure the Beast front-end web server to use the OpenSSL library to provide Transport Layer Security (TLS).
When using HAProxy and keepalived to terminate SSL connections, the HAProxy and keepalived components use encryption keys.
When using HAProxy and keepalived to terminate SSL, the connection between the load balancer and the Ceph Object Gateway is NOT encrypted.
Reference
For more information, see the following:
Configuring SSL for Beast
High availability service
Messenger v2 protocol
The second version of Ceph’s on-wire protocol, msgr2, has the following features:
A secure mode encrypting all data moving through the network.
Encapsulation improvement of authentication payloads, enabling future integration of new authentication modes.
Improvements to feature advertisement and negotiation.
The Ceph daemons bind to multiple ports allowing both the legacy v1-compatible, and the new, v2-compatible Ceph clients to connect to the same storage cluster. Ceph
clients or other Ceph daemons connecting to the Ceph Monitor daemon uses the v2 protocol first, if possible, but if not, then the legacy v1 protocol is used. By default,
both messenger protocols, v1 and v2, are enabled. The new v2 port is 3300, and the legacy v1 port is 6789, by default.
The messenger v2 protocol has two configuration options that control whether the v1 or the v2 protocol is used:
ms_bind_msgr1 - This option controls whether a daemon binds to a port speaking the v1 protocol; it is true by default.
ms_bind_msgr2 - This option controls whether a daemon binds to a port speaking the v2 protocol; it is true by default.
Similarly, two options control based on IPv4 and IPv6 addresses used:
ms_bind_ipv4 - This option controls whether a daemon binds to an IPv4 address; it is true by default.
ms_bind_ipv6 - This option controls whether a daemon binds to an IPv6 address; it is true by default.
Note: The ability to bind to multiple ports has paved the way for dual-stack IPv4 and IPv6 support.
The msgr2 protocol supports two connection modes:
crc
Provides strong initial authentication when a connection is established with cephx.
Provides a crc32c integrity check to protect against bit flips.
Does not provide protection against a malicious man-in-the-middle attack.
Does not prevent an eavesdropper from seeing all post-authentication traffic.
secure
IBM Storage Ceph 51
Provides strong initial authentication when a connection is established with cephx.
Provides full encryption of all post-authentication traffic.
Provides a cryptographic integrity check.
The default mode is crc.
Ceph Object Gateway Encryption
Also, the Ceph Object Gateway supports encryption with customer-provided keys using its S3 API.
Important: To comply with regulatory compliance standards requiring strict encryption in transit, administrators MUST deploy the Ceph Object Gateway with client-side
encryption.
Ceph Block Device Encryption
System administrators integrating Ceph as a backend for OpenStack Platform 13 MUST encrypt Ceph Block Device volumes using dm_crypt for RBD Cinder to ensure onwire encryption within the Ceph storage cluster.
Important: To comply with regulatory compliance standards requiring strict encryption in transit, system administrators MUST use dmcrypt for RBD Cinder to ensure onwire encryption within the Ceph storage cluster.
Reference
For more information, see the following:
Configuring
Encryption in transit
Starting with IBM Storage Ceph 5 and later, encryption for all Ceph traffic over the network is enabled by default, with the introduction of the messenger version 2
protocol. The secure mode setting for messenger v2 encrypts communication between Ceph daemons and Ceph clients, providing end-to-end encryption.
You can check for encryption of the messenger v2 protocol with the ceph config
dump command, netstat -Ip | grep ceph-osd command, or verify the Ceph daemon on the v2 ports.
Reference
For more information, see the following:
SSL Termination
S3 server-side encryption
Compression modes of messenger v2 protocol
The messenger v2 protocol supports the compression feature.
This feature is not enabled by default. Compressing and encrypting the same message is not recommended as the level of security of messages between peers is reduced.
If encryption is enabled, a request to enable compression is ignored until the configuration option ms_osd_compress_mode is set to true.
It supports two compression modes:
force
In multi-availability zones deployment, compresses replication messages between OSDs saves latency.
In the public cloud, minimizes message size, thereby reducing network costs to cloud provider.
Instances on public clouds with NVMe provides low network bandwidth relative to the device bandwidth. Does not provide protection against a malicious
man-in-the-middle attack.
none
The messages are transmitted without compression.
To ensure that compression of the message is enabled, run the debug_ms command and check some debug entries for connections. Also, you can run the ceph config
get command to get details about the different configuration options for the network messages.
Reference
Network configuration options
Encryption at Rest
IBM Storage Ceph supports encryption at rest in a few scenarios:
52 IBM Storage Ceph
1. Ceph Storage Cluster: The Ceph Storage Cluster supports Linux Unified Key Setup or LUKS encryption of Ceph OSDs and their corresponding journals, write-ahead
logs, and metadata databases. In this scenario, Ceph will encrypt all data at rest irrespective of whether the client is a Ceph Block Device, Ceph Filesystem, or a
custom application built on librados.
2. Ceph Object Gateway: The Ceph storage cluster supports encryption of client objects. Additionally, the data transmitted is between the Ceph Object Gateway and
the Ceph Storage Cluster is in encrypted form.
Ceph Storage Cluster Encryption
The Ceph storage cluster supports encrypting data stored in Ceph OSDs. IBM Storage Ceph can encrypt logical volumes with lvm by specifying dmcrypt; that is, lvm,
invoked by ceph-volume, encrypts an OSD’s logical volume, not its physical volume. It can encrypt non-LVM devices like partitions using the same OSD key. Encrypting
logical volumes allows for more configuration flexibility.
Ceph uses LUKS v1 rather than LUKS v2, because LUKS v1 has the broadest support among Linux distributions.
When creating an OSD, lvm will generate a secret key and pass the key to the Ceph Monitors securely in a JSON payload via stdin. The attribute name for the encryption
key is dmcrypt_key.
Important: System administrators must explicitly enable encryption.
By default, Ceph does not encrypt data stored in Ceph OSDs. System administrators must enable dmcrypt to encrypt data stored in Ceph OSDs. When using a Ceph
Orchestrator service specification file for adding Ceph OSDs to the storage cluster, set the following option in the file to encrypt Ceph OSDs:
Example
...
encrypted: true
...
Note: LUKS and dmcrypt only address encryption for data at rest, not encryption for data in transit.
Ceph Object Gateway Encryption
The Ceph Object Gateway supports encryption with customer-provided keys using its S3 API. When using customer-provided keys, the S3 client passes an encryption key
along with each request to read or write encrypted data. It is the customer’s responsibility to manage those keys. Customers must remember which key the Ceph Object
Gateway used to encrypt each object.
Reference
For more information, see the following:
S3 API server-side encryption
Enabling key rotation
Ceph and Ceph Object Gateway daemons within the Ceph cluster have a secret key. This key is used to connect to and authenticate with the cluster. You can update an
active security key within an active Ceph cluster with minimal service interruption, by using the key rotation feature.
Note: The active Ceph cluster includes nodes within the Ceph client role with parallel key changes.
Key rotation helps ensure that current industry and security compliance requirements are met.
For more information about enabling key rotation on Ceph dashboard, see Ceph client authentication keys.
Prerequisites
Be sure that you have a running IBM Storage Ceph cluster and admin user privileges.
Procedure
1. Rotate the key:
Syntax
ceph orch daemon rotate-key NAME
Example
[ceph: root@host01 /]# ceph orch daemon rotate-key mgr.ceph-key-host01
Scheduled to rotate-key mgr.ceph-key-host01 on host 'my-host-host01-installer'
2. If using a daemon other than MDS, OSD, or MGR, restart the daemon to switch to the new key. MDS, OSD, and MGR daemons do not require daemon restart.
Syntax
ceph orch restart SERVICE_TYPE
Example
[ceph: root@host01 /]# ceph orch restart rgw
Identity and Access Management
IBM Storage Ceph provides identity and access management for:
IBM Storage Ceph 53
Ceph Storage Cluster User Access
Ceph Object Gateway User Access
Ceph Object Gateway LDAP or AD authentication
Ceph Object Gateway OpenStack Keystone authentication
Ceph Storage Cluster User Access
To identify users and protect against man-in-the-middle attacks, Ceph provides its cephx authentication system to authenticate users and daemons. For more
information about cephx, see Ceph user management.
Important: The cephx protocol DOES NOT address data encryption in transport or encryption at rest.
Cephx uses shared secret keys for authentication, meaning both the client and the monitor cluster have a copy of the client’s secret key. The authentication protocol is
such that both parties are able to prove to each other they have a copy of the key without actually revealing it. This provides mutual authentication, which means the
cluster is sure the user possesses the secret key, and the user is sure that the cluster has a copy of the secret key.
In the figure below, users are either individuals or system actors such as applications, which use Ceph clients to interact with the IBM Storage Ceph cluster daemons.
Figure 1. OSD states
Ceph runs with authentication and authorization enabled by default. Ceph clients may specify a user name and a keyring containing the secret key of the specified user,
usually by using the command line. If the user and keyring are not provided as arguments, Ceph will use the client.admin administrative user as the default. If a keyring
is not specified, Ceph will look for a keyring by using the keyring setting in the Ceph configuration.
Important: To harden a Ceph cluster, keyrings SHOULD ONLY have read and write permissions for the current user and root. The keyring containing the client.admin
administrative user key must be restricted to the root user.
Reference
For more information, see the following:
Configuring
Ceph authentication configuration
Ceph Object Gateway User Access
The Ceph Object Gateway provides a RESTful application programming interface (API) service with its own user management that authenticates and authorizes users to
access S3 and Swift APIs containing user data. Authentication consists of:
S3 User: An access key and secret for a user of the S3 API.
Swift User: An access key and secret for a user of the Swift API. The Swift user is a subuser of an S3 user. Deleting the S3 parent user will delete the Swift user.
Administrative User: An access key and secret for a user of the administrative API. Administrative users should be created sparingly, as the administrative user will
be able to access the Ceph Admin API and execute its functions, such as creating users, and giving them permissions to access buckets or containers and their
objects among other things.
The Ceph Object Gateway stores all user authentication information in Ceph Storage cluster pools. Additional information may be stored about users including names,
email addresses, quotas, and usage.
Reference
For more information, see the following:
User management
Creating an administrative user
Ceph Object Gateway LDAP or AD authentication
IBM Storage Ceph supports Light-weight Directory Access Protocol (LDAP) servers for authenticating Ceph Object Gateway users. When configured to use LDAP or Active
Directory (AD), Ceph Object Gateway defers to an LDAP server to authenticate users of the Ceph Object Gateway.
Ceph Object Gateway controls whether to use LDAP. However, once configured, it is the LDAP server that is responsible for authenticating users.
54 IBM Storage Ceph
To secure communications between the Ceph Object Gateway and the LDAP server, IBM recommends deploying configurations with LDAP Secure or LDAPS.
Important: When using LDAP, ensure that access to the rgw_ldap_secret =
_PATH_TO_SECRET_FILE_ secret file is secure.
Reference
For more information, see the following:
Configuring LDAP and Ceph Object Gateway
Configuring Active Directory and Ceph Object Gateway
Ceph Object Gateway OpenStack Keystone authentication
IBM Storage Ceph supports using OpenStack Keystone to authenticate Ceph Object Gateway Swift API users. The Ceph Object Gateway can accept a Keystone token,
authenticate the user and create a corresponding Ceph Object Gateway user. When Keystone validates a token, the Ceph Object Gateway considers the user
authenticated.
Ceph Object Gateway controls whether to use OpenStack Keystone for authentication. However, once configured, it is the OpenStack Keystone service that is responsible
for authenticating users.
Configuring the Ceph Object Gateway to work with Keystone requires converting the OpenSSL certificates that Keystone uses for creating the requests to the nss db
format.
Reference
For more information, see the following:
Ceph Object Gateway and OpenStack Keystone
Infrastructure Security
The scope of this guide is IBM Storage Ceph. However, a proper IBM Storage Ceph security plan requires consideration of the following prerequisites.
Prerequisites
Red Hat Enterprise Linux 9 Security Hardening Guide
Red Hat Enterprise Linux 9 Using SELinux Guide
Administration
Network Communication
Hardening the Network Service
Reporting
Auditing Administrator Actions
Administration
Administering an IBM Storage Ceph cluster involves using command line tools. The CLI tools require an administrator key for administrator access privileges to the cluster.
By default, Ceph stores the administrator key in the /etc/ceph directory. The default file name is ceph.client.admin.keyring. Take steps to secure the keyring so
that only a user with administrative privileges to the cluster may access the keyring.
Network Communication
IBM Storage Ceph provides two networks:
A public network.
A cluster network.
All Ceph daemons and Ceph clients require access to the public network, which is part of the storage access security zone. By contrast, ONLY the OSD daemons require
access to the cluster network, which is part of the Ceph cluster security zone.
Figure 1. Network architecture
IBM Storage Ceph 55
The Ceph configuration contains public_network and cluster_network settings. For hardening purposes, specify the IP address and the netmask using CIDR
notation. Specify multiple comma-delimited IP address and netmask entries if the cluster will have multiple sub-nets.
public_network = <public-network/netmask>[,<public-network/netmask>]
cluster_network = <cluster-network/netmask>[,<cluster-network/netmask>]
Reference
For more information, see the following:
Ceph network configuration
Hardening the Network Service
System administrators deploy IBM Storage Ceph clusters on Red Hat Enterprise Linux 9 Server. SELinux is on by default and the firewall blocks all inbound traffic except
for the SSH service port 22; however, you MUST ensure that this is the case so that no other unauthorized ports are open or unnecessary services are enabled.
On each server node, execute the following:
1. Start the firewalld service, enable it to run on boot, and ensure that it is running:
# systemctl enable firewalld
# systemctl start firewalld
# systemctl status firewalld
2. Take an inventory of all open ports.
# firewall-cmd --list-all
On a new installation, the sources: section should be blank indicating that no ports have been opened specifically. The services section should indicate ssh
indicating that the SSH service (and port 22) and dhcpv6-client are enabled.
sources:
services: ssh dhcpv6-client
56 IBM Storage Ceph
3. Ensure SELinux is running and Enforcing.
# getenforce
Enforcing
If SELinux is Permissive, set it to Enforcing.
# setenforce 1
If SELinux is not running, enable it. For more information, see Red Hat Enterprise Linux 9 Using SELinux.
Each Ceph daemon uses one or more ports to communicate with other daemons in the IBM Storage Ceph cluster. In some cases, you may change the default port
settings. Administrators typically only change the default port with the Ceph Object Gateway or ceph-radosgw daemon.
Table 1. Ceph Ports
TCP/UDP Port
6789, 3300
Daemon
ceph-mon
N/A
Configuration Option
6800-7300
ceph-osd
ms_bind_port_min to ms_bind_port_max
6800-7300
ceph-mgr
ms_bind_port_min to ms_bind_port_max
6800
ceph-mds
N/A
ceph-radosgw rgw_frontends
8080
The Ceph Storage Cluster daemons include ceph-mon, ceph-mgr, and ceph-osd. These daemons and their hosts comprise the Ceph cluster security zone, which should
use its own subnet for hardening purposes.
The Ceph clients include ceph-radosgw, ceph-mds, ceph-fuse, libcephfs, rbd, librbd, and librados. These daemons and their hosts comprise the storage
access security zone, which should use its own subnet for hardening purposes.
On the Ceph Storage Cluster zone’s hosts, consider enabling only hosts running Ceph clients to connect to the Ceph Storage Cluster daemons. For example:
firewall-cmd --zone=<zone-name> --add-rich-rule="rule family="ipv4" \
source address="<ip-address>/<netmask>" port protocol="tcp" \
port="<port-number>" accept"
Replace <zone-name> with the zone name, <ipaddress> with the IP address, <netmask> with the subnet mask in CIDR notation, and <port-number> with the port
number or range. Repeat the process with the --permanent flag so that the changes persist after reboot. For example:
firewall-cmd --zone=<zone-name> --add-rich-rule="rule family="ipv4" \
source address="<ip-address>/<netmask>" port protocol="tcp" \
port="<port-number>" accept" --permanent
Reporting
IBM Storage Ceph provides basic system monitoring and reporting with the ceph-mgr daemon plug-ins, namely, the RESTful API, the dashboard, and other plug-ins such
as Prometheus and Zabbix. Ceph collects this information using collectd and sockets to retrieve settings, configuration details, and statistical information.
In addition to default system behavior, system administrators may configure collectd to report on security matters, such as configuring the IP-Tables or ConnTrack
plug-ins to track open ports and connections respectively.
System administrators may also retrieve configuration settings at runtime.
Reference
For more information, see the following:
Viewing the Ceph configuration at runtime
Auditing Administrator Actions
An important aspect of system security is to periodically audit administrator actions on the cluster. IBM Storage Ceph stores a history of administrator actions in the
/var/log/ceph/CLUSTER_FSID/ceph.audit.log file.
Example
[root@host04 ~]# cat /var/log/ceph/6c58dfb8-4342-11ee-a953-fa163e843234/ceph.audit.log
Each entry will contain:
Timestamp: Indicates when the command was executed.
Monitor Address: Identifies the monitor modified.
Client Node: Identifies the client node initiating the change.
Entity: Identifies the user making the change.
Command: Identifies the command executed.
The following is an output of the Ceph audit log:
IBM Storage Ceph 57
2023-09-01T10:20:21.445990+0000 mon.host01 (mon.0) 122301 : audit [DBG] from='mgr.14189 10.0.210.22:0/1157748332'
entity='mgr.host01.mcadea' cmd=[{"prefix": "config generate-minimal-conf"}]: dispatch
2023-09-01T10:20:21.446972+0000 mon.host01 (mon.0) 122302 : audit [INF] from='mgr.14189 10.0.210.22:0/1157748332'
entity='mgr.host01.mcadea' cmd=[{"prefix": "auth get", "entity": "client.admin"}]: dispatch
2023-09-01T10:20:21.453790+0000 mon.host01 (mon.0) 122303 : audit [INF] from='mgr.14189 10.0.210.22:0/1157748332'
entity='mgr.host01.mcadea'
2023-09-01T10:20:21.457119+0000 mon.host01 (mon.0) 122304 : audit [DBG] from='mgr.14189 10.0.210.22:0/1157748332'
entity='mgr.host01.mcadea' cmd=[{"prefix": "osd tree", "states": ["destroyed"], "format": "json"}]: dispatch
2023-09-01T10:20:30.671816+0000 mon.host01 (mon.0) 122305 : audit [DBG] from='mgr.14189 10.0.210.22:0/1157748332'
entity='mgr.host01.mcadea' cmd=[{"prefix": "osd blocklist ls", "format": "json"}]: dispatch
In distributed systems such as Ceph, actions may begin on one instance and get propagated to other nodes in the cluster. When the action begins, the log indicates
dispatch. When the action ends, the log indicates finished.
Data Retention
IBM Storage Ceph stores user data, but usually in an indirect manner. Customer data retention may involve other applications, such as the OpenStack Platform.
Ceph Storage Cluster
Ceph Block Device
Ceph File System
Ceph Object Gateway
Ceph Storage Cluster
The Ceph Storage Cluster, often referred to as the Reliable Autonomic Distributed Object Store or RADOS, stores data as objects within pools. In most cases, these objects
are the atomic units representing client data, such as Ceph Block Device images, Ceph Object Gateway objects, or Ceph Filesystem files. However, custom applications
built on top of librados may bind to a pool and store data too.
Cephx controls access to the pools storing object data. However, Ceph Storage Cluster users are typically Ceph clients, and not users. Consequently, users generally DO
NOT have the ability to write, read or delete objects directly in a Ceph Storage Cluster pool.
Ceph Block Device
The most popular use of IBM Storage Ceph, the Ceph Block Device interface, also referred to as RADOS Block Device or RBD, creates virtual volumes, images, and
compute instances and stores them as a series of objects within pools. Ceph assigns these objects to placement groups and distributes or places them pseudo-randomly
in OSDs throughout the cluster.
Depending upon the application consuming the Ceph Block Device interface, usually OpenStack Platform, users may create, modify, and delete volumes and images. Ceph
handles the create, retrieve, update, and delete operations of each individual object.
Deleting volumes and images destroys the corresponding objects in an unrecoverable manner. However, residual data artifacts may continue to reside on storage media
until overwritten. Data may also remain in backup archives.
Ceph File System
The Ceph File System interface creates virtual file systems and stores them as a series of objects within pools. Ceph assigns these objects to placement groups and
distributes or places them pseudo-randomly in OSDs throughout the cluster.
Typically, the Ceph File System uses two pools:
Metadata: The metadata pool stores the data of the Ceph Metadata Server (MDS), which generally consists of inodes; that is, the file ownership, permissions,
creation date and time, last modified or accessed date and time, parent directory, and so on.
Data: The data pool stores file data. Ceph may store a file as one or more objects, typically representing smaller chunks of file data such as extents.
Depending upon the application consuming the Ceph File System interface, usually OpenStack Platform, users may create, modify, and delete files in a Ceph File System.
Ceph handles the create, retrieve, update, and delete operations of each individual object representing the file.
Deleting files destroys the corresponding objects in an unrecoverable manner. However, residual data artifacts may continue to reside on storage media until overwritten.
Data may also remain in backup archives.
Ceph Object Gateway
From a data security and retention perspective, the Ceph Object Gateway interface has some important differences when compared to the Ceph Block Device and Ceph
Filesystem interfaces. The Ceph Object Gateway provides a service to users. The Ceph Object Gateway may store:
User Authentication Information: User authentication information generally consists of user IDs, user access keys, and user secrets. It may also comprise a user’s
name and email address if provided. Ceph Object Gateway will retain user authentication data unless the user is explicitly deleted from the system.
58 IBM Storage Ceph
User Data: User data generally comprises user- or administrator-created buckets or containers, and the user-created S3 or Swift objects contained within them.
The Ceph Object Gateway interface creates one or more Ceph Storage cluster objects for each S3 or Swift object and stores the corresponding Ceph Storage cluster
objects within a data pool. Ceph assigns the Ceph Storage cluster objects to placement groups and distributes or places them pseudo-randomly in OSDs throughout
the cluster. The Ceph Object Gateway may also store an index of the objects contained within a bucket or index to enable services such as listing the contents of an
S3 bucket or Swift container. Additionally, when implementing multi-part uploads, the Ceph Object Gateway may temporarily store partial uploads of S3 or Swift
objects.
Users may create, modify, and delete buckets or containers, and the objects contained within them in a Ceph Object Gateway. Ceph handles the create, retrieve,
update, and delete operations of each individual Ceph Storage cluster object representing the S3 or Swift object.
Deleting S3 or Swift objects destroys the corresponding Ceph Storage cluster objects in an unrecoverable manner. However, residual data artifacts may continue to
reside on storage media until overwritten. Data may also remain in backup archives.
Logging: Ceph Object Gateway also stores logs of user operations that the user intends to accomplish and operations that have been executed. This data provides
traceability about who created, modified or deleted a bucket or container, or an S3 or Swift object residing in an S3 bucket or Swift container. When users delete
their data, the logging information is not affected and will remain in storage until deleted by a system administrator or removed automatically by expiration policy.
Bucket Lifecycle
Ceph Object Gateway also supports bucket lifecycle features, including object expiration. Data retention regulations like the General Data Protection Regulation may
require administrators to set object expiration policies and disclose them to users among other compliance factors.
multi-site
Ceph Object Gateway is often deployed in a multi-site context whereby a user stores an object at one site and the Ceph Object Gateway creates a replica of the object in
another cluster possibly at another geographic location. For example, if a primary cluster fails, a secondary cluster may resume operations. In another example, a
secondary cluster may be in a different geographic location, such as an edge network or content-delivery network such that a client may access the closest cluster to
improve response time, throughput, and other performance characteristics. In multi-site scenarios, administrators must ensure that each site has implemented security
measures. Additionally, if geographic distribution of data would occur in a multi-site scenario, administrators must be aware of any regulatory implications when the data
crosses political boundaries.
Federal Information Processing Standard (FIPS)
IBM Storage Ceph uses FIPS validated cryptography modules when run on the latest certified Red Hat Enterprise Linux version.
Enable FIPS mode on Red Hat Enterprise Linux either during system installation or after it.
For container deployments, follow the instructions in the Red Hat Enterprise Linux 9 Security Hardening Guide on the Red Hat Customer Portal.
Reference
For more information, see the following:
US Government Standards
Summary
This document has provided only a general introduction to security for IBM Storage Ceph. Contact IBM Support for additional help.
Planning
Planning involves considering the supported compatibility, physical configuration, and various storage strategy prerequisites before working with IBM Storage Ceph.
Considerations and recommendations
Use this information to understand all considerations and recommendations before you run your IBM Storage Ceph cluster.
Hardware
Storage Strategies
Considerations and recommendations
Use this information to understand all considerations and recommendations before you run your IBM Storage Ceph cluster.
IBM Storage Ceph can be used for different workloads based on a particular business need or set of requirements. Doing the necessary planning before installing an IBM
Storage Ceph is critical to the success of running a Ceph storage cluster efficiently and achieving the business requirements.
Note: For more IBM Storage Ceph planning help, contact your IBM representative.
Basic considerations
Use these basic considerations for running your IBM Storage Ceph cluster.
Workload considerations
One of the key benefits of a Ceph storage cluster is the ability to support different types of workloads within the same storage cluster using performance domains.
Different hardware configurations can be associated with each performance domain. Storage administrators can deploy storage pools on the appropriate
IBM Storage Ceph 59
performance domain, providing applications with storage tailored to specific performance and cost profiles. Selecting the correct sized and optimized servers for
these performance domains is an essential aspect of designing an IBM Storage Ceph cluster.
Network considerations for IBM Storage Ceph
An important aspect of a cloud storage solution is that storage clusters can run out of IOPS due to network latency, and other factors. The storage cluster can run
out of throughput due to bandwidth constraints long before the storage clusters run out of storage capacity. As a result, the network hardware configuration must
support the chosen workloads to meet price versus performance requirements.
Considerations for using a RAID controller with OSD hosts
Use this information when planning for RAID controller usage with OSD hosts.
Tuning considerations for the Linux kernel when running Ceph
Colocation
Use this information to understand how colocation works and its advantages.
Operating system requirements
Before running IBM Storage Ceph, be sure to comply with all operating system requirements listed here.
Accessing Red Hat entitlements from IBM Storage Ceph
IBM Storage Ceph can include entitlement to use Red Hat® OpenShift® Container Platform, Red Hat Enterprise Linux® CoreOS (RHCOS), and Red Hat Enterprise
Linux (RHEL). To access these entitlements, you must link your IBM Storage Ceph to your Red Hat account. You can link your IBM Storage Ceph to its Red Hat
entitlement through IBM Passport Advantage.
Minimum hardware considerations
Before running IBM Storage Ceph, be sure to comply with all minimum hardware requirements listed here.
Basic considerations
Use these basic considerations for running your IBM Storage Ceph cluster.
Before using IBM Storage Ceph, it is important to develop a proper storage strategy for your data. A storage strategy is a method of storing data that serves a particular
use case. If storing volumes and images for a cloud platform, such as OpenStack, is necessary, you can store data on faster Serial Attached SCSI (SAS) drives with Solid
State Drives (SSD) for journals. By contrast, if you need to store object data for an S3- or Swift-compliant gateway, you can choose to use something more economical, like
traditional Serial Advanced Technology Attachment (SATA) drives.
Both cloud platform and S3- and Swift-compliant gateway scenarios are supported within the same Ceph storage cluster. When a storage cluster needs to support both
scenarios, a fast storage strategy to the cloud platform and traditional storage for the object store must be provided.
One of the most important steps in a successful Ceph deployment is identifying a price-to-performance profile suitable for the storage cluster’s use case and workload. It
is important to choose the right hardware for the use case. For example:
An IOPS-optimized hardware for a cold storage application can increase hardware costs unnecessarily.
Capacity-optimitized hardware might have a more attractive price point but can lead to slow performance with IOPS-intensive workloads.
For more hardware considerations, see Minimum hardware considerations.
IBM Storage Ceph use cases can support multiple storage strategies. Use cases, cost versus benefit performance tradeoffs, and data durability are the primary
considerations that help develop a sound storage strategy.
Use cases
Ceph provides massive storage capacity, and it supports numerous use cases, such as:
The Ceph Block Device client is a leading storage backend for cloud platforms that provides limitless storage for volumes and images with high-performance
features like copy-on-write cloning.
The Ceph Object Gateway client is a leading storage backend for cloud platforms that provides a RESTful S3-compliant and Swift-compliant object storage for
objects like audio, bitmap, video, and other data.
The Ceph File System for traditional file storage.
Cost versus benefit of performance
Faster is better. Bigger is better. High durability is better. However, there is a price for each superlative quality, and a corresponding cost versus benefit tradeoff. Consider
the following use cases from a performance perspective:
SSDs can provide fast storage for relatively small amounts of data and journaling. Storing a database or object index can benefit from a pool of very fast SSDs, but
proves too expensive for other data.
SAS drives with SSD journaling provide fast performance at an economical price for volumes and images.
SATA drives without SSD journaling provide cheap storage with lower overall performance.
When you create a CRUSH hierarchy of OSDs, you need to consider the use case and an acceptable cost versus performance tradeoff.
Data durability
In large-scale storage clusters, hardware failure is an expectation, not an exception. However, data loss and service interruption remain unacceptable. For this reason,
data durability is very important. Ceph addresses data durability with multiple replica copies of an object or with erasure coding and multiple coding chunks. Multiple
copies or multiple coding chunks present an extra cost versus benefit tradeoff: it is cheaper to store fewer copies or coding chunks, but it can lead to the inability to
service write requests in a degraded state. Generally, one object with two more copies, or two coding chunks can allow a storage cluster to service writes in a degraded
state while the storage cluster recovers.
Replication stores one or more redundant copies of the data across failure domains in case of a hardware failure. However, redundant copies of data can become
expensive at scale. For example, to store 1 petabyte of data with triple replication would require a cluster with at least 3 petabytes of storage capacity.
Erasure coding stores data as data chunks and coding chunks. If there is a lost data chunk, erasure coding can recover the lost data chunk with the remaining data chunks
and coding chunks. Erasure coding is substantially more economical than replication. For example, by using erasure coding with 8 data chunks and 3 coding chunks
60 IBM Storage Ceph
provides the same redundancy as 3 copies of the data. However, such an encoding scheme uses approximately 1.5x the initial data stored compared to 3x with
replication.
The CRUSH algorithm aids this process by ensuring that Ceph stores extra copies or coding chunks in different locations within the storage cluster. This ensures that the
failure of a single storage device or host does not lead to losing all copies or coding chunks necessary to preclude data loss. You can plan a storage strategy with cost
versus benefit tradeoffs, and data durability in mind, then present it to a Ceph client as a storage pool.
Important:
ONLY the data storage pool can use erasure coding. Pools storing service data and bucket indexes use replication.
Ceph’s object copies or coding chunks make RAID solutions obsolete. Do not use RAID, as Ceph already handles data durability. A degraded RAID has a negative
impact on performance and recovering data that uses RAID is substantially slower than using deep copies or erasure coding chunks.
Workload considerations
One of the key benefits of a Ceph storage cluster is the ability to support different types of workloads within the same storage cluster using performance domains.
Different hardware configurations can be associated with each performance domain. Storage administrators can deploy storage pools on the appropriate performance
domain, providing applications with storage tailored to specific performance and cost profiles. Selecting the correct sized and optimized servers for these performance
domains is an essential aspect of designing an IBM Storage Ceph cluster.
To the Ceph client interface that reads and writes data, a Ceph storage cluster appears as a simple pool where the client stores data. However, the storage cluster
performs many complex operations in a manner that is visible to the client interface. Ceph clients and Ceph object storage daemons, referred to as Ceph OSDs, or simply
OSDs, both use the Controlled Replication Under Scalable Hashing (CRUSH) algorithm for the storage and retrieval of objects. Ceph OSDs can run in containers within the
storage cluster.
A CRUSH map describes a topography of cluster resources, and the map exists both on client hosts as well as Ceph Monitor hosts within the cluster. Ceph clients and Ceph
OSDs both use the CRUSH map and the CRUSH algorithm. Ceph clients communicate directly with OSDs, eliminating a centralized object lookup and a potential
performance bottleneck. With awareness of the CRUSH map and communication with their peers, OSDs can handle replication, backfilling, and recovery—allowing for
dynamic failure recovery.
Ceph uses the CRUSH map to implement failure domains. Ceph also uses the CRUSH map to implement performance domains, which take the performance profile of the
underlying hardware into consideration. The CRUSH map describes how Ceph stores data, and it is implemented as a simple hierarchy, specifically an acyclic graph, and a
ruleset. The CRUSH map can support multiple hierarchies to separate one type of hardware performance profile from another. Ceph implements performance domains
with device "classes".
For example, you can have these performance domains coexisting in the same IBM Storage Ceph cluster:
Hard disk drives (HDDs) are typically appropriate for cost and capacity-focused workloads.
Throughput-sensitive workloads typically use HDDs with Ceph write journals on solid-state drives (SSDs).
IOPS-intensive workloads, such as MySQL and MariaDB, often use SSDs.
Figure 1. Performance and Failure Domains
IBM Storage Ceph 61
Workloads
IBM Storage Ceph is optimized for three primary workloads:
Important: Carefully consider the workload being run by IBM Storage Ceph clusters BEFORE considering what hardware to purchase because it can significantly impact
the price and performance of the storage cluster. For example, if the workload is capacity-optimized and the hardware is better suited to a throughput-optimized
workload, then hardware will be more expensive than necessary. Conversely, if the workload is throughput-optimized and the hardware is better suited to a capacityoptimized workload, then the storage cluster can suffer from poor performance.
IOPS optimized
Input, output per second (IOPS) optimization deployments are suitable for cloud computing operations, such as running MYSQL or MariaDB instances as virtual
machines on OpenStack. IOPS optimized deployments require higher performance storage such as 15k RPM SAS drives and separate SSD journals to handle
frequent write operations. Some high IOPS scenarios use all flash storage to improve IOPS and total throughput.
An IOPS-optimized storage cluster has the following properties:
Lowest cost per IOPS.
Highest IOPS per GB.
99th percentile latency consistency.
Uses for an IOPS-optimized storage cluster are:
Typically block storage.
3x replication for hard disk drives (HDDs) or 2x replication for solid-state drives (SSDs).
MySQL on OpenStack clouds.
Throughput optimized
Throughput-optimized deployments are suitable for serving up significant amounts of data, such as graphic, audio, and video content. Throughput-optimized
deployments require high-bandwidth networking hardware, controllers, and hard disk drives with fast sequential read and write characteristics. If fast data access
is a requirement, then use a throughput-optimized storage strategy. Also, if fast write performance is a requirement, using SSDs for journals can substantially
improve write performance.
A throughput-optimized storage cluster has the following properties:
Lowest cost per MBps (throughput).
Highest MBps per TB.
Highest MBps per BTU.
62 IBM Storage Ceph
Highest MBps per Watt.
97th percentile latency consistency.
Uses for a throughput-optimized storage cluster are:
Block or object storage.
3x replication.
Active performance storage for video, audio, and images.
Streaming media, such as 4k video.
Capacity optimized
Capacity-optimized deployments are suitable for storing significant amounts of data as inexpensively as possible. Capacity-optimized deployments typically trade
performance for a more attractive price point. For example, capacity-optimized deployments often use slower and less expensive SATA drives and colocate journals
rather than using SSDs for journaling.
A cost and capacity-optimized storage cluster has the following properties:
Lowest cost per TB.
Lowest BTU per TB.
Lowest Watts required per TB.
A cost and capacity-optimized storage cluster has the following properties:
Typically object storage.
Erasure coding for maximizing usable capacity
Object archive.
Video, audio, and image object repositories.
Network considerations for IBM Storage Ceph
An important aspect of a cloud storage solution is that storage clusters can run out of IOPS due to network latency, and other factors. The storage cluster can run out of
throughput due to bandwidth constraints long before the storage clusters run out of storage capacity. As a result, the network hardware configuration must support the
chosen workloads to meet price versus performance requirements.
Storage administrators prefer that a storage cluster recovers as quickly as possible. Carefully consider bandwidth requirements for the storage cluster network, be mindful
of network link oversubscription, and separate the intra-cluster traffic from the client-to-cluster traffic. Network performance is increasingly important when considering
the use of Solid State Disks (SSD), flash, NVMe, and other high performing storage devices.
Ceph supports a public network and a storage cluster network. The public network handles client traffic and communication with Ceph Monitors. The storage cluster
network handles Ceph OSD heartbeats, replication, backfilling, and recovery traffic. Use a minimum of a single 10 Gb/s Ethernet link for storage hardware, and another 10
Gb/s Ethernet links can be added for connectivity and throughput.
Important:
Allocate bandwidth to the storage cluster network, such that it is a multiple of the public network by using the osd_pool_default_size parameter as the basis for the
multiple on replicated pools. Run the public and storage cluster networks on separate network cards.
Use 10 Gb/s Ethernet for IBM Storage Ceph deployments in production. A 1 Gb/s Ethernet network is not suitable for production storage clusters.
In the case of a drive failure, replicating 1 TB of data across a 1 Gb/s network takes 3 hours and replicating 10 TB across a 1 Gb/s network takes 30 hours. Using 10 TB is
the typical drive configuration. By contrast, with a 10 Gb/s Ethernet network, the replication times would be 20 minutes for 1 TB and 1 hour for 10 TB.
Note: When a Ceph OSD fails, the storage cluster recovers by replicating the data that it contained to other Ceph OSDs within the pool.
The failure of a larger domain such as a rack means that the storage cluster uses considerably more bandwidth. When building a storage cluster consisting of multiple
racks, which is common for large storage implementations, consider using as much network bandwidth between switches in a "fat tree" design for optimal performance. A
typical 10 Gb/s Ethernet switch has 48 10 Gb/s ports and four 40 Gb/s ports. Use the 40 Gb/s ports on the spine for maximum throughput. Alternatively, consider
aggregating unused 10 Gb/s ports with QSFP+ and SFP+ cables into more 40 Gb/s ports to connect to other rack and spine routers. LACP mode 4 can be used to bond
network interfaces. Use jumbo frames with a maximum transmission unit (MTU) of 9000, especially on the backend or cluster network.
Before installing and testing an IBM Storage Ceph cluster, verify the network throughput. Most performance-related problems in Ceph usually begin with a networking
issue. Simple network issues like a kinked or bent Cat-6 cable could result in degraded bandwidth. Use a minimum of 10 Gb/s ethernet for the front side network. For large
clusters, consider using 40 Gb/s ethernet for the backend or cluster network.
Important: For network optimization, use jumbo frames for a better CPU per bandwidth ratio, and a non-blocking network switch back-plane. IBM Storage Ceph requires
the same MTU value throughout all networking devices in the communication path, end-to-end for both public and cluster networks. Verify that the MTU value is the same
on all hosts and networking equipment in the environment before using an IBM Storage Ceph cluster in production.
Reference
Configuring multiple public networks to the cluster
Configuring a private network
Configuring a public network
Considerations for using a RAID controller with OSD hosts
Use this information when planning for RAID controller usage with OSD hosts.
IBM Storage Ceph 63
If an OSD host has a RAID controller with 1 - 2 Gb of cache installed, enabling the write-back cache might result in increased small I/O write throughput. However,
the cache must be nonvolatile.
Most modern RAID controllers have super capacitors that provide enough power to drain volatile memory to nonvolatile NAND memory during a power-loss event.
It is important to understand how a particular controller and its firmware behave after power is restored.
Some RAID controllers require manual intervention. Hard disks typically advertise to the operating system whether their disk caches can be enabled or disabled by
default. However, certain RAID controllers and some firmware do not provide such information. Verify that disk level caches are disabled to avoid file system
corruption.
Create a single RAID 0 volume with write-back for each Ceph OSD data drive with write-back cache enabled.
If Serial Attached SCSI (SAS) or SATA connected Solid-state Drive (SSD) disks are also present on the RAID controller, then investigate whether the controller and
firmware support pass-through mode. Enabling pass-through mode helps avoid caching logic, and generally results in lower latency for fast media.
Tuning considerations for the Linux kernel when running Ceph
Production IBM Storage Ceph clusters generally benefit from tuning the operating system, specifically around limits and memory allocation. Ensure that adjustments are
set for all hosts within the storage cluster. For more information, contact IBM Support.
Increase the file descriptors
The Ceph Object Gateway can hang if it runs out of file descriptors. You can modify the /etc/security/limits.conf file on Ceph Object Gateway hosts to increase the file
descriptors for the Ceph Object Gateway.
ceph
soft
nofile
unlimited
Adjusting the ulimit value for large storage clusters
When running Ceph administrative commands on large storage clusters. For example, with 1024 Ceph OSDs or more, create a /etc/security/limits.d/50-ceph.conf file on
each host that runs administrative commands.
Include the following contents:
USER_NAME
soft
nproc
unlimited
Replace USER_NAME with the name of the non-root user account that runs the Ceph administrative commands.
Note: The root user’s ulimit value is already set to unlimited by default on Red Hat Enterprise Linux.
Colocation
Use this information to understand how colocation works and its advantages.
You can colocate containerized Ceph daemons on the same host.
The following are the advantages of colocating some of Ceph’s services:
Significant improvement in total cost of ownership (TCO) at small scale
Reduction from six hosts to three for the minimum configuration
Easier upgrade
Better resource isolation
How colocation works
With the help of the Cephadm orchestrator, you can colocate one daemon from the following list with one or more OSD daemons (ceph-osd):
Ceph Monitor (ceph-mon) and Ceph Manager (ceph-mgr) daemons
NFS Ganesha (nfs-ganesha) for Ceph Object Gateway (nfs-ganesha)
RBD Mirror (rbd-mirror)
Observability Stack (Grafana)
Additionally, for Ceph Object Gateway (radosgw) and Ceph File System (ceph-mds), you can colocate either with an OSD daemon plus a daemon from the previous list,
excluding RBD mirror.
Note: Collocating two of the same kind of daemons on a given node is not supported.
Note: Because ceph-mon and ceph-mgr work together closely they do not count as two separate daemons for the purposes of colocation.
Note: IBM recommends colocating the Ceph Object Gateway with Ceph OSD containers to increase performance.
With the colocation rules shared above, we have the following minimum clusters sizes that comply with these rules:
Example 1
1. Media: Full flash systems (SSDs)
64 IBM Storage Ceph
2. Use case: Block (RBD) and File (CephFS), or Object (Ceph Object Gateway)
3. Number of nodes: 3
4. Replication scheme: 2
Table 1. Colocated Daemons Example 1
Host
Daemon
Daemon
Daemon
host1
OSD
Monitor/Manager Grafana
host2
OSD
Monitor/Manager RGW or CephFS
host3
OSD
Monitor/Manager RGW or CaphFS
Note: The minimum size for a storage cluster with three replicas is four nodes. Similarly, the size of a storage cluster with two replicas is a three node cluster. It is a
requirement to have a certain number of nodes for the replication factor with an extra node in the cluster to avoid extended periods with the cluster in a degraded state.
Figure 1. Colocated Daemons Example 1
Example 2
1. Media: Full flash systems (SSDs) or spinning devices (HDDs)
2. Use case: Block (RBD), File (CephFS), and Object (Ceph Object Gateway)
3. Number of nodes: 4
4. Replication scheme: 3
Table 2. Colocated Daemons Example 2
Host
Daemon
Daemon
Daemon
host1
OSD
Grafana
CephFS
host2
OSD
Monitor/Manager RGW
host3
OSD
Monitor/Manager RGW
host4
OSD
Monitor/Manager CephFS
Figure 2. Colocated Daemons Example 2
IBM Storage Ceph 65
Example 3
1. Media: Full flash systems (SSDs) or spinning devices (HDDs)
2. Use case: Block (RBD), Object (Ceph Object Gateway), and NFS for Ceph Object Gateway
3. Number of nodes: 4
4. Replication scheme: 3
Table 3. Colocated Daemons Example 3
Host
Daemon
Daemon
Daemon
host1
OSD
Grafana
host2
OSD
Monitor/Manager RGW
host3
OSD
Monitor/Manager RGW
host4
OSD
Monitor/Manager NFS (RGW)
Figure 3. Colocated Daemons Example 3
66 IBM Storage Ceph
The diagrams below shows the differences between storage clusters with colocated and non-colocated daemons.
Figure 4. Colocated Daemons
IBM Storage Ceph 67
Figure 5. Non-colocated Daemons
68 IBM Storage Ceph
Operating system requirements
Before running IBM Storage Ceph, be sure to comply with all operating system requirements listed here.
For the latest supported Red Hat Enterprise Linux versions, see Compatibility matrix.
IBM Storage Ceph 6.1 clusters support kernel clients from any of the supported RHEL versions, but RHEL 8.x clients must install IBM Storage Ceph 6.x user-end
packages.
Use the same architecture and deployment type across all nodes.
For example, do not use a mixture of nodes with both AMD64 and Intel 64 architectures or a mixture of nodes with container-based deployments.
Important: IBM Storage Ceph does not support clusters with heterogeneous architectures or deployment types.
By default, SELinux is set to Enforcing mode and the ceph-selinux packages are installed. For details, see Using SELinux for your OS version, on the Red Hat
Customer Portal.
Red Hat Enterprise Linux (RHEL) entitlements
Be aware of the following Red Hat Enterprise Linux entitlement requirements for your IBM Storage Ceph edition:
IBM Storage Ceph Premium Edition
Includes a Red Hat Enterprise operating system entitlement and Red Hat Enterprise Premium Service Level Agreement.
IBM Storage Ceph 69
To access your entitlements, link your IBM Storage Ceph to your Red Hat account. You can link your IBM Storage Ceph to its Red Hat entitlement through
IBM Passport Advantage. See Accessing Red Hat entitlements from IBM Storage Ceph for instructions on how to link your entitlements.
IBM Storage Ceph Pro Edition
Does NOT include a Red Hat Enterprise operating system entitlement or Red Hat Enterprise Premium Service Level Agreement.
A Red Hat Enterprise operating system subscription and Red Hat Enterprise Premium Service Level Agreement are required.
A Red Hat Enterprise Premium Service Level Agreement is required to align with the IBM Storage Ceph Service Level Agreement.
Note: See Installing Pro edition for free to install IBM Storage Ceph for a free 60 day trial.
Accessing Red Hat entitlements from IBM Storage Ceph
IBM Storage Ceph can include entitlement to use Red Hat® OpenShift® Container Platform, Red Hat Enterprise Linux® CoreOS (RHCOS), and Red Hat Enterprise Linux
(RHEL). To access these entitlements, you must link your IBM Storage Ceph to your Red Hat account. You can link your IBM Storage Ceph to its Red Hat entitlement
through IBM Passport Advantage.
About this task
Complete the following procedure to access your Red Hat entitlements:
1. Go to the IBM Passport Advantage Online tab at IBM Passport Advantage, click Sign in to your PAO Site, and log in with your IBMid.
Important:
The Customer Primary Contact name on your IBM order is passed to Red Hat to accept the Red Hat terms and conditions and link entitlements, then are tied
to that Red Hat account number.
If your company has multiple sites, an extra Sign in screen appears for you to select a specific company site number.
2. On the Passport Advantage Online page, click Download Software.
3. On the Software downloads page, confirm your Site Name and Site number and enter your IBM Storage Ceph part number in the search field, then press Enter. You
can find your IBM Storage Ceph in your Proof of Entitlement (POE) document. If you do not have your specific part number handy, or, if you have more than one IBM
Storage Ceph that you want to link, search for Storage Ceph.
4. Scroll over the product name of the IBM Storage Ceph that has the entitlement that you want to link, a View more link appears. Click View more for that IBM
Storage Ceph.
5. Verify the IBM Storage Ceph product description and click Continue.
Important: Do not change or enter data in the Version, Operating System, or Language fields.
6. Locate the Link with Red Hat heading in the Specification section. Click the order number that corresponds to the order number that you want to link your
entitlement to. Click Okay to navigate away from Passport Advantage Online to the Red Hat login page to map your IBM entitlement to your Red Hat account.
Previously linked orders show up under the Red Hat account row under Specifications.
Important: Do not click Container Install or Software downloads. These are not part of the entitlement linking process.
7. On the Red Hat login page, either log in with your existing Red Hat account, or create a new account. You must have a Red Hat account to access the OpenShift
Cluster Manager. You do not need a paid Red Hat subscription entitlement to access any IBM® offering.
8. On the Red Hat Review order summary page, verify that the information is correct and click Next.
9. On the Red Hat Link your Red Hat Account page, select Assign the Red Hat subscriptions to this Red Hat account and link my IBM order, accept the Enterprise
agreement terms, then click Confirm.
A message appears confirming that your Red Hat account is linked with your IBM Order. Your entitlement is now accessible. You can link additional orders on the Software
downloads page of Passport Advantage Online by clicking the SDMA tab in your browser.
For more information about viewing and managing your Red Hat subscriptions, see the Red Hat Customer Portal page.
If these steps do not resolve your Red Hat product entitlement issues, contact IBM eCustomer Care.
Minimum hardware considerations
Before running IBM Storage Ceph, be sure to comply with all minimum hardware requirements listed here.
IBM Storage Ceph can run on non-proprietary commodity hardware. Small production clusters and development clusters can run without performance optimization with
modest hardware.
Note: Disk space requirements are based on the Ceph daemons' default path under /var/lib/ceph/ directory.
For more information about the IBM Storage Ceph internal components and the strategies around those components, see Storage Strategies.
Table 1. Containers
Process
ceph-osdcontainer
Criteria
Minimum Recommended
Processor
1x AMD64 or Intel 64 CPU CORE per OSD container
RAM
Minimum of 5 GB of RAM per OSD container
OS Disk
1x OS disk per host
OSD Storage
block.db
1x storage drive per OSD container. Cannot be shared with OS Disk.
block.wal
Optionally, 1x SSD or NVMe or Optane partition or logical volume per daemon. Use a small size, for example 10 GB, and only if it’s
faster than the block.db device.
70 IBM Storage Ceph
Optional, but IBM recommended, 1x SSD or NVMe or Optane partition or lvm per daemon. Sizing is 4% of block.data for
BlueStore for object, file, and mixed workloads and 1% of block.data for the BlueStore for Block Device, Openstack cinder, and
Openstack cinder workloads.
Process
ceph-moncontainer
ceph-mgrcontainer
cephradosgwcontainer
ceph-mdscontainer
Criteria
Minimum Recommended
Network
2x 10 GB Ethernet NICs
Processor
1x AMD64 or Intel 64 CPU CORE per mon-container
RAM
3 GB per mon-container
Disk Space
10 GB per mon-container, 50 GB Recommended
Monitor Disk
Optionally, 1x SSD disk for Monitor rocksdb data
Network
2x 1 GB Ethernet NICs, 10 GB Recommended
Prometheus
20 GB to 50 GB under /var/lib/ceph/ directory created as a separate file system to protect the contents under /var/ directory.
Processor
1x AMD64 or Intel 64 CPU CORE per mgr-container
RAM
3 GB per mgr-container
Network
2x 1 GB Ethernet NICs, 10 GB Recommended
Processor
1x AMD64 or Intel 64 CPU CORE per radosgw-container
RAM
1 GB per daemon
Disk Space
5 GB per daemon
Network
1x 1 GB Ethernet NICs
Processor
1x AMD64 or Intel 64 CPU CORE per mds-container
RAM
3 GB per mds-container
This number is highly dependent on the configurable MDS cache size. The RAM requirement is typically twice as much as the
amount set in the mds_cache_memory_limit configuration setting. Note also that this is the memory for your daemon, not the
overall system memory.
Disk Space
2 GB per mds-container, plus considering any additional space required for possible debug logging, 20 GB is a good start.
Network
2x 1 GB Ethernet NICs, 10 GB Recommended
Note that this is the same network as the OSD containers. If you have a 10 GB network on your OSDs you should use the same on
your MDS so that the MDS is not disadvantaged when it comes to latency.
Hardware
Get a high level guidance on selecting hardware for use with IBM Storage Ceph.
Executive summary
General principles for selecting hardware
Optimize workload performance domains
Server and rack solutions
Minimum hardware recommendations for containerized Ceph
Ceph can run on non-proprietary commodity hardware. Small production clusters and development clusters can run without performance optimization with modest
hardware.
Recommended minimum hardware requirements for the IBM Storage Ceph Dashboard
Executive summary
Many hardware vendors now offer both Ceph-optimized servers and rack-level solutions designed for distinct workload profiles. To simplify the hardware selection
process and reduce risk for organizations, IBM has worked with multiple storage server vendors to test and evaluate specific cluster options for different cluster sizes and
workload profiles. IBM’s exacting methodology combines performance testing with proven guidance for a broad range of cluster capabilities and sizes.
With appropriate storage servers and rack-level solutions, IBM Storage Ceph can provide storage pools serving a variety of workloads—from throughput-sensitive and cost
and capacity-focused workloads to emerging IOPS-intensive workloads.
IBM Storage Ceph significantly lowers the cost of storing enterprise data and helps organizations manage exponential data growth. The software is a robust and modern
petabyte-scale storage platform for public or private cloud deployments. IBM Storage Ceph offers mature interfaces for enterprise block and object storage, making it an
optimal solution for active archive, rich media, and cloud infrastructure workloads characterized by tenant-agnostic OpenStack® environments1. Delivered as a unified,
software-defined, scale-out storage platform, IBM Storage Ceph lets businesses focus on improving application innovation and availability by offering capabilities such as:
Scaling to hundreds of petabytes2.
No single point of failure in the cluster.
Lower capital expenses (CapEx) by running on commodity server hardware.
Lower operational expenses (OpEx) with self-managing and self-healing properties.
IBM Storage Ceph can run on myriad industry-standard hardware configurations to satisfy diverse needs. To simplify and accelerate the cluster design process, IBM
conducts extensive performance and suitability testing with participating hardware vendors. This testing allows evaluation of selected hardware under load and generates
essential performance and sizing data for diverse workloads—ultimately simplifying Ceph storage cluster hardware selection. As discussed in this guide, multiple hardware
vendors now provide server and rack-level solutions optimized for IBM Storage Ceph deployments with IOPS-, throughput-, and cost and capacity-optimized solutions as
available options.
Software-defined storage presents many advantages to organizations seeking scale-out solutions to meet demanding applications and escalating storage needs. With a
proven methodology and extensive testing performed with multiple vendors, IBM simplifies the process of selecting hardware to meet the demands of any environment.
Importantly, the guidelines and example systems listed in this document are not a substitute for quantifying the impact of production workloads on sample systems.
1 Ceph is and has been the leading storage for OpenStack according to several semi-annual OpenStack user surveys.
2 See Yahoo Cloud Object Store - Object Storage at Exabyte Scale for details.
IBM Storage Ceph 71
General principles for selecting hardware
As a storage administrator, you must select the appropriate hardware for running a production IBM Storage Ceph cluster. When selecting hardware for IBM Storage Ceph,
review these following general principles. These principles will help save time, avoid common mistakes, save money and achieve a more effective solution.
Identify performance use case
Consider storage density
Identical hardware configuration
Network considerations for IBM Storage Ceph
Avoid using RAID and SAN solutions
Summary of common mistakes when selecting hardware
Reference
Prerequisites
A planned use for IBM Storage Ceph.
Linux System Administration Advance level with Red Hat Enterprise Linux certification.
Storage administrator with Ceph Certification.
Identify performance use case
One of the most important steps in a successful Ceph deployment is identifying a price-to-performance profile suitable for the cluster’s use case and workload. It is
important to choose the right hardware for the use case. For example, choosing IOPS-optimized hardware for a cloud storage application increases hardware costs
unnecessarily. Whereas, choosing capacity-optimized hardware for its more attractive price point in an IOPS-intensive workload will likely lead to unhappy users
complaining about slow performance.
The primary use cases for Ceph are:
IOPS optimized: IOPS optimized deployments are suitable for cloud computing operations, such as running MYSQL or MariaDB instances as virtual machines on
OpenStack. IOPS optimized deployments require higher performance storage such as 15k RPM SAS drives and separate SSD journals to handle frequent write
operations. Some high IOPS scenarios use all flash storage to improve IOPS and total throughput.
Throughput optimized: Throughput-optimized deployments are suitable for serving up significant amounts of data, such as graphic, audio and video content.
Throughput-optimized deployments require networking hardware, controllers and hard disk drives with acceptable total throughput characteristics. In cases where
write performance is a requirement, SSD journals will substantially improve write performance.
Capacity optimized: Capacity-optimized deployments are suitable for storing significant amounts of data as inexpensively as possible. Capacity-optimized
deployments typically trade performance for a more attractive price point. For example, capacity-optimized deployments often use slower and less expensive SATA
drives and co-locate journals rather than using SSDs for journaling.
This document provides examples of IBM tested hardware suitable for these use cases.
Consider storage density
Hardware planning should include distributing Ceph daemons and other processes that use Ceph across many hosts to maintain high availability in the event of hardware
faults. Balance storage density considerations with the need to rebalance the cluster in the event of hardware faults. A common hardware selection mistake is to use very
high storage density in small clusters, which can overload networking during backfill and recovery operations.
Identical hardware configuration
Create pools and define CRUSH hierarchies such that the OSD hardware within the pool is identical.
Same controller.
Same drive size.
Same RPMs.
Same seek times.
Same I/O.
Same network throughput.
Same journal configuration.
Using the same hardware within a pool provides a consistent performance profile, simplifies provisioning and streamlines troubleshooting.
Network considerations for IBM Storage Ceph
72 IBM Storage Ceph
An important aspect of a cloud storage solution is that storage clusters can run out of IOPS due to network latency, and other factors. Also, the storage cluster can run out
of throughput due to bandwidth constraints long before the storage clusters run out of storage capacity. This means that the network hardware configuration must support
the chosen workloads to meet price versus performance requirements.
Storage administrators prefer that a storage cluster recovers as quickly as possible. Carefully consider bandwidth requirements for the storage cluster network, be mindful
of network link oversubscription, and segregate the intra-cluster traffic from the client-to-cluster traffic. Also consider that network performance is increasingly important
when considering the use of Solid State Disks (SSD), flash, NVMe, and other high performing storage devices.
Ceph supports a public network and a storage cluster network. The public network handles client traffic and communication with Ceph Monitors. The storage cluster
network handles Ceph OSD heartbeats, replication, backfilling, and recovery traffic. At a minimum, a single 10 GB Ethernet link should be used for storage hardware, and
you can add additional 10 GB Ethernet links for connectivity and throughput.
Important: IBM recommends allocating bandwidth to the storage cluster network, such that it is a multiple of the public network using the osd_pool_default_size as
the basis for the multiple on replicated pools. IBM also recommends running the public and storage cluster networks on separate network cards.
IBM recommends using 10 GB Ethernet for IBM Storage Ceph deployments in production. A 1 GB Ethernet network is not suitable for production storage clusters.
In the case of a drive failure, replicating 1 TB of data across a 1 GB Ethernet network takes 3 hours, and 3 TB takes 9 hours. Using 3 TB is the typical drive configuration.
By contrast, with a 10 GB Ethernet network, the replication times would be 20 minutes and 1 hour. Remember that when a Ceph OSD fails, the storage cluster will recover
by replicating the data it contained to other Ceph OSDs within the pool.
The failure of a larger domain such as a rack means that the storage cluster utilizes considerably more bandwidth. When building a storage cluster consisting of multiple
racks, which is common for large storage implementations, consider utilizing as much network bandwidth between switches in a "fat tree" design for optimal performance.
A typical 10 GB Ethernet switch has 48 10 GB ports and four 40 GB ports. Use the 40 GB ports on the spine for maximum throughput. Alternatively, consider aggregating
unused 10 GB ports with QSFP+ and SFP+ cables into more 40 GB ports to connect to other rack and spine routers. Also, consider using LACP mode 4 to bond network
interfaces. Additionally, use jumbo frames, with a maximum transmission unit (MTU) of 9000, especially on the backend or cluster network.
Before installing and testing a IBM Storage Ceph cluster, verify the network throughput. Most performance-related problems in Ceph usually begin with a networking
issue. Simple network issues like a kinked or bent Cat-6 cable could result in degraded bandwidth. Use a minimum of 10 GB ethernet for the front side network. For large
clusters, consider using 40 GB ethernet for the backend or cluster network.
Important: For network optimization, IBM recommends using jumbo frames for a better CPU per bandwidth ratio, and a non-blocking network switch back-plane. IBM
Storage Ceph requires the same MTU value throughout all networking devices in the communication path, end-to-end for both public and cluster networks. Verify that the
MTU value is the same on all hosts and networking equipment in the environment before using a IBM Storage Ceph cluster in production.
Avoid using RAID and SAN solutions
Ceph can replicate or erasure code objects. RAID duplicates this functionality on the block level and reduces available capacity. Consequently, RAID is an unnecessary
expense. Additionally, a degraded RAID will have a negative impact on performance.
Important: IBM does not support SAN as a storage backend for OSDs.
Important: IBM recommends that each hard drive be exported separately from the RAID controller as a single volume with write-back caching enabled.
This requires a battery-backed, or a non-volatile flash memory device on the storage controller. It is important to make sure the battery is working, as most controllers will
disable write-back caching if the memory on the controller can be lost as a result of a power failure. Periodically, check the batteries and replace them if necessary, as they
do degrade over time. See the storage controller vendor’s documentation for details. Typically, the storage controller vendor provides storage management utilities to
monitor and adjust the storage controller configuration without any downtime.
Using Just a Bunch of Drives (JBOD) in independent drive mode with Ceph is supported when using all Solid State Drives (SSDs), or for configurations with high numbers of
drives per controller. For example, 60 drives attached to one controller. In this scenario, the write-back caching can become a source of I/O contention. Since JBOD
disables write-back caching, it is ideal in this scenario. One advantage of using JBOD mode is the ease of adding or replacing drives and then exposing the drive to the
operating system immediately after it is physically plugged in.
Summary of common mistakes when selecting hardware
Repurposing underpowered legacy hardware for use with Ceph.
Using dissimilar hardware in the same pool.
Using 1Gbps networks instead of 10Gbps or greater.
Neglecting to setup both public and cluster networks.
Using RAID instead of JBOD.
Selecting drives on a price basis without regard to performance or throughput.
Journaling on OSD data drives when the use case calls for an SSD journal.
Having a disk controller with insufficient throughput characteristics.
Use the examples in this document of IBM tested configurations for different workloads to avoid some of the foregoing hardware selection mistakes.
Reference
Supported configurations IBM Storage Ceph Supported configurations on the portal.
IBM Storage Ceph 73
Optimize workload performance domains
One of the key benefits of Ceph storage is the ability to support different types of workloads within the same cluster using Ceph performance domains. Dramatically
different hardware configurations can be associated with each performance domain. Ceph system administrators can deploy storage pools on the appropriate
performance domain, providing applications with storage tailored to specific performance and cost profiles. Selecting appropriately sized and optimized servers for these
performance domains is an essential aspect of designing a IBM Storage Ceph cluster.
The following lists provide the criteria IBM uses to identify optimal IBM Storage Ceph cluster configurations on storage servers. These categories are provided as general
guidelines for hardware purchases and configuration decisions, and can be adjusted to satisfy unique workload blends. Actual hardware configurations chosen will vary
depending on specific workload mix and vendor capabilities.
IOPS optimized
An IOPS-optimized storage cluster typically has the following properties:
Lowest cost per IOPS.
Highest IOPS per GB.
99th percentile latency consistency.
Typically uses for an IOPS-optimized storage cluster are:
Typically block storage.
3x replication for hard disk drives (HDDs) or 2x replication for solid state drives (SSDs).
MySQL on OpenStack clouds.
Throughput optimized
A throughput-optimized storage cluster typically has the following properties:
Lowest cost per MBps (throughput).
Highest MBps per TB.
Highest MBps per BTU.
Highest MBps per Watt.
97th percentile latency consistency.
Typically uses for an throughput-optimized storage cluster are:
Block or object storage.
3x replication.
Active performance storage for video, audio, and images.
Streaming media.
Cost and capacity optimized
A cost- and capacity-optimized storage cluster typically has the following properties:
Lowest cost per TB.
Lowest BTU per TB.
Lowest Watts required per TB.
Typically uses for an cost- and capacity-optimized storage cluster are:
Typically object storage.
Erasure coding common for maximizing usable capacity
Object archive.
Video, audio, and image object repositories.
How performance domains work
To the Ceph client interface that reads and writes data, a Ceph storage cluster appears as a simple pool where the client stores data. However, the storage cluster
performs many complex operations in a manner that is completely transparent to the client interface. Ceph clients and Ceph object storage daemons (Ceph OSDs, or
simply OSDs) both use the controlled replication under scalable hashing (CRUSH) algorithm for storage and retrieval of objects. OSDs run on OSD hosts—the storage
servers within the cluster.
A CRUSH map describes a topography of cluster resources, and the map exists both on client nodes as well as Ceph Monitor (MON) nodes within the cluster. Ceph clients
and Ceph OSDs both use the CRUSH map and the CRUSH algorithm. Ceph clients communicate directly with OSDs, eliminating a centralized object lookup and a potential
performance bottleneck. With awareness of the CRUSH map and communication with their peers, OSDs can handle replication, backfilling, and recovery—allowing for
dynamic failure recovery.
Ceph uses the CRUSH map to implement failure domains. Ceph also uses the CRUSH map to implement performance domains, which simply take the performance profile
of the underlying hardware into consideration. The CRUSH map describes how Ceph stores data, and it is implemented as a simple hierarchy (acyclic graph) and a ruleset.
74 IBM Storage Ceph
The CRUSH map can support multiple hierarchies to separate one type of hardware performance profile from another.
The following examples describe performance domains.
Hard disk drives (HDDs) are typically appropriate for cost- and capacity-focused workloads.
Throughput-sensitive workloads typically use HDDs with Ceph write journals on solid state drives (SSDs).
IOPS-intensive workloads such as MySQL and MariaDB often use SSDs.
All of these performance domains can coexist in a Ceph storage cluster.
Server and rack solutions
Ceph provides both optimized server-level and rack-level solution SKUs. Hardware vendors have responded to the enthusiasm around Ceph by providing both optimized
server-level and rack-level solution SKUs. Validated through joint testing with IBM, these solutions offer predictable price-to-performance ratios for Ceph deployments,
with a convenient modular approach to expand Ceph storage for specific workloads.
Typical rack-level solutions include:
Network switching
Redundant network switching interconnects the cluster and provides access to clients.
Ceph MON nodes
The Ceph monitor is a datastore for the health of the entire cluster, and contains the cluster log. A minimum of three monitor nodes are strongly recommended for a
cluster quorum in production.
Ceph OSD hosts
Ceph OSD hosts house the storage capacity for the cluster, with one or more OSDs running per individual storage device. OSD hosts are selected and configured
differently depending on both workload optimization and the data devices installed: HDDs, SSDs, or NVMe SSDs.
IBM Storage Ceph
Many vendors provide a capacity-based subscription for IBM Storage Ceph bundled with both server and rack-level solution SKUs.
Note: For more information and assistance, contact IBM Support.
IOPS-optimized solutions
With the growing use of flash storage, organizations increasingly host IOPS-intensive workloads on Ceph storage clusters to let them emulate high-performance public
cloud solutions with private cloud storage. These workloads commonly involve structured data from MySQL-, MariaDB-, or PostgreSQL-based applications.
Typical servers include the following elements:
CPU: 10 cores per NVMe SSD, assuming a 2 GHz CPU.
RAM: 16 GB baseline, plus 5 GB per OSD.
Networking: 10 Gigabit Ethernet (GbE) per 2 OSDs.
OSD media: High-performance, high-endurance enterprise NVMe SSDs.
OSDs: Two per NVMe SSD.
Bluestore WAL/DB: High-performance, high-endurance enterprise NVMe SSD, co-located with OSDs.
Controller: Native PCIe bus.
Note: For Non-NVMe SSDs, for CPU, use two cores per SSD OSD.
Table 1. Solutions SKUs for IOPS-optimized Ceph workloads, by
cluster size
Vendor
SuperMicro 1
Small (250TB)
Medium (1PB)
SYS-5038MR-OSD006P N/A
Large (2PB+)
N/A
1See Supermicro® Total Solution for Ceph.
Throughput-optimized solutions
Throughput-optimized Ceph solutions are usually centered around semi-structured or unstructured data. Large-block sequential I/O is typical.
Typical server elements include:
CPU: 0.5 cores per HDD, assuming a 2 GHz CPU.
RAM: 16 GB baseline, plus 5 GB per OSD.
Networking: 10 GbE per 12 OSDs each for client- and cluster-facing networks.
OSD media: 7,200 RPM enterprise HDDs.
OSDs: One per HDD.
Bluestore WAL/DB: High-performance, high-endurance enterprise NVMe SSD, co-located with OSDs.
Host bus adapter (HBA): Just a bunch of disks (JBOD).
IBM Storage Ceph 75
Several vendors provide pre-configured server and rack-level solutions for throughput-optimized Ceph workloads. IBM has conducted extensive testing and evaluation of
servers from Supermicro and Quanta Cloud Technologies (QCT).
Table 2. Rack-level SKUs for Ceph OSDs, MONs, and top-of-rack (TOR)
switches.
Vendor
SuperMicro
Small (250TB)
Medium (1PB)
Large (2PB+)
SRS-42E112-Ceph-03 SRS-42E136-Ceph-03 SRS-42E136-Ceph-03
Table 3. Individual OSD servers
Vendor
Small (250TB)
Medium (1PB)
Large (2PB+)
SuperMicro
SSG-6028R-OSD072P SSG-6048-OSD216P SSG-6048-OSD216P
QCT1
QxStor RCT-200
QxStor RCT-400
QxStor RCT-400
1See QCT: QxStor IBM Storage Ceph Edition.
Table 4. Additional servers configurable for throughput-optimized Ceph OSD workloads
Small (250TB)
Medium (1PB)
Large (2PB+)
Dell
Vendor
PowerEdge R730XD 1
DSS 7000 2, twin node
DSS 7000, twin node
Cisco
UCS C240 M4
UCS C3260 3
UCS C3260 4
Lenovo
System x3650 M5
System x3650 M5
N/A
1See Dell PowerEdge R730xd Performance and Sizing Guide for IBM Storage Ceph - A Dell IBM Technical White Paper.
2See Dell EMC DSS 7000 Performance & Sizing Guide for IBM Storage Ceph.
3See Hardware.
4See UCS C3260.
Cost and capacity-optimized solutions
Cost- and capacity-optimized solutions typically focus on higher capacity, or longer archival scenarios. Data can be either semi-structured or unstructured. Workloads
include media archives, big data analytics archives, and machine image backups. Large-block sequential I/O is typical.
Solutions typically include the following elements:
CPU. 0.5 cores per HDD, assuming a 2 GHz CPU.
RAM. 16 GB baseline, plus 5 GB per OSD.
Networking. 10 GbE per 12 OSDs (each for client- and cluster-facing networks).
OSD media. 7,200 RPM enterprise HDDs.
OSDs. One per HDD.
Bluestore WAL/DB Co-located on the HDD.
HBA. JBOD.
Supermicro and QCT provide pre-configured server and rack-level solution SKUs for cost- and capacity-focused Ceph workloads.
Table 5. Pre-configured Rack-level SKUs for Cost- and Capacityoptimized Workloads
Vendor
SuperMicro
Small (250TB)
N/A
Medium (1PB)
Large (2PB+)
SRS-42E136-Ceph-03 SRS-42E172-Ceph-03
Table 6. Pre-configured server-level SKUs for cost- and capacityoptimized workloads
Vendor
Small (250TB)
Medium (1PB)
Large (2PB+)
SuperMicro
N/A
SSG-6048R-OSD216P 1 SSD-6048R-OSD360P
QCT
N/A
QxStor RCC-400
QxStor RCC-400
1See Supermicro’s Total Solution for Ceph.
Table 7. Additional servers configurable for cost- and capacityoptimized workloads
Vendor
Small (250TB)
Medium (1PB)
Large (2PB+)
Dell
N/A
DSS 7000, twin node DSS 7000, twin node
Cisco
N/A
UCS C3260
UCS C3260
Lenovo
N/A
System x3650 M5
N/A
Red Hat Ceph Storage on Samsung NVMe SSDs
Deploying MySQL Databases on Red Hat Ceph Storage
Intel® Data Center Blocks for Cloud – IBM OpenStack Platform with Red Hat Ceph Storage
Red Hat Ceph Storage on QCT Servers
Red Hat Ceph Storage on Servers with Intel Processors and SSDs
Minimum hardware recommendations for containerized Ceph
76 IBM Storage Ceph
Ceph can run on non-proprietary commodity hardware. Small production clusters and development clusters can run without performance optimization with modest
hardware.
Table 1. Minimum hardware recommendations for containerized Ceph
Process
ceph-osdcontainer
ceph-moncontainer
ceph-mgrcontainer
cephradosgwcontainer
ceph-mdscontainer
Criteria
Minimum Recommended
Processor
1x AMD64 or Intel 64 CPU CORE per OSD container
RAM
Minimum of 5 GB of RAM per OSD container
OS Disk
1x OS disk per host
OSD Storage
block.db
1x storage drive per OSD container. Cannot be shared with OS Disk.
block.wal
Optionally, 1x SSD or NVMe or Optane partition or logical volume per daemon. Use a small size, for example 10 GB, and only if it’s
faster than the block.db device.
Network
2x 10 GB Ethernet NICs
Processor
1x AMD64 or Intel 64 CPU CORE per mon-container
RAM
3 GB per mon-container
Disk Space
10 GB per mon-container, 50 GB Recommended
Monitor Disk
Optionally, 1x SSD disk for Monitor rocksdb data
Network
2x 1GB Ethernet NICs, 10 GB Recommended
Processor
1x AMD64 or Intel 64 CPU CORE per mgr-container
RAM
3 GB per mgr-container
Network
2x 1GB Ethernet NICs, 10 GB Recommended
Processor
1x AMD64 or Intel 64 CPU CORE per radosgw-container
RAM
1 GB per daemon
Disk Space
5 GB per daemon
Network
1x 1GB Ethernet NICs
Processor
1x AMD64 or Intel 64 CPU CORE per mds-container
RAM
3 GB per mds-container
Optional, but IBM recommended, 1x SSD or NVMe or Optane partition or lvm per daemon. Sizing is 4% of block.data for
BlueStore for object, file and mixed workloads and 1% of block.data for the BlueStore for Block Device, Openstack cinder, and
Openstack cinder workloads.
This number is highly dependent on the configurable MDS cache size. The RAM requirement is typically twice as much as the
amount set in the mds_cache_memory_limit configuration setting. Note also that this is the memory for your daemon, not the
overall system memory.
Disk Space
2 GB per mds-container, plus taking into consideration any additional space required for possible debug logging, 20GB is a good
start.
Network
2x 1GB Ethernet NICs, 10 GB Recommended
Note that this is the same network as the OSD containers. If you have a 10 GB network on your OSDs you should use the same on
your MDS so that the MDS is not disadvantaged when it comes to latency.
Recommended minimum hardware requirements for the IBM Storage Ceph
Dashboard
The IBM Storage Ceph Dashboard has minimum hardware requirements.
Minimum requirements
4 core processor at 2.5 GHz or higher
8 GB RAM
50 GB hard disk drive
1 Gigabit Ethernet network interface
Reference
For more information, see High-level monitoring of a Ceph storage cluster.
Storage Strategies
Creating storage strategies for IBM Storage Ceph clusters
This section of the document provides instructions for creating storage strategies, including creating CRUSH hierarchies, estimating the number of placement groups,
determining which type of storage pool to create, and managing pools.
Overview
Crush admin overview
Placement Groups
Pools overview
Erasure code pools overview
IBM Storage Ceph 77
Overview
From the perspective of a Ceph client, interacting with the Ceph storage cluster is remarkably simple:
1. Connect to the Cluster
2. Create a Pool I/O Context
This remarkably simple interface is how a Ceph client selects one of the storage strategies you define. Storage strategies are invisible to the Ceph client in all but storage
capacity and performance.
The diagram below shows the logical data flow starting from the client into the IBM Storage Ceph cluster.
Figure 1. Ceph storage architecture data flow
What are storage strategies?
Configuring storage strategies
What are storage strategies?
A storage strategy is a method of storing data that serves a particular use case. For example, if you need to store volumes and images for a cloud platform like OpenStack,
you might choose to store data on reasonably performant SAS drives with SSD-based journals. By contrast, if you need to store object data for an S3- or Swift-compliant
gateway, you might choose to use something more economical, like SATA drives. Ceph can accommodate both scenarios in the same Ceph cluster, but you need a means
of providing the SAS/SSD storage strategy to the cloud platform (for example, Glance and Cinder in OpenStack), and a means of providing SATA storage for your object
store.
Storage strategies include the storage media (hard drives, SSDs, and the rest), the CRUSH maps that set up performance and failure domains for the storage media, the
number of placement groups, and the pool interface. Ceph supports multiple storage strategies. Use cases, cost/benefit performance tradeoffs and data durability are the
primary considerations that drive storage strategies.
78 IBM Storage Ceph
1. Use Cases: Ceph provides massive storage capacity, and it supports numerous use cases. For example, the Ceph Block Device client is a leading storage backend
for cloud platforms like OpenStack—providing limitless storage for volumes and images with high performance features like copy-on-write cloning. Likewise, Ceph
can provide container-based storage for OpenShift environments. By contrast, the Ceph Object Gateway client is a leading storage backend for cloud platforms that
provides RESTful S3-compliant and Swift-compliant object storage for objects like audio, bitmap, video and other data.
2. Cost/Benefit of Performance: Faster is better. Bigger is better. High durability is better. However, there is a price for each superlative quality, and a corresponding
cost/benefit trade off. Consider the following use cases from a performance perspective: SSDs can provide very fast storage for relatively small amounts of data and
journaling. Storing a database or object index might benefit from a pool of very fast SSDs, but prove too expensive for other data. SAS drives with SSD journaling
provide fast performance at an economical price for volumes and images. SATA drives without SSD journaling provide cheap storage with lower overall
performance. When you create a CRUSH hierarchy of OSDs, you need to consider the use case and an acceptable cost/performance trade off.
3. Durability: In large scale clusters, hardware failure is an expectation, not an exception. However, data loss and service interruption remain unacceptable. For this
reason, data durability is very important. Ceph addresses data durability with multiple deep copies of an object or with erasure coding and multiple coding chunks.
Multiple copies or multiple coding chunks present an additional cost/benefit tradeoff: it’s cheaper to store fewer copies or coding chunks, but it might lead to the
inability to service write requests in a degraded state. Generally, one object with two additional copies (that is, size = 3) or two coding chunks might allow a
cluster to service writes in a degraded state while the cluster recovers. The CRUSH algorithm aids this process by ensuring that Ceph stores additional copies or
coding chunks in different locations within the cluster. This ensures that the failure of a single storage device or node doesn’t lead to a loss of all of the copies or
coding chunks necessary to preclude data loss.
You can capture use cases, cost/benefit performance tradeoffs and data durability in a storage strategy and present it to a Ceph client as a storage pool.
Important: Ceph’s object copies or coding chunks make RAID obsolete. Do not use RAID, because Ceph already handles data durability, a degraded RAID has a negative
impact on performance, and recovering data using RAID is substantially slower than using deep copies or erasure coding chunks.
Configuring storage strategies
Configuring storage strategies is about assigning Ceph OSDs to a CRUSH hierarchy, defining the number of placement groups for a pool, and creating a pool. The general
steps are:
1. Define a Storage Strategy: Storage strategies require you to analyze your use case, cost/benefit performance tradeoffs and data durability. Then, you create OSDs
suitable for that use case. For example, you can create SSD-backed OSDs for a high performance pool; SAS drive/SSD journal-backed OSDs for high-performance
block device volumes and images; or, SATA-backed OSDs for low cost storage. Ideally, each OSD for a use case should have the same hardware configuration so
that you have a consistent performance profile.
2. Define a CRUSH Hierarchy: Ceph rules select a node, usually the root, in a CRUSH hierarchy, and identify the appropriate OSDs for storing placement groups and
the objects they contain. You must create a CRUSH hierarchy and a CRUSH rule for your storage strategy. CRUSH hierarchies get assigned directly to a pool by the
CRUSH rule setting.
3. Calculate Placement Groups: Ceph shards a pool into placement groups. You do not have to manually set the number of placement groups for your pool. PG
autoscaler sets an appropriate number of placement groups for your pool that remains within a healthy maximum number of placement groups in the event that
you assign multiple pools to the same CRUSH rule.
4. Create a Pool: Finally, you must create a pool and determine whether it uses replicated or erasure-coded storage. You must set the number of placement groups
for the pool, the rule for the pool and the durability, such as size or K+M coding chunks.
Remember, the pool is the Ceph client’s interface to the storage cluster, but the storage strategy is completely transparent to the Ceph client, except for capacity and
performance.
Crush admin overview
The Controlled Replication Under Scalable Hashing (CRUSH) algorithm determines how to store and retrieve data by computing data storage locations.
Any sufficiently advanced technology is indistinguishable from magic.
— Arthur C. Clarke
CRUSH introduction
The CRUSH map for your storage cluster describes your device locations within CRUSH hierarchies and a rule for each hierarchy that determines how Ceph stores
data.
CRUSH hierarchy
Ceph OSDs in CRUSH
Device class
The Ceph CRUSH map provides a lot of flexibility in controlling data placement.
CRUSH weights
Primary affinity
CRUSH rules
CRUSH tunables overview
Edit a CRUSH map
Generally, modifying your CRUSH map at runtime with the Ceph CLI is more convenient than editing the CRUSH map manually. However, there are times when you
might choose to edit it, such as changing the default bucket types, or using a bucket algorithm other than straw2.
CRUSH storage strategies examples
CRUSH introduction
The CRUSH map for your storage cluster describes your device locations within CRUSH hierarchies and a rule for each hierarchy that determines how Ceph stores data.
IBM Storage Ceph 79
The CRUSH map contains at least one hierarchy of nodes and leaves. The nodes of a hierarchy, called "buckets" in Ceph, are any aggregation of storage locations as
defined by their type. For example, rows, racks, chassis, hosts, and devices. Each leaf of the hierarchy consists essentially of one of the storage devices in the list of
storage devices. A leaf is always contained in one node or "bucket." A CRUSH map also has a list of rules that determine how CRUSH stores and retrieves data.
Note: Storage devices are added to the CRUSH map when adding an OSD to the cluster.
The CRUSH algorithm distributes data objects among storage devices according to a per-device weight value, approximating a uniform probability distribution. CRUSH
distributes objects and their replicas or erasure-coding chunks according to the hierarchical cluster map an administrator defines. The CRUSH map represents the
available storage devices and the logical buckets that contain them for the rule, and by extension each pool that uses the rule.
To map placement groups to OSDs across failure domains or performance domains, a CRUSH map defines a hierarchical list of bucket types; that is, under types in the
generated CRUSH map. The purpose of creating a bucket hierarchy is to segregate the leaf nodes by their failure domains or performance domains or both. Failure
domains include hosts, chassis, racks, power distribution units, pods, rows, rooms, and data centers. Performance domains include failure domains and OSDs of a
particular configuration. For example, SSDs, SAS drives with SSD journals, SATA drives, and so on. Devices have the notion of a class, such as hdd, ssd and nvme to more
rapidly build CRUSH hierarchies with a class of devices.
With the exception of the leaf nodes representing OSDs, the rest of the hierarchy is arbitrary, and you can define it according to your own needs if the default types do not
suit your requirements. We recommend adapting your CRUSH map bucket types to your organization’s hardware naming conventions and using instance names that
reflect the physical hardware names. Your naming practice can make it easier to administer the cluster and troubleshoot problems when an OSD or other hardware
malfunctions and the administrator needs remote or physical access to the host or other hardware.
In the following example, the bucket hierarchy has four leaf buckets (osd 1-4), two node buckets (host 1-2) and one rack node (rack 1).
Figure 1. CRUSH hierarchy
Since leaf nodes reflect storage devices declared under the devices list at the beginning of the CRUSH map, there is no need to declare them as bucket instances. The
second lowest bucket type in the hierarchy usually aggregates the devices; that is, it is usually the computer containing the storage media, and uses whatever term
administrators prefer to describe it, such as "node", "computer", "server," "host", "machine", and so on. In high density environments, it is increasingly common to see
multiple hosts/nodes per card and per chassis. Make sure to account for card and chassis failure too, for example, the need to pull a card or chassis if a node fails can
result in bringing down numerous hosts/nodes and their OSDs.
When declaring a bucket instance, specify its type, give it a unique name as a string, assign it an optional unique ID expressed as a negative integer, specify a weight
relative to the total capacity or capability of its items, specify the bucket algorithm such as straw2, and the hash that is usually 0 reflecting hash algorithm rjenkins1. A
bucket can have one or more items. The items can consist of node buckets or leaves. Items can have a weight that reflects the relative weight of the item.
Dynamic data placement
CRUSH failure domain
CRUSH performance domain
Using different device classes
To create performance domains, use device classes and a single CRUSH hierarchy.
Dynamic data placement
Ceph Clients and Ceph OSDs both use the CRUSH map and the CRUSH algorithm.
80 IBM Storage Ceph
Ceph Clients: By distributing CRUSH maps to Ceph clients, CRUSH empowers Ceph clients to communicate with OSDs directly. This means that Ceph clients avoid
a centralized object look-up table that could act as a single point of failure, a performance bottleneck, a connection limitation at a centralized look-up server and a
physical limit to the storage cluster’s scalability.
Ceph OSDs: By distributing CRUSH maps to Ceph OSDs, Ceph empowers OSDs to handle replication, backfilling and recovery. This means that the Ceph OSDs
handle storage of object replicas (or coding chunks) on behalf of the Ceph client. It also means that Ceph OSDs know enough about the cluster to re-balance the
cluster (backfilling) and recover from failures dynamically.
CRUSH failure domain
Having multiple object replicas or M erasure coding chunks helps prevent data loss, but it is not sufficient to address high availability. By reflecting the underlying physical
organization of the Ceph Storage Cluster, CRUSH can model—and thereby address—potential sources of correlated device failures. By encoding the cluster’s topology into
the cluster map, CRUSH placement policies can separate object replicas or erasure coding chunks across different failure domains while still maintaining the desired
pseudo-random distribution.
For example, to address the possibility of concurrent failures, it might be desirable to ensure that data replicas or erasure coding chunks are on devices using different
shelves, racks, power supplies, controllers or physical locations. This helps to prevent data loss and allows the cluster to operate in a degraded state.
CRUSH performance domain
Ceph can support multiple hierarchies to separate one type of hardware performance profile from another type of hardware performance profile. For example, CRUSH can
create one hierarchy for hard disk drives and another hierarchy for SSDs. Performance domains—hierarchies that take the performance profile of the underlying hardware
into consideration—are increasingly popular due to the need to support different performance characteristics. Operationally, these are just CRUSH maps with more than
one root type bucket. Use case examples include:
Object Storage: Ceph hosts that serve as an object storage back end for S3 and Swift interfaces might take advantage of less expensive storage media such as
SATA drives that might not be suitable for VMs—reducing the cost per gigabyte for object storage, while separating more economical storage hosts from more
performing ones intended for storing volumes and images on cloud platforms. HTTP tends to be the bottleneck in object storage systems.
Cold Storage: Systems designed for cold storage—infrequently accessed data, or data retrieval with relaxed performance requirements—might take advantage of
less expensive storage media and erasure coding. However, erasure coding might require a bit of additional RAM and CPU, and thus differ in RAM and CPU
requirements from a host used for object storage or VMs.
SSD-backed Pools: SSDs are expensive, but they provide significant advantages over hard disk drives. SSDs have no seek time and they provide high total
throughput. In addition to using SSDs for journaling, a cluster can support SSD-backed pools. Common use cases include high performance SSD pools. For example,
it is possible to map the .rgw.buckets.index pool for the Ceph Object Gateway to SSDs instead of SATA drives.
A CRUSH map supports the notion of a device class. Ceph can discover aspects of a storage device and automatically assign a class such as hdd, ssd or nvme. However,
CRUSH is not limited to these defaults. For example, CRUSH hierarchies might also be used to separate different types of workloads. For example, an SSD might be used
for a journal or write-ahead log, a bucket index or for raw object storage. CRUSH can support different device classes, such as ssd-bucket-index or ssd-objectstorage so Ceph does not use the same storage media for different workloads—making performance more predictable and consistent.
Behind the scenes, Ceph generates a crush root for each device-class. These roots should only be modified by setting or changing device classes on OSDs. You can view
the generated roots using the following command:
Example
[ceph: root@host01 /]# ceph osd crush tree --show-shadow
ID
-24
-19
8
-20
7
-21
3
-22
5
-23
6
-2
-4
10
-12
0
12
-6
4
11
-10
1
-8
2
-1
-3
10
8
-11
0
12
7
CLASS WEIGHT
ssd
4.54849
ssd
0.90970
ssd
0.90970
ssd
0.90970
ssd
0.90970
ssd
0.90970
ssd
0.90970
ssd
0.90970
ssd
0.90970
ssd
0.90970
ssd
0.90970
hdd 50.94173
hdd
7.27739
hdd
7.27739
hdd 14.55478
hdd
7.27739
hdd
7.27739
hdd 14.55478
hdd
7.27739
hdd
7.27739
hdd
7.27739
hdd
7.27739
hdd
7.27739
hdd
7.27739
55.49022
8.18709
hdd
7.27739
ssd
0.90970
15.46448
hdd
7.27739
hdd
7.27739
ssd
0.90970
TYPE NAME
root default~ssd
host ceph01~ssd
osd.8
host ceph02~ssd
osd.7
host ceph03~ssd
osd.3
host ceph04~ssd
osd.5
host ceph05~ssd
osd.6
root default~hdd
host ceph01~hdd
osd.10
host ceph02~hdd
osd.0
osd.12
host ceph03~hdd
osd.4
osd.11
host ceph04~hdd
osd.1
host ceph05~hdd
osd.2
root default
host ceph01
osd.10
osd.8
host ceph02
osd.0
osd.12
osd.7
IBM Storage Ceph 81
-5
4
11
3
-9
1
5
-7
2
6
hdd
hdd
ssd
hdd
ssd
hdd
ssd
15.46448
7.27739
7.27739
0.90970
8.18709
7.27739
0.90970
8.18709
7.27739
0.90970
host ceph03
osd.4
osd.11
osd.3
host ceph04
osd.1
osd.5
host ceph05
osd.2
osd.6
Using different device classes
To create performance domains, use device classes and a single CRUSH hierarchy.
To create performance domains, add OSDs to the CRUSH hierarchy, then do the following:
1. Add a class to each device. For example:
Syntax
ceph osd crush set-device-class <class> <osdId> [<osdId>]
ceph osd crush set-device-class hdd osd.0 osd.1 osd.4 osd.5
ceph osd crush set-device-class ssd osd.2 osd.3 osd.6 osd.7
2. Then, create rules to use the devices.
Syntax
ceph osd crush rule create-replicated <rule-name> <root> <failure-domain-type> <class>
ceph osd crush rule create-replicated cold default host hdd
ceph osd crush rule create-replicated hot default host ssd
3. Finally, set pools to use the rules.
Syntax
ceph osd pool set <poolname> crush_rule <rule-name>
ceph osd pool set cold_tier crush_rule cold
ceph osd pool set hot_tier crush_rule hot
Note: There is no need to manually edit the CRUSH map.
CRUSH hierarchy
The CRUSH map is a directed acyclic graph, so it can accommodate multiple hierarchies, for example, performance domains. The easiest way to create and modify a
CRUSH hierarchy is with the Ceph CLI; however, you can also decompile a CRUSH map, edit it, recompile it, and activate it.
When declaring a bucket instance with the Ceph CLI, you must specify its type and give it a unique string name. Ceph automatically assigns a bucket ID, sets the algorithm
to straw2, sets the hash to 0 reflecting rjenkins1 and sets a weight. When modifying a decompiled CRUSH map, assign the bucket a unique ID expressed as a negative
integer (optional), specify a weight relative to the total capacity/capability of its item(s), specify the bucket algorithm (usually straw2), and the hash (usually 0, reflecting
hash algorithm rjenkins1).
A bucket can have one or more items. The items can consist of node buckets (for example, racks, rows, hosts) or leaves (for example, an OSD disk). Items can have a
weight that reflects the relative weight of the item.
When modifying a decompiled CRUSH map, you can declare a node bucket with the following syntax:
[bucket-type] [bucket-name] {
id [a unique negative numeric ID]
weight [the relative capacity/capability of the item(s)]
alg [the bucket type: uniform | list | tree | straw2 ]
hash [the hash type: 0 by default]
item [item-name] weight [weight]
}
For example, using the diagram above, we would define two host buckets and one rack bucket. The OSDs are declared as items within the host buckets:
host node1 {
id -1
alg straw2
hash 0
item osd.0 weight 1.00
item osd.1 weight 1.00
}
host node2 {
id -2
alg straw2
hash 0
item osd.2 weight 1.00
item osd.3 weight 1.00
}
rack rack1 {
id -3
82 IBM Storage Ceph
}
alg straw2
hash 0
item node1 weight 2.00
item node2 weight 2.00
Note: In the foregoing example, note that the rack bucket does not contain any OSDs. Rather it contains lower level host buckets, and includes the sum total of their
weight in the item entry.
CRUSH location
A CRUSH location is the position of an OSD in terms of the CRUSH map’s hierarchy.
Adding a bucket
Moving a bucket
Removing a bucket
Learn how to remove a bucket.
CRUSH Bucket algorithms
CRUSH location
A CRUSH location is the position of an OSD in terms of the CRUSH map’s hierarchy.
When you express a CRUSH location on the command line interface, a CRUSH location specifier takes the form of a list of name/value pairs describing the OSD’s position.
For example, if an OSD is in a particular row, rack, chassis and host, and is part of the default CRUSH tree, its crush location could be described as:
root=default row=a rack=a2 chassis=a2a host=a2a1
Note:
1. The order of the keys does not matter.
2. The key name (left of = ) must be a valid CRUSH type. By default these include root, datacenter, room, row, pod, pdu, rack, chassis and host. You might
edit the CRUSH map to change the types to suit your needs.
3. You do not need to specify all the buckets/keys. For example, by default, Ceph automatically sets a ceph-osd daemon’s location to be root=default host=
{HOSTNAME} (based on the output from hostname -s).
Adding a bucket
To add a bucket instance to your CRUSH hierarchy, specify the bucket name and its type. Bucket names must be unique in the CRUSH map.
ceph osd crush add-bucket {name} {type}
If you plan to use multiple hierarchies, for example, for different hardware performance profiles, consider naming buckets based on their type of hardware or use case.
For example, you could create a hierarchy for solid state drives (ssd), a hierarchy for SAS disks with SSD journals (hdd-journal), and another hierarchy for SATA drives
(hdd):
ceph osd crush add-bucket ssd-root root
ceph osd crush add-bucket hdd-journal-root root
ceph osd crush add-bucket hdd-root root
The Ceph CLI outputs:
added bucket ssd-root type root to crush map
added bucket hdd-journal-root type root to crush map
added bucket hdd-root type root to crush map
Important: Using colons (:) in bucket names is not supported.
Add an instance of each bucket type you need for your hierarchy. The following example demonstrates adding buckets for a row with a rack of SSD hosts and a rack of
hosts for object storage.
ceph osd crush add-bucket ssd-row1 row
ceph osd crush add-bucket ssd-row1-rack1 rack
ceph osd crush add-bucket ssd-row1-rack1-host1 host
ceph osd crush add-bucket ssd-row1-rack1-host2 host
ceph osd crush add-bucket hdd-row1 row
ceph osd crush add-bucket hdd-row1-rack2 rack
ceph osd crush add-bucket hdd-row1-rack1-host1 host
ceph osd crush add-bucket hdd-row1-rack1-host2 host
ceph osd crush add-bucket hdd-row1-rack1-host3 host
ceph osd crush add-bucket hdd-row1-rack1-host4 host
Once you have completed these steps, view your tree.
ceph osd tree
Notice that the hierarchy remains flat. You must move your buckets into a hierarchical position after you add them to the CRUSH map.
Moving a bucket
IBM Storage Ceph 83
When you create your initial cluster, Ceph has a default CRUSH map with a root bucket named default and your initial OSD hosts appear under the default bucket.
When you add a bucket instance to your CRUSH map, it appears in the CRUSH hierarchy, but it does not necessarily appear under a particular bucket.
To move a bucket instance to a particular location in your CRUSH hierarchy, specify the bucket name and its type.
Example
ceph osd crush move ssd-row1 root=ssd-root
ceph osd crush move ssd-row1-rack1 row=ssd-row1
ceph osd crush move ssd-row1-rack1-host1 rack=ssd-row1-rack1
ceph osd crush move ssd-row1-rack1-host2 rack=ssd-row1-rack1
Once you have completed these steps, you can view your tree.
ceph osd tree
Note: You can also use ceph osd crush create-or-move to create a location while moving an OSD.
Removing a bucket
Learn how to remove a bucket.
To remove a bucket instance from your CRUSH hierarchy, specify the bucket name. For example:
ceph osd crush remove {bucket-name}
Or:
ceph osd crush rm {bucket-name}
Note: The bucket must be empty in order to remove it.
If you are removing higher level buckets (for example, a root like default), check to see if a pool uses a CRUSH rule that selects that bucket. If so, you need to modify
your CRUSH rules; otherwise, peering fails.
CRUSH Bucket algorithms
When you create buckets using the Ceph CLI, Ceph sets the algorithm to straw2 by default. Ceph supports four bucket algorithms, each representing a tradeoff between
performance and reorganization efficiency. If you are unsure of which bucket type to use, we recommend using a straw2 bucket. The bucket algorithms are:
1. Uniform: Uniform buckets aggregate devices with exactly the same weight. For example, when firms commission or decommission hardware, they typically do so
with many machines that have exactly the same physical configuration (for example, bulk purchases). When storage devices have exactly the same weight, you can
use the uniform bucket type, which allows CRUSH to map replicas into uniform buckets in constant time. With non-uniform weights, you should use another
bucket algorithm.
2. List: List buckets aggregate their content as linked lists. Based on the RUSH (Replication Under Scalable Hashing) P algorithm, a list is a natural and intuitive choice
for an expanding cluster: either an object is relocated to the newest device with some appropriate probability, or it remains on the older devices as before. The
result is optimal data migration when items are added to the bucket. Items removed from the middle or tail of the list, however, can result in a significant amount of
unnecessary movement, making list buckets most suitable for circumstances in which they never, or very rarely shrink.
3. Tree: Tree buckets use a binary search tree. They are more efficient than listing buckets when a bucket contains a larger set of items. Based on the RUSH
(Replication Under Scalable Hashing) R algorithm, tree buckets reduce the placement time to zero (log n), making them suitable for managing much larger sets of
devices or nested buckets.
4. Straw2 (default): List and Tree buckets use a divide and conquer strategy in a way that either gives certain items precedence, for example, those at the beginning
of a list or obviates the need to consider entire subtrees of items at all. That improves the performance of the replica placement process, but can also introduce
suboptimal reorganization behavior when the contents of a bucket change due an addition, removal, or re-weighting of an item. The straw2 bucket type allows all
items to fairly “compete” against each other for replica placement through a process analogous to a draw of straws.
Ceph OSDs in CRUSH
Once you have a CRUSH hierarchy for the OSDs, add OSDs to the CRUSH hierarchy. You can also move or remove OSDs from an existing hierarchy. The Ceph CLI usage has
the following values:
id
Description
The numeric ID of the OSD.
Type
Integer
Required
Yes
Example
0
name
Description
84 IBM Storage Ceph
The full name of the OSD.
Type
String
Required
Yes
Example
osd.0
weight
Description
The CRUSH weight for the OSD.
Type
Double
Required
Yes
Example
2.0
root
Description
The name of the root bucket of the hierarchy or tree in which the OSD resides.
Type
Key-value pair.
Required
Yes
Example
root=default, root=replicated_rule, and so on
bucket-type
Description
One or more name-value pairs, where the name is the bucket type and the value is the bucket’s name. You can specify a CRUSH location for an OSD in the CRUSH
hierarchy.
Type
Key-value pairs.
Required
No
Example
datacenter=dc1 room=room1 row=foo rack=bar host=foo-bar-1
Viewing OSDs in CRUSH
Adding an OSD to CRUSH
Moving an OSD within a CRUSH Hierarchy
If the storage cluster topology changes, you can move an OSD in the CRUSH hierarchy to reflect its actual location.
Removing an OSD from a CRUSH Hierarchy
Viewing OSDs in CRUSH
The ceph osd crush tree command prints CRUSH buckets and items in a tree view. Use this command to determine a list of OSDs in a particular bucket. It will print
output similar to ceph osd tree.
To return additional details, execute the following:
# ceph osd crush tree -f json-pretty
The command returns an output similar to the following:
[
{
"id": -2,
"name": "ssd",
"type": "root",
"type_id": 10,
"items": [
{
"id": -6,
"name": "dell-per630-11-ssd",
"type": "host",
"type_id": 1,
"items": [
{
"id": 6,
"name": "osd.6",
"type": "osd",
"type_id": 0,
IBM Storage Ceph 85
},
{
},
{
},
{
]
}
]
}
"crush_weight": 0.099991,
"depth": 2
"id": -7,
"name": "dell-per630-12-ssd",
"type": "host",
"type_id": 1,
"items": [
{
"id": 7,
"name": "osd.7",
"type": "osd",
"type_id": 0,
"crush_weight": 0.099991,
"depth": 2
}
]
"id": -8,
"name": "dell-per630-13-ssd",
"type": "host",
"type_id": 1,
"items": [
{
"id": 8,
"name": "osd.8",
"type": "osd",
"type_id": 0,
"crush_weight": 0.099991,
"depth": 2
}
]
"id": -1,
"name": "default",
"type": "root",
"type_id": 10,
"items": [
{
"id": -3,
"name": "dell-per630-11",
"type": "host",
"type_id": 1,
"items": [
{
"id": 0,
"name": "osd.0",
"type": "osd",
"type_id": 0,
"crush_weight": 0.449997,
"depth": 2
},
{
"id": 3,
"name": "osd.3",
"type": "osd",
"type_id": 0,
"crush_weight": 0.289993,
"depth": 2
}
]
},
{
"id": -4,
"name": "dell-per630-12",
"type": "host",
"type_id": 1,
"items": [
{
"id": 1,
"name": "osd.1",
"type": "osd",
"type_id": 0,
"crush_weight": 0.449997,
"depth": 2
},
{
"id": 4,
"name": "osd.4",
"type": "osd",
"type_id": 0,
"crush_weight": 0.289993,
"depth": 2
}
]
},
{
"id": -5,
"name": "dell-per630-13",
86 IBM Storage Ceph
]
}
]
}
"type": "host",
"type_id": 1,
"items": [
{
"id": 2,
"name": "osd.2",
"type": "osd",
"type_id": 0,
"crush_weight": 0.449997,
"depth": 2
},
{
"id": 5,
"name": "osd.5",
"type": "osd",
"type_id": 0,
"crush_weight": 0.289993,
"depth": 2
}
]
Adding an OSD to CRUSH
Adding a Ceph OSD to a CRUSH hierarchy is the final step before you might start an OSD (rendering it up and in) and Ceph assigns placement groups to the OSD.
You must prepare a Ceph OSD before you add it to the CRUSH hierarchy. Deployment utilities, such as the Ceph Orchestrator, can perform this step for you. For example
creating a Ceph OSD on a single node:
Syntax
ceph orch daemon add osd HOST:DEVICE,[DEVICE]
The CRUSH hierarchy is notional, so the ceph osd crush add command allows you to add OSDs to the CRUSH hierarchy wherever you wish. The location you specify
should reflect its actual location. If you specify at least one bucket, the command places the OSD into the most specific bucket you specify, and it moves that bucket
underneath any other buckets you specify.
To add an OSD to a CRUSH hierarchy:
Syntax
ceph osd crush add ID_OR_NAME WEIGHT [BUCKET_TYPE=BUCKET_NAME ...]
Important: If you specify only the root bucket, the command attaches the OSD directly to the root. However, CRUSH rules expect OSDs to be inside of hosts or chassis, and
host or chassis should be inside of other buckets reflecting your cluster topology.
If you specify only the root bucket, the command attaches the OSD directly to the root. However, CRUSH rules expect OSDs to be inside of hosts or chassis, and host or
chassis should be inside of other buckets reflecting your cluster topology.
The following example adds osd.0 to the hierarchy:
ceph osd crush add osd.0 1.0 root=default datacenter=dc1 room=room1 row=foo rack=bar host=foo-bar-1
Note: You can also use ceph osd crush set or ceph osd crush
create-or-move to add an OSD to the CRUSH hierarchy.
Moving an OSD within a CRUSH Hierarchy
If the storage cluster topology changes, you can move an OSD in the CRUSH hierarchy to reflect its actual location.
Important: Moving an OSD in the CRUSH hierarchy means that Ceph will recompute which placement groups get assigned to the OSD, potentially resulting in significant
redistribution of data
To move an OSD within the CRUSH hierarchy:
Syntax
ceph osd crush set ID_OR_NAME WEIGHT root=POOL_NAME
[BUCKET_TYPE=BUCKET_NAME...]
Note: You can also use ceph osd crush create-or-move to move an OSD within the CRUSH hierarchy.
Removing an OSD from a CRUSH Hierarchy
Removing an OSD from a CRUSH hierarchy is the first step when you want to remove an OSD from your cluster. When you remove the OSD from the CRUSH map, CRUSH
recomputes which OSDs get the placement groups and data re-balances accordingly. See Adding/Removing OSDs for additional details.
To remove an OSD from the CRUSH map of a running cluster, execute the following:
Syntax
ceph osd crush remove NAME
IBM Storage Ceph 87
Device class
The Ceph CRUSH map provides a lot of flexibility in controlling data placement.
The flexibility in controlling data placement is one of Ceph’s greatest strengths. Early Ceph deployments used hard disk drives almost exclusively. Today, Ceph clusters are
frequently built with multiple types of storage devices: HDD, SSD, NVMe, or even various classes of the foregoing. For example, it is common in Ceph Object Gateway
deployments to have storage policies where clients can store data on slower HDDs and other storage policies for storing data on fast SSDs. Ceph Object Gateway
deployments might even have a pool backed by fast SSDs for bucket indices. Additionally, OSD nodes also frequently have SSDs used exclusively for journals or writeahead logs that do NOT appear in the CRUSH map. These complex hardware scenarios historically required manually editing the CRUSH map, which can be timeconsuming and tedious. It is not required to have different CRUSH hierarchies for different classes of storage devices.
CRUSH rules work in terms of the CRUSH hierarchy. However, if different classes of storage devices reside in the same hosts, the process becomes more complicated—
requiring users to create multiple CRUSH hierarchies for each class of device, and then disable the osd crush update
on start option that automates much of the CRUSH hierarchy management. Device classes eliminate this tediousness by telling the CRUSH rule what class of device to
use, dramatically simplifying CRUSH management tasks.
Note: The ceph osd tree command has a column reflecting a device class.
Setting a device class
Removing a device class
Renaming a device class
Listing a device class
Listing OSDs of a device class
Listing CRUSH Rules by Class
Reference
For more information, see Using different device classes and CRUSH storage strategies examples.
Setting a device class
To set a device class for an OSD, execute the following:
Syntax
ceph osd crush set-device-class
CLASS OSD_ID [OSD_ID..]
Example
[ceph: root@host01 /]# ceph osd crush set-device-class hdd osd.0 osd.1
[ceph: root@host01 /]# ceph osd crush set-device-class ssd osd.2 osd.3
[ceph: root@host01 /]# ceph osd crush set-device-class bucket-index osd.4
Note: Ceph might assign a class to a device automatically. However, class names are simply arbitrary strings. There is no requirement to adhere to hdd, ssd or nvme. In
the foregoing example, a device class named bucket-index might indicate an SSD device that a Ceph Object Gateway pool uses exclusively bucket index workloads. To
change a device class that was already set, use ceph osd crush
rm-device-class first.
Removing a device class
To remove a device class for an OSD, execute the following:
Syntax
ceph osd crush rm-device-class CLASS OSD_ID [OSD_ID..]
Example
[ceph: root@host01 /]# ceph osd crush rm-device-class hdd osd.0 osd.1
[ceph: root@host01 /]# ceph osd crush rm-device-class ssd osd.2 osd.3
[ceph: root@host01 /]# ceph osd crush rm-device-class bucket-index osd.4
Renaming a device class
To rename a device class for all OSDs that use that class, execute the following:
Syntax
ceph osd crush class rename OLD_NAME NEW_NAME
Example
[ceph: root@host01 /]# ceph osd crush class rename hdd sas15k
88 IBM Storage Ceph
Listing a device class
To list device classes in the CRUSH map, execute the following:
Syntax
ceph osd crush class ls
The output will look something like this:
Example
[
]
"hdd",
"ssd",
"bucket-index"
Listing OSDs of a device class
To list all OSDs that belong to a particular class, execute the following:
Syntax
ceph osd crush class ls-osd CLASS
Example
[ceph: root@host01 /]# ceph osd crush class ls-osd hdd
The output is simply a list of OSD numbers. For example:
0
1
2
3
4
5
6
Listing CRUSH Rules by Class
To list all crush rules that reference the same class, execute the following:
Syntax
ceph osd crush rule ls-by-class CLASS
Example
[ceph: root@host01 /]# ceph osd crush rule ls-by-class hdd
CRUSH weights
The CRUSH algorithm assigns a weight value in terabytes (by convention) per OSD device with the objective of approximating a uniform probability distribution for write
requests that assign new data objects to PGs and PGs to OSDs. For this reason, as a best practice, we recommend creating CRUSH hierarchies with devices of the same
type and size, and assigning the same weight. We also recommend using devices with the same I/O and throughput characteristics so that you will also have uniform
performance characteristics in your CRUSH hierarchy, even though performance characteristics do not affect data distribution.
Since using uniform hardware is not always practical, you might incorporate OSD devices of different sizes and use a relative weight so that Ceph will distribute more data
to larger devices and less data to smaller devices.
Setting CRUSH weights of OSDs
Setting a Bucket’s OSD Weights
Set an OSD’s in Weight
Setting the OSDs weight by utilization
Setting an OSD’s Weight by PG distribution
Recalculating a CRUSH Tree’s weights
Setting CRUSH weights of OSDs
To set an OSD CRUSH weight in Terabytes within the CRUSH map, execute the following command
ceph osd crush reweight NAME WEIGHT
IBM Storage Ceph 89
Where:
name
Description
The full name of the OSD.
Type
String
Required
Yes
Example
osd.0
weight
Description
The CRUSH weight for the OSD. This should be the size of the OSD in Terabytes, where 1.0 is 1 Terabyte.
Type
Double
Required
Yes
Example
2.0
This setting is used when creating an OSD or adjusting the CRUSH weight immediately after adding the OSD. It usually does not change over the life of the OSD.
Setting a Bucket’s OSD Weights
Using ceph osd crush reweight can be time-consuming. You can set (or reset) all Ceph OSD weights under a bucket (row, rack, node, and so on) by executing:
Syntax
osd crush reweight-subtree NAME
Where,
name is the name of the CRUSH bucket.
Set an OSD’s in Weight
For the purposes of ceph osd in and ceph osd out, an OSD is either in the cluster or out of the cluster. That is how a monitor records an OSD’s status. However,
even though an OSD is in the cluster, it might be experiencing a malfunction such that you do not want to rely on it as much until you fix it (for example, replace a storage
drive, change out a controller, and so on).
You can increase or decrease the in weight of a particular OSD (that is, without changing its weight in Terabytes) by executing:
Syntax
ceph osd reweight ID WEIGHT
Where:
id is the OSD number.
weight is a range from 0.0-1.0, where 0 is not in the cluster (that is, it does not have any PGs assigned to it) and 1.0 is in the cluster (that is, the OSD receives the
same number of PGs as other OSDs).
Setting the OSDs weight by utilization
CRUSH is designed to approximate a uniform probability distribution for write requests that assign new data objects PGs and PGs to OSDs. However, a cluster might
become imbalanced anyway. This can happen for a number of reasons. For example:
Multiple Pools: You can assign multiple pools to a CRUSH hierarchy, but the pools might have different numbers of placement groups, size (number of replicas to
store), and object size characteristics.
Custom Clients: Ceph clients such as block device, object gateway and filesystem share data from their clients and stripe the data as objects across the cluster as
uniform-sized smaller RADOS objects. So except for the foregoing scenario, CRUSH usually achieves its goal. However, there is another case where a cluster can
become imbalanced: namely, using librados to store data without normalizing the size of objects. This scenario can lead to imbalanced clusters (for example,
storing 100 1 MB objects and 10 4 MB objects will make a few OSDs have more data than the others).
Probability: A uniform distribution will result in some OSDs with more PGs and some with less. For clusters with a large number of OSDs, the statistical outliers will
be further out.
You can reweight OSDs by utilization by executing the following:
90 IBM Storage Ceph
Syntax
ceph osd reweight-by-utilization [THRESHOLD] [WEIGHT_CHANGE_AMOUNT] [NUMBER_OF_OSDS] [--no-increasing]
Example
[ceph: root@host01 /]# ceph osd test-reweight-by-utilization 110 .5 4 --no-increasing
Where:
threshold is a percentage of utilization such that OSDs facing higher data storage loads will receive a lower weight and thus fewer PGs assigned to them. The
default value is 120, reflecting 120%. Any value from 100+ is a valid threshold. Optional.
weight_change_amount is the amount to change the weight. Valid values are greater than 0.0 - 1.0. The default value is 0.05. Optional.
number_of_OSDs is the maximum number of OSDs to reweight. For large clusters, limiting the number of OSDs to reweight prevents significant rebalancing.
Optional.
no-increasing is off by default. Increasing the osd weight is allowed when using the reweight-by-utilization or test-reweight-by-utilization
commands. If this option is used with these commands, it prevents the OSD weight from increasing, even if the OSD is underutilized. Optional.
Important: Executing reweight-by-utilization is recommended and somewhat inevitable for large clusters. Utilization rates might change over time, and as your
cluster size or hardware changes, the weightings might need to be updated to reflect changing utilization. If you elect to reweight by utilization, you might need to re-run
this command as utilization, hardware or cluster size change.
Executing this or other weight commands that assign a weight will override the weight assigned by this command (for example, osd reweight-by-utilization, osd
crush
weight, osd weight, in or out).
Setting an OSD’s Weight by PG distribution
In CRUSH hierarchies with a smaller number of OSDs, it’s possible for some OSDs to get more PGs than other OSDs, resulting in a higher load. You can reweight OSDs by
PG distribution to address this situation by executing the following:
Syntax
osd reweight-by-pg POOL_NAME
Where:
poolname is the name of the pool. Ceph will examine how the pool assigns PGs to OSDs and reweight the OSDs according to this pool’s PG distribution. Note that
multiple pools could be assigned to the same CRUSH hierarchy. Reweighting OSDs according to one pool’s distribution could have unintended effects for other
pools assigned to the same CRUSH hierarchy if they do not have the same size (number of replicas) and PGs.
Recalculating a CRUSH Tree’s weights
CRUSH tree buckets should be the sum of their leaf weights. If you manually edit the CRUSH map weights, you should execute the following to ensure that the CRUSH
bucket tree accurately reflects the sum of the leaf OSDs under the bucket.
Syntax
osd crush reweight-all
Primary affinity
When a Ceph Client reads or writes data, it always contacts the primary OSD in the acting set. For set [2, 3, 4], osd.2 is the primary. Sometimes an OSD is not well
suited to act as a primary compared to other OSDs (for example, it has a slow disk or a slow controller). To prevent performance bottlenecks (especially on read
operations) while maximizing utilization of your hardware, you can set a Ceph OSD’s primary affinity so that CRUSH is less likely to use the OSD as a primary in an acting
set. :
Syntax
ceph osd primary-affinity OSD_ID WEIGHT
Primary affinity is 1 by default (that is, an OSD might act as a primary). You might set the OSD primary range from 0-1, where 0 means that the OSD might NOT be used as
a primary and 1 means that an OSD might be used as a primary. When the weight is < 1, it is less likely that CRUSH will select the Ceph OSD Daemon to act as a primary.
CRUSH rules
CRUSH rules define how a Ceph client selects buckets and the primary OSD within them to store objects, and how the primary OSD selects buckets and the secondary
OSDs to store replicas or coding chunks. For example, you might create a rule that selects a pair of target OSDs backed by SSDs for two object replicas, and another rule
that selects three target OSDs backed by SAS drives in different data centers for three replicas.
A rule takes the following form:
IBM Storage Ceph 91
rule <rulename> {
}
id <unique number>
type [replicated | erasure]
min_size <min-size>
max_size <max-size>
step take <bucket-type> [class <class-name>]
step [choose|chooseleaf] [firstn|indep] <N> <bucket-type>
step emit
id
Description
A unique whole number for identifying the rule.
Purpose
A component of the rule mask.
Type
Integer
Required
Yes
Default
0
type
Description
Describes a rule for either a storage drive replicated or erasure coded.
Purpose
A component of the rule mask.
Type
String
Required
Yes
Default
replicated
Valid Values
Currently only replicated
min_size
Description
If a pool makes fewer replicas than this number, CRUSH will not select this rule.
Type
Integer
Purpose
A component of the rule mask.
Required
Yes
Default
1
max_size
Description
If a pool makes more replicas than this number, CRUSH will not select this rule.
Type
Integer
Purpose
A component of the rule mask.
Required
Yes
Default
10
step take <bucket-name> [class
<class-name>]
Description
Takes a bucket name, and begins iterating down the tree.
Purpose
A component of the rule.
Required
Yes
92 IBM Storage Ceph
Example
step take data step take data class ssd
step choose firstn <num> type
<bucket-type>
Description
Selects the number of buckets of the given type. The number is usually the number of replicas in the pool (that is, pool size).
If <num> == 0, choose pool-num-replicas buckets (all available).
If <num> > 0 && < pool-num-replicas, choose that many buckets.
If <num> < 0, it means pool-num-replicas - {num}.
Purpose
A component of the rule.
Prerequisite
Follow step take or step
choose.
Example
step choose firstn 1 type row
step chooseleaf firstn <num> type
<bucket-type>
Description
Selects a set of buckets of {bucket-type} and chooses a leaf node from the subtree of each bucket in the set of buckets. The number of buckets in the set is usually the
number of replicas in the pool (that is, pool size).
If <num> == 0, choose pool-num-replicas buckets (all available).
If <num> > 0 && < pool-num-replicas, choose that many buckets.
If <num> < 0, it means pool-num-replicas <num>.
Purpose
A component of the rule. Usage removes the need to select a device using two steps.
Prerequisite
Follows step take or step
choose.
Example
step chooseleaf firstn 0 type row
step emit
Description
Outputs the current value and empties the stack. Typically used at the end of a rule, but might also be used to pick from different trees in the same rule.
Purpose
A component of the rule.
Prerequisite
Follows step choose.
Example
step emit
firstn versus indep
Description
Controls the replacement strategy CRUSH uses when OSDs are marked down in the CRUSH map. If this rule is to be used with replicated pools it should be firstn and if
it is for erasure-coded pools it should be indep.
Example
You have a PG stored on OSDs 1, 2, 3, 4, 5 in which 3 goes down.. In the first scenario, with the firstn mode, CRUSH adjusts its calculation to select 1 and 2, then selects
3 but discovers it is down, so it retries and selects 4 and 5, and then goes on to select a new OSD 6. The final CRUSH mapping change is from 1, 2, 3, 4, 5 to 1, 2, 4, 5, 6. In
the second scenario, with indep mode on an erasure-coded pool, CRUSH attempts to select the failed OSD 3, tries again and picks out 6, for a final transformation from 1,
2, 3, 4, 5 to 1, 2, 6, 4, 5.
Important: A given CRUSH rule can be assigned to multiple pools, but it is not possible for a single pool to have multiple CRUSH rules.
Listing CRUSH rules
Dumping CRUSH rules
Adding CRUSH rules
Creating CRUSH rules for replicated pools
Creating CRUSH rules for erasure coded pools
Removing CRUSH rules
Listing CRUSH rules
To list CRUSH rules from the command line, execute the following:
Syntax
IBM Storage Ceph 93
ceph osd crush rule list
ceph osd crush rule ls
Dumping CRUSH rules
To dump the contents of a specific CRUSH rule, execute the following:
Syntax
ceph osd crush rule dump NAME
Adding CRUSH rules
To add a CRUSH rule, you must specify a rule name, the root node of the hierarchy you wish to use, the type of bucket you want to replicate across (for example, rack, row,
and so on and the mode for choosing the bucket.
Syntax
ceph osd crush rule create-simple RUENAME ROOT BUCKET_NAME FIRSTN_OR_INDEP
Ceph creates a rule with chooseleaf and one bucket of the type you specify.
Example
[ceph: root@host01 /]# ceph osd crush rule create-simple deleteme default host firstn
Create the following rule:
{ "id": 1,
"rule_name": "deleteme",
"type": 1,
"min_size": 1,
"max_size": 10,
"steps": [
{ "op": "take",
"item": -1,
"item_name": "default"},
{ "op": "chooseleaf_firstn",
"num": 0,
"type": "host"},
{ "op": "emit"}]}
Creating CRUSH rules for replicated pools
To create a CRUSH rule for a replicated pool, execute the following:
Syntax
ceph osd crush rule create-replicated NAME ROOT FAILURE_DOMAIN CLASS
Where:
<name>: The name of the rule.
<root>: The root of the CRUSH hierarchy.
<failure-domain>: The failure domain. For example: host or rack.
<class>: The storage device class. For example: hdd or ssd.
Example
[ceph: root@host01 /]# ceph osd crush rule create-replicated fast default host ssd
Creating CRUSH rules for erasure coded pools
To add a CRUSH rule for use with an erasure coded pool, you might specify a rule name and an erasure code profile.
Syntax
ceph osd crush rule create-erasure RULE_NAME PROFILE_NAME
For example,
[ceph: root@host01 /]# ceph osd crush rule create-erasure default default
See Erasure code profiles for more information.
94 IBM Storage Ceph
Removing CRUSH rules
To remove a rule, execute the following and specify the CRUSH rule name:
Syntax
ceph osd crush rule rm NAME
CRUSH tunables overview
The Ceph project has grown exponentially with many changes and many new features. Beginning with the first commercially supported major release of Ceph, v0.48
(Argonaut), Ceph provides the ability to adjust certain parameters of the CRUSH algorithm, that is, the settings are not frozen in the source code.
A few important points to consider:
Adjusting CRUSH values might result in the shift of some PGs between storage nodes. If the Ceph cluster is already storing a lot of data, be prepared for some
fraction of the data to move.
The ceph-osd and ceph-mon daemons will start requiring the feature bits of new connections as soon as they receive an updated map. However, alreadyconnected clients are effectively grandfathered in, and will misbehave if they do not support the new feature. Make sure when you upgrade your Ceph Storage
Cluster daemons that you also update your Ceph clients.
If the CRUSH tunables are set to non-legacy values and then later changed back to the legacy values, ceph-osd daemons will not be required to support the
feature. However, the OSD peering process requires examining and understanding old maps. Therefore, you should not run old versions of the ceph-osd daemon if
the cluster has previously used non-legacy CRUSH values, even if the latest version of the map has been switched back to using the legacy defaults.
CRUSH tuning
CRUSH tuning, the hard way
CRUSH legacy values
CRUSH tuning
Before you tune CRUSH, you should ensure that all Ceph clients and all Ceph daemons use the same version. If you have recently upgraded, ensure that you have
restarted daemons and reconnected clients.
The simplest way to adjust the CRUSH tunables is by changing to a known profile. Those are:
legacy: The legacy behavior from v0.47 (pre-Argonaut) and earlier.
argonaut: The legacy values supported by v0.48 (Argonaut) release.
bobtail: The values supported by the v0.56 (Bobtail) release.
firefly: The values supported by the v0.80 (Firefly) release.
hammer: The values supported by the v0.94 (Hammer) release.
jewel: The values supported by the v10.0.2 (Jewel) release.
optimal: The current best values.
default: The current default values for a new cluster.
You can select a profile on a running cluster with the command:
Syntax
# ceph osd crush tunables PROFILE
Note: This might result in some data movement.
Generally, you should set the CRUSH tunables after you upgrade, or if you receive a warning. Starting with version v0.74, Ceph issues a health warning if the CRUSH
tunables are not set to their optimal values, the optimal values are the default as of v0.73.
To make this warning go away, you have two options:
1. Adjust the tunables on the existing cluster. Note that this will result in some data movement (possibly as much as 10%). This is the preferred route, but should be
taken with care on a production cluster where the data movement might affect performance. You can enable optimal tunables with:
# ceph osd crush tunables optimal
If things go poorly (for example, too much load) and not very much progress has been made, or there is a client compatibility problem (old kernel cephfs or rbd
clients, or pre-bobtail librados clients), you can switch back to an earlier profile:
# ceph osd crush tunables <profile>
For example, to restore the pre-v0.48 (Argonaut) values, execute:
# ceph osd crush tunables legacy
2. You can make the warning go away without making any changes to CRUSH by adding the following option to the mon section of the ceph.conf file:
IBM Storage Ceph 95
# mon warn on legacy crush tunables = false
For the change to take effect, restart the monitors, or apply the option to running monitors with:
# ceph tell mon.\* injectargs --no-mon-warn-on-legacy-crush-tunables
CRUSH tuning, the hard way
If you can ensure that all clients are running recent code, you can adjust the tunables by extracting the CRUSH map, modifying the values, and reinjecting it into the
cluster.
Extract the latest CRUSH map:
ceph osd getcrushmap -o /tmp/crush
Adjust tunables. These values appear to offer the best behavior for both large and small clusters we tested with. You will need to additionally specify the -enable-unsafe-tunables argument to crushtool for this to work. Please use this option with extreme care.:
crushtool -i /tmp/crush --set-choose-local-tries 0 --set-choose-local-fallback-tries 0 --set-choose-total-tries 50 -o
/tmp/crush.new
Reinject modified map:
ceph osd setcrushmap -i /tmp/crush.new
CRUSH legacy values
For reference, the legacy values for the CRUSH tunables can be set with:
crushtool -i /tmp/crush --set-choose-local-tries 2 --set-choose-local-fallback-tries 5 --set-choose-total-tries 19 --setchooseleaf-descend-once 0 --set-chooseleaf-vary-r 0 -o /tmp/crush.legacy
Again, the special --enable-unsafe-tunables option is required. Further, as noted above, be careful running old versions of the ceph-osd daemon after reverting to
legacy values as the feature bit is not perfectly enforced.
Edit a CRUSH map
Generally, modifying your CRUSH map at runtime with the Ceph CLI is more convenient than editing the CRUSH map manually. However, there are times when you might
choose to edit it, such as changing the default bucket types, or using a bucket algorithm other than straw2.
To edit an existing CRUSH map:
1. Get the CRUSH map, as described in Getting the CRUSH map.
2. Decompile the CRUSH map, as described in Decompiling the CRUSH map .
3. Edit at least one of the devices, and buckets and rules.
4. Recompile the CRUSH map, as described in Compiling the CRUSH map.
5. Set the CRUSH map, as described in Setting a CRUSH map.
To activate a CRUSH Map rule for a specific pool, identify the common rule number and specify that rule number for the pool when creating the pool.
Getting the CRUSH map
Decompiling the CRUSH map
Compiling the CRUSH map
Setting a CRUSH map
Getting the CRUSH map
To get the CRUSH map for your cluster, execute the following:
Syntax
ceph osd getcrushmap -o COMPILED_CRUSHMAP_FILENAME
Ceph will output (-o) a compiled CRUSH map to the file name you specified. Since the CRUSH map is in a compiled form, you must decompile it first before you can edit it.
Decompiling the CRUSH map
To decompile a CRUSH map, execute the following:
Syntax
crushtool -d COMPILED_CRUSHMAP_FILENAME -o DECOMPILED_CRUSHMAP_FILENAME
96 IBM Storage Ceph
Ceph decompiles (-d) the compiled CRUSH map and send the output (-o) to the file name you specified.
Compiling the CRUSH map
To compile a CRUSH map, execute the following:
Syntax
crushtool -c DECOMPILED_CRUSHMAP_FILENAME -o COMPILED_CRUSHMAP_FILENAME
Ceph will store a compiled CRUSH map to the file name you specified.
Setting a CRUSH map
To set the CRUSH map for your cluster, execute the following:
Syntax
ceph osd setcrushmap -i COMPILED_CRUSHMAP_FILENAME
Ceph inputs the compiled CRUSH map of the file name you specified as the CRUSH map for the cluster.
CRUSH storage strategies examples
If you want to have most pools default to OSDs backed by large hard drives, but have some pools mapped to OSDs backed by fast solid-state drives (SSDs). CRUSH can
handle these scenarios easily.
Use device classes. The process is simple to add a class to each device.
Syntax
ceph osd crush set-device-class CLASS OSD_ID [OSD_ID]
Example
[ceph:root@host01 /]# ceph osd crush set-device-class hdd osd.0 osd.1 osd.4 osd.5
[ceph:root@host01 /]# ceph osd crush set-device-class ssd osd.2 osd.3 osd.6 osd.7
Then, create rules to use the devices.
Syntax
ceph osd crush rule create-replicated RULENAME ROOT FAILURE_DOMAIN_TYPE DEVICE_CLASS
Example
[ceph:root@host01 /]# ceph osd crush rule create-replicated cold default host hdd
[ceph:root@host01 /]# ceph osd crush rule create-replicated hot default host ssd
Finally, set pools to use the rules.
Syntax
ceph osd pool set POOL_NAME crush_rule RULENAME
Example
[ceph:root@host01 /]# ceph osd pool set cold crush_rule hdd
[ceph:root@host01 /]# ceph osd pool set hot crush_rule ssd
There is no need to manually edit the CRUSH map, because one hierarchy can serve multiple classes of devices.
device 0 osd.0 class hdd
device 1 osd.1 class hdd
device 2 osd.2 class ssd
device 3 osd.3 class ssd
device 4 osd.4 class hdd
device 5 osd.5 class hdd
device 6 osd.6 class ssd
device 7 osd.7 class ssd
host ceph-osd-server-1 {
id -1
alg straw2
hash 0
item osd.0 weight 1.00
item osd.1 weight 1.00
item osd.2 weight 1.00
item osd.3 weight 1.00
}
host ceph-osd-server-2 {
id -2
IBM Storage Ceph 97
}
alg straw2
hash 0
item osd.4 weight 1.00
item osd.5 weight 1.00
item osd.6 weight 1.00
item osd.7 weight 1.00
root default {
id -3
alg straw2
hash 0
item ceph-osd-server-1 weight 4.00
item ceph-osd-server-2 weight 4.00
}
rule cold {
ruleset 0
type replicated
min_size 2
max_size 11
step take default class hdd
step chooseleaf firstn 0 type host
step emit
}
rule hot {
ruleset 1
type replicated
min_size 2
max_size 11
step take default class ssd
step chooseleaf firstn 0 type host
step emit
}
Placement Groups
Placement Groups (PGs) are invisible to Ceph clients, but they play an important role in Ceph Storage Clusters.
A Ceph Storage Cluster might require many thousands of OSDs to reach an exabyte level of storage capacity. Ceph clients store objects in pools, which are a logical subset
of the overall cluster. The number of objects stored in a pool might easily run into the millions and beyond. A system with millions of objects or more cannot realistically
track placement on a per-object basis and still perform well. Ceph assigns objects to placement groups, and placement groups to OSDs to make re-balancing dynamic and
efficient.
All problems in computer science can be solved by another level of indirection, except of course for the problem of too many indirections.
— David Wheeler
About placement groups
Placement group states
Placement group tradeoffs
Placement group count
Auto-scaling placement groups
The number of placement groups (PGs) in a pool plays a significant role in how a cluster peers, distributes data, and rebalances.
Updating noautoscale flag
Specifying target pool size
Specify target pool size.
Placement group command line interface
About placement groups
Tracking object placement on a per-object basis within a pool is computationally expensive at scale. To facilitate high performance at scale, Ceph subdivides a pool into
placement groups, assigns each individual object to a placement group, and assigns the placement group to a primary OSD. If an OSD fails or the cluster re-balances,
Ceph can move or replicate an entire placement group—that is, all of the objects in the placement groups—without having to address each object individually. This allows a
Ceph cluster to re-balance or recover efficiently.
Figure 1. About PGs
98 IBM Storage Ceph
When CRUSH assigns a placement group to an OSD, it calculates a series of OSDs—the first being the primary. The osd_pool_default_size setting minus 1 for
replicated pools, and the number of coding chunks M for erasure-coded pools determine the number of OSDs storing a placement group that can fail without losing data
permanently. Primary OSDs use CRUSH to identify the secondary OSDs and copy the placement group’s contents to the secondary OSDs. For example, if CRUSH assigns an
object to a placement group, and the placement group is assigned to OSD 5 as the primary OSD, if CRUSH calculates that OSD 1 and OSD 8 are secondary OSDs for the
placement group, the primary OSD 5 will copy the data to OSDs 1 and 8. By copying data on behalf of clients, Ceph simplifies the client interface and reduces the client
workload. The same process allows the Ceph cluster to recover and rebalance dynamically.
Figure 2. CRUSH hierarchy
When the primary OSD fails and gets marked out of the cluster, CRUSH assigns the placement group to another OSD, which receives copies of objects in the placement
group. Another OSD in the Up Set will assume the role of the primary OSD.
When you increase the number of object replicas or coding chunks, CRUSH will assign each placement group to additional OSDs as required.
Note: PGs do not own OSDs. CRUSH assigns many placement groups to each OSD pseudo-randomly to ensure that data gets distributed evenly across the cluster.
Placement group states
When you check the storage cluster’s status with the ceph -s or ceph
-w commands, Ceph reports on the status of the placement groups (PGs). A PG has one or more states. The optimum state for PGs in the PG map is an active + clean
state.
activating
The PG is peered, but not yet active.
active
Ceph processes requests to the PG.
backfill_toofull
A backfill operation is waiting because the destination OSD is over the backfillfull ratio.
backfill_unfound
Backfill stopped due to unfound objects.
backfill_wait
The PG is waiting in line to start backfill.
IBM Storage Ceph 99
backfilling
Ceph is scanning and synchronizing the entire contents of a PG instead of inferring what contents need to be synchronized from the logs of recent operations. Backfill is a
special case of recovery.
clean
Ceph replicated all objects in the PG accurately.
creating
Ceph is still creating the PG.
deep
Ceph is checking the PG data against stored checksums.
degraded
Ceph has not replicated some objects in the PG accurately yet.
down
A replica with necessary data is down, so the PG is offline. A PG with less than min_size replicas is marked as down. Use ceph health
detail to understand the backing OSD state.
forced_backfill
High backfill priority of that PG is enforced by user.
forced_recovery
High recovery priority of that PG is enforced by user.
incomplete
Ceph detects that a PG is missing information about writes that might have occurred, or does not have any healthy copies. If you see this state, try to start any failed OSDs
that might contain the needed information. In the case of an erasure coded pool, temporarily reducing min_size might allow recovery.
inconsistent
Ceph detects inconsistencies in one or more replicas of an object in the PG, such as objects are the wrong size, objects are missing from one replica after recovery
finished.
peering
The PG is undergoing the peering process. A peering process should clear off without much delay, but if it stays and the number of PGs in a peering state does not reduce
in number, the peering might be stuck.
peered
The PG has peered, but cannot serve client IO due to not having enough copies to reach the pool’s configured min_size parameter. Recovery might occur in this state, so
the PG might heal up to min_size eventually.
recovering
Ceph is migrating or synchronizing objects and their replicas.
recovery_toofull
A recovery operation is waiting because the destination OSD is over its full ratio.
recovery_unfound
Recovery stopped due to unfound objects.
recovery_wait
The PG is waiting in line to start recovery.
remapped
The PG is temporarily mapped to a different set of OSDs from what CRUSH specified.
repair
Ceph is checking the PG and repairing any inconsistencies it finds, if possible.
replay
The PG is waiting for clients to replay operations after an OSD crashed.
snaptrim
Trimming snaps.
snaptrim_error
Error stopped trimming snaps.
snaptrim_wait
Queued to trim snaps.
scrubbing
Ceph is checking the PG metadata for inconsistencies.
splitting
Ceph is splitting the PG into multiple PGs.
stale
The PG is in an unknown state; the monitors have not received an update for it since the PG mapping changed.
undersized
The PG has fewer copies than the configured pool replication level.
unknown
The ceph-mgr has not yet received any information about the PG’s state from an OSD since Ceph Manager started up.
100 IBM Storage Ceph
References
See the knowledge base What are the possible Placement Group states in an Ceph cluster for more information.
Placement group tradeoffs
Data durability and data distribution among all OSDs call for more placement groups but their number should be reduced to the minimum required for maximum
performance to conserve CPU and memory resources.
Data durability
Data distribution
Resource usage
Data durability
Ceph strives to prevent the permanent loss of data. However, after an OSD fails, the risk of permanent data loss increases until the data it had is fully recovered.
Permanent data loss, though rare, is still possible. The following scenario describes how Ceph could permanently lose data in a single placement group with three copies
of the data:
An OSD fails and all copies of the object it contains are lost. For all objects within a placement group stored on the OSD, the number of replicas suddenly drops from
three to two.
Ceph starts recovery for each placement group stored on the failed OSD by choosing a new OSD to re-create the third copy of all objects for each placement group.
The second OSD containing a copy of the same placement group fails before the new OSD is fully populated with the third copy. Some objects will then only have
one surviving copy.
Ceph picks yet another OSD and keeps copying objects to restore the desired number of copies.
The third OSD containing a copy of the same placement group fails before recovery is complete. If this OSD contained the only remaining copy of an object, the
object is lost permanently.
Hardware failure isn’t an exception, but an expectation. To prevent the foregoing scenario, ideally the recovery process should be as fast as reasonably possible. The size
of your cluster, your hardware configuration and the number of placement groups play an important role in total recovery time.
Small clusters don’t recover as quickly.
In a cluster containing 10 OSDs with 512 placement groups in a three replica pool, CRUSH will give each placement group three OSDs. Each OSD will end up hosting (512
* 3) / 10 =
~150 placement groups. When the first OSD fails, the cluster will start recovery for all 150 placement groups simultaneously.
It is likely that Ceph stored the remaining 150 placement groups randomly across the 9 remaining OSDs. Therefore, each remaining OSD is likely to send copies of objects
to all other OSDs and also receive some new objects, because the remaining OSDs become responsible for some of the 150 placement groups now assigned to them.
The total recovery time depends upon the hardware supporting the pool. For example, in a 10 OSD cluster, if a host contains one OSD with a 1 TB SSD, and a 10 GB/s
switch connects each of the 10 hosts, the recovery time will take M minutes. By contrast, if a host contains two SATA OSDs and a 1 GB/s switch connects the five hosts,
recovery will take substantially longer. Interestingly, in a cluster of this size, the number of placement groups has almost no influence on data durability. The placement
group count could be 128 or 8192 and the recovery would not be slower or faster.
However, growing the same Ceph cluster to 20 OSDs instead of 10 OSDs is likely to speed up recovery and therefore improve data durability significantly. Why? Each OSD
now participates in only 75 placement groups instead of 150. The 20 OSD cluster will still require all 19 remaining OSDs to perform the same amount of copy operations in
order to recover. In the 10 OSD cluster, each OSDs had to copy approximately 100 GB. In the 20 OSD cluster each OSD only has to copy 50 GB each. If the network was
the bottleneck, recovery will happen twice as fast. In other words, recovery time decreases as the number of OSDs increases.
In large clusters, PG count is important!
If the exemplary cluster grows to 40 OSDs, each OSD will only host 35 placement groups. If an OSD dies, recovery time will decrease unless another bottleneck precludes
improvement. However, if this cluster grows to 200 OSDs, each OSD will only host approximately 7 placement groups. If an OSD dies, recovery will happen between at
most of 21 (7 * 3) OSDs in these placement groups: recovery will take longer than when there were 40 OSDs, meaning the number of placement groups should be
increased!
Important: No matter how short the recovery time, there is a chance for another OSD storing the placement group to fail while recovery is in progress.
In the 10 OSD cluster described above, if any OSD fails, then approximately 8 placement groups (that is 75 pgs / 9 osds being recovered) will only have one surviving
copy. And if any of the 8 remaining OSDs fail, the last objects of one placement group are likely to be lost (that is 8 pgs / 8 osds with only one remaining copy being
recovered). This is why starting with a somewhat larger cluster is preferred (for example, 50 OSDs).
When the size of the cluster grows to 20 OSDs, the number of placement groups damaged by the loss of three OSDs drops. The second OSD lost will degrade
approximately 2 (that is 35 pgs / 19
osds being recovered) instead of 8 and the third OSD lost will only lose data if it is one of the two OSDs containing the surviving copy. In other words, if the probability of
losing one OSD is 0.0001% during the recovery time frame, it goes from 8 *
0.0001% in the cluster with 10 OSDs to 2 * 0.0001% in the cluster with 20 OSDs. Having 512 or 4096 placement groups is roughly equivalent in a cluster with less than
50 OSDs as far as data durability is concerned.
Tip: More OSDs means faster recovery and a lower risk of cascading failures leading to the permanent loss of a placement group and its objects.
When you add an OSD to the cluster, it might take a long time to populate the new OSD with placement groups and objects. However there is no degradation of any object
and adding the OSD has no impact on data durability.
Data distribution
IBM Storage Ceph 101
Ceph seeks to avoid hot spots—that is, some OSDs receive substantially more traffic than other OSDs. Ideally, CRUSH assigns objects to placement groups evenly so that
when the placement groups get assigned to OSDs (also pseudo randomly), the primary OSDs store objects such that they are evenly distributed across the cluster and hot
spots and network over-subscription problems cannot develop because of data distribution.
Since CRUSH computes the placement group for each object, but does not actually know how much data is stored in each OSD within this placement group, the ratio
between the number of placement groups and the number of OSDs might influence the distribution of the data significantly.
For instance, if there was only one placement group with ten OSDs in a three replica pool, Ceph would only use three OSDs to store data because CRUSH would have no
other choice. When more placement groups are available, CRUSH is more likely to evenly spread objects across OSDs. CRUSH also evenly assigns placement groups to
OSDs.
As long as there are one or two orders of magnitude more placement groups than OSDs, the distribution should be even. For instance, 256 placement groups for 3 OSDs,
512 or 1024 placement groups for 10 OSDs, and so forth.
The ratio between OSDs and placement groups usually solves the problem of uneven data distribution for Ceph clients that implement advanced features like object
striping. For example, a 4 TB block device might get sharded up into 4 MB objects.
The ratio between OSDs and placement groups does not address uneven data distribution in other cases, because CRUSH does not take object size into account.
Using the librados interface to store some relatively small objects and some very large objects can lead to uneven data distribution. For example, one million 4K objects
totaling 4 GB are evenly spread among 1000 placement groups on 10 OSDs. They will use 4 GB / 10 = 400 MB on each OSD. If one 400 MB object is added to the
pool, the three OSDs supporting the placement group in which the object has been placed will be filled with 400 MB + 400 MB = 800 MB while the seven others will
remain occupied with only 400 MB.
Resource usage
For each placement group, OSDs and Ceph monitors need memory, network and CPU at all times, and even more during recovery. Sharing this overhead by clustering
objects within a placement group is one of the main reasons placement groups exist.
Minimizing the number of placement groups saves significant amounts of resources.
Placement group count
The number of placement groups in a pool plays a significant role in how a cluster peers, distributes data and rebalances. Small clusters don’t see as many performance
improvements compared to large clusters by increasing the number of placement groups. However, clusters that have many pools accessing the same OSDs might need to
carefully consider PG count so that Ceph OSDs use resources efficiently.
TIP IBM recommends 100 to 200 PGs per OSD.
Placement group calculator
The placement group (PG) calculator calculates the number of placement groups for you and addresses specific use cases.
Configuring default placement group count
Placement group count for small clusters
Small clusters don’t benefit from large numbers of placement groups, therefore it is important to use the PG calculator with small clusters.
Calculating placement group count
Maximum placement group count
Placement group calculator
The placement group (PG) calculator calculates the number of placement groups for you and addresses specific use cases.
The PG calculator is helpful when using Ceph clients like the Ceph Object Gateway where there are many pools typically using the same rule (CRUSH hierarchy). You might
still calculate PGs manually using the guidelines in Placement group count for small clusters and Calculating placement group count. However, the PG calculator is the
preferred method of calculating PGs.
See Ceph Placement Groups (PGs) per Pool Calculator on the Red Hat Customer Portal for details.
Configuring default placement group count
When you create a pool, you also create a number of placement groups for the pool. If you don’t specify the number of placement groups, Ceph will use the default value
of 8, which is unacceptably low. You can increase the number of placement groups for a pool, but we recommend setting reasonable default values too.
osd pool default pg num = 100
osd pool default pgp num = 100
You need to set both the number of placement groups (total), and the number of placement groups used for objects (used in PG splitting). They should be equal.
Placement group count for small clusters
Small clusters don’t benefit from large numbers of placement groups, therefore it is important to use the PG calculator with small clusters.
102 IBM Storage Ceph
As the number of OSDs increase, choosing the right value for pg_num and pgp_num becomes more important because it has a significant influence on the behavior of the
cluster as well as the durability of the data when something goes wrong (that is the probability that a catastrophic event leads to data loss).
For more information about the PG calculator, see Placement group calculator.
Calculating placement group count
If you have more than 50 OSDs, we recommend approximately 50-100 placement groups per OSD to balance out resource usage, data durability and distribution. If you
have less than 50 OSDs, choosing among the PG Count for Small Clusters is ideal. For a single pool of objects, you can use the following formula to get a baseline:
Total PGs =
(OSDs * 100)
-----------pool size
Where pool size is either the number of replicas for replicated pools or the K+M sum for erasure coded pools (as returned by ceph osd
erasure-code-profile get).
You should then check if the result makes sense with the way you designed your Ceph cluster to maximize data durability, data distribution and minimize resource usage.
The result should be rounded up to the nearest power of two. Rounding up is optional, but recommended for CRUSH to evenly balance the number of objects among
placement groups.
For a cluster with 200 OSDs and a pool size of 3 replicas, you would estimate your number of PGs as follows:
(200 * 100)
----------- = 6667. Nearest power of 2: 8192
3
With 8192 placement groups distributed across 200 OSDs, that evaluates to approximately 41 placement groups per OSD. You also need to consider the number of pools
you are likely to use in your cluster, since each pool will create placement groups too. Ensure that you have a reasonable maximum placement group count (see Maximum
placement group count).
Maximum placement group count
When using multiple data pools for storing objects, you need to ensure that you balance the number of placement groups per pool with the number of placement groups
per OSD so that you arrive at a reasonable total number of placement groups. The aim is to achieve reasonably low variance per OSD without taxing system resources or
making the peering process too slow.
In an exemplary Ceph Storage Cluster consisting of 10 pools, each pool with 512 placement groups on ten OSDs, there are a total of 5,120 placement groups spread over
ten OSDs, or 512 placement groups per OSD. That might not use too many resources depending on your hardware configuration. By contrast, if you create 1,000 pools
with 512 placement groups each, the OSDs will handle ~50,000 placement groups each and it would require significantly more resources. Operating with too many
placement groups per OSD can significantly reduce performance, especially during rebalancing or recovery.
The Ceph Storage Cluster has a default maximum value of 300 placement groups per OSD. You can set a different maximum value in your Ceph configuration file.
mon pg warn max per osd
TIP Ceph Object Gateways deploy with 10-15 pools, so you might consider using less than 100 PGs per OSD to arrive at a reasonable maximum number.
Auto-scaling placement groups
The number of placement groups (PGs) in a pool plays a significant role in how a cluster peers, distributes data, and rebalances.
Auto-scaling the number of PGs can make managing the cluster easier. The pg-autoscaling command provides recommendations for scaling PGs, or automatically
scales PGs based on how the cluster is being used.
To learn more about how auto-scaling works, see Placement group auto-scaling.
To enable, or disable auto-scaling, see Setting placement group auto-scaling modes.
To view placement group scaling recommendations, see Viewing placement group scaling recommendations.
To set placement group auto-scaling, see Setting placement group auto-scaling.
To manually update autoscaler profile, see Manually updating autoscaler profile for a pool
To update the autoscaler globally, see Updating noautoscale flag
To set target pool size, see Specifying target pool size
Placement group auto-scaling
Placement group splitting and merging
Setting placement group auto-scaling modes
Setting minimum and maximum number of placement groups for pools
Viewing placement group scaling recommendations
Setting placement group auto-scaling
IBM Storage Ceph 103
Placement group auto-scaling
How the auto-scaler works
The auto-scaler analyzes pools and adjusts on a per-subtree basis. Because each pool can map to a different CRUSH rule, and each rule can distribute data across
different devices, Ceph considers utilization of each subtree of the hierarchy independently. For example, a pool that maps to OSDs of class ssd, and a pool that maps to
OSDs of class hdd, will each have optimal PG counts that depend on the number of those respective device types.
Placement group splitting and merging
Splitting
IBM Storage Ceph can split existing placement groups (PGs) into smaller PGs, which increases the total number of PGs for a given pool. Splitting existing placement
groups (PGs) allows a small IBM Storage Ceph cluster to scale over time as storage requirements increase. The PG auto-scaling feature can increase the pg_num value,
which causes the existing PGs to split as the storage cluster expands. If the PG auto-scaling feature is disabled, then you can manually increase the pg_num value, which
triggers the PG split process to begin. For example, increasing the pg_num value from 4 to 16, will split into four pieces. Increasing the pg_num value will also increase the
pgp_num value, but the pgp_num value increases at a gradual rate. This gradual increase is done to minimize the impact to a storage cluster’s performance and to a
client’s workload, because migrating object data adds a significant load to the system. By default, Ceph queues and moves no more than 5% of the object data that is in a
"misplaced" state. This default percentage can be adjusted with the target_max_misplaced_ratio option.
Figure 1. Splitting
Merging
IBM Storage Ceph can also merge two existing PGs into a larger PG, which decreases the total number of PGs. Merging two PGs together can be useful, especially when
the relative amount of objects in a pool decreases over time, or when the initial number of PGs chosen was too large. While merging PGs can be useful, it is also a complex
and delicate process. When doing a merge, pausing I/O to the PG occurs, and only one PG is merged at a time to minimize the impact to a storage cluster’s performance.
Ceph works slowly on merging the object data until the new pg_num value is reached.
Figure 2. Merging
104 IBM Storage Ceph
Setting placement group auto-scaling modes
Each pool in the IBM Storage Ceph cluster has a pg_autoscale_mode property for PGs that you can set to off, on, or warn.
off: Disables auto-scaling for the pool. It is up to the administrator to choose an appropriate PG number for each pool. For more information, see Placement group
count.
on: Enables automated adjustments of the PG count for the given pool.
warn: Raises health alerts when the PG count needs adjustment.
Note: pg_autoscale_mode is on by default. Upgraded storage clusters retain the existing pg_autoscale_mode setting. The pg_auto_scale mode is on for the newly
created pools. PG count is automatically adjusted, and ceph status might display a recovering state during PG count adjustment.
The autoscaler uses the bulk flag to determine which pool should start with a full complement of PGs and only scales down when the usage ratio across the pool is not
even. However, if the pool does not have the bulk flag, the pool starts with minimal PGs and only when there is more usage in the pool.
Note: The autoscaler identifies any overlapping roots and prevents the pools with such roots from scaling because overlapping roots can cause problems with the scaling
process.
Procedure
Enable auto-scaling on an existing pool:
Syntax
ceph osd pool set POOL_NAME pg_autoscale_mode on
Example
[ceph: root@host01 /]# ceph osd pool set testpool pg_autoscale_mode on
Enable auto-scaling on a newly created pool:
Syntax
ceph config set global osd_pool_default_pg_autoscale_mode MODE
Example
[ceph: root@host01 /]# ceph config set global osd_pool_default_pg_autoscale_mode on
Create a pool with the bulk flag:
Syntax
ceph osd pool create POOL_NAME [--bulk]
Example
[ceph: root@host01 /]#
ceph osd pool create testpool --bulk
Set or unset the bulk flag for an existing pool:
Important: The values must be written as true, false, 1, or 0. 1 is equivalent to true and 0 is equivalent to false. If written with different capitalization, or with
other content, an error is emitted. The following is an example of the command written with the wrong syntax:
IBM Storage Ceph 105
[ceph: root@host01 /]# ceph osd pool set ec_pool_overwrite bulk
TrueError EINVAL: expecting value 'true', 'false', '0', or '1'
Syntax
ceph osd pool set POOL_NAME bulk true/false/1/0
Example
[ceph: root@host01 /]#
ceph osd pool set testpool bulk true
Get the bulk flag of an existing pool:
Syntax
ceph osd pool get POOL_NAME bulk
Example
[ceph: root@host01 /]# ceph osd pool get testpool bulk
bulk: true
Setting minimum and maximum number of placement groups for pools
Specify the minimum and maximum value of placement groups (PGs) in order to limit the auto-scaling range.
If a minimum value is set, Ceph does not automatically reduce, or recommend to reduce, the number of PGs to a value below the set minimum value.
If a minimum value is set, Ceph does not automatically increase, or recommend to increase, the number of PGs to a value above the set maximum value.
The minimum and maximum values can be set together, or separately.
In addition to the this procedure, the ceph osd pool create command has two command-line options that can be used to specify the minimum or maximum PG count
at the time of pool creation.
Syntax
ceph osd pool create --pg-num-min NUMBER
ceph osd pool create --pg-num-max NUMBER
Example
ceph osd pool create --pg-num-min 50
ceph osd pool create --pg-num-max 150
Prerequisites
A running IBM Storage Ceph cluster
Root-level access to the node
Procedure
Set the minimum number of PGs for a pool.
Syntax
ceph osd pool set POOL_NAME pg_num_min NUMBER
Example
[ceph: root@host01 /]# ceph osd pool set testpool pg_num_min 50
Set the maximum number of PGs for a pool.
Syntax
ceph osd pool set POOL_NAME pg_num_max NUMBER
Example
[ceph: root@host01 /]# ceph osd pool set testpool pg_num_max 150
Reference
For more information, see:
Setting placement group auto-scaling modes
Placement Group count
Viewing placement group scaling recommendations
106 IBM Storage Ceph
You can view the pool, it’s relative utilization and any suggested changes to the PG count in the storage cluster.
Prerequisites
A running IBM Storage Ceph cluster
Root-level access to all the nodes.
Procedure
You can view each pool, its relative utilization, and any suggested changes to the PG count using:
[ceph: root@host01 /]# ceph osd pool autoscale-status
Output will look similar to the following:
POOL
PG_NUM AUTOSCALE BULK
device_health_metrics
on
False
cephfs.cephfs.meta
on
False
cephfs.cephfs.data
on
False
.rgw.root
on
False
default.rgw.log
on
False
default.rgw.control
on
False
default.rgw.meta
on
False
SIZE
TARGET SIZE
RATE
RAW CAPACITY
RATIO
0
3.0
374.9G
24632
3.0
0
3.0
1323
TARGET RATIO
EFFECTIVE RATIO
BIAS
PG_NUM
0.0000
1.0
1
374.9G
0.0000
4.0
32
374.9G
0.0000
1.0
32
3.0
374.9G
0.0000
1.0
32
3702
3.0
374.9G
0.0000
1.0
32
0
3.0
374.9G
0.0000
1.0
32
382
3.0
374.9G
0.0000
4.0
8
NEW
SIZE is the amount of data stored in the pool.
TARGET SIZE, if present, is the amount of data the administrator has specified they expect to eventually be stored in this pool. The system uses the larger of the two
values for its calculation.
RATE is the multiplier for the pool that determines how much raw storage capacity the pool uses. For example, a 3 replica pool has a ratio of 3.0, while a k=4,m=2
erasure coded pool has a ratio of 1.5.
RAW CAPACITY is the total amount of raw storage capacity on the OSDs that are responsible for storing the pool’s data.
RATIO is the ratio of the total capacity that the pool is consuming, that is, ratio = size * rate / raw capacity.
TARGET RATIO, if present, is the ratio of storage the administrator has specified that they expect the pool to consume relative to other pools with target ratios set. If both
target size bytes and ratio are specified, the ratio takes precedence. The default value of TARGET RATIO is 0 unless it was specified while creating the pool. The more the
--target_ratio you give in a pool, the larger the PGs you are expecting the pool to have.
EFFECTIVE RATIO, is the target ratio after adjusting in two ways: 1. subtracting any capacity expected to be used by pools with target size set. 2. normalizing the target
ratios among pools with target ratio set so they collectively target the rest of the space. For example, 4 pools with target ratio 1.0 would have an effective ratio
of 0.25. The system uses the larger of the actual ratio and the effective ratio for its calculation.
BIAS, is used as a multiplier to manually adjust a pool’s PG based on prior information about how much PGs a specific pool is expected to have. By default, the value if 1.0
unless it was specified when creating a pool. The more --bias you give in a pool, the larger the PGs you are expecting the pool to have.
PG_NUM is the current number of PGs for the pool, or the current number of PGs that the pool is working towards, if a pg_num change is in progress. NEW
PG_NUM, if present, is the suggested number of PGs (pg_num). It is always a power of 2, and is only present if the suggested value varies from the current value by more
than a factor of 3.
AUTOSCALE, is the pool pg_autoscale_mode, and is either on, off, or warn.
BULK, is used to determine which pool should start out with a full complement of PGs. BULK only scales down when the usage ratio cross the pool is not even. If the pool
does not have this flag the pool starts out with a minimal amount of PGs and only used when there is more usage in the pool. The BULK values are true, false, 1, or 0,
where 1 is equivalent to true and 0 is equivalent to false. The default value is false. Set the BULK value either during or after pool creation.
For more information on using the bulk flag, see Creating a pool
and Setting placement group auto-scaling.
Setting placement group auto-scaling
Allowing the cluster to automatically scale PGs based on cluster usage is the simplest approach to scaling PGs. IBM Storage Ceph takes the total available storage and the
target number of PGs for the whole system, compares how much data is stored in each pool, and apportions the PGs accordingly. The command only makes changes to a
pool whose current number of PGs (pg_num) is more than three times off from the calculated or suggested PG number.
The target number of PGs per OSD is based on the mon_target_pg_per_osd configurable. The default value is set to 100.
To adjust mon_target_pg_per_osd:
Syntax
ceph config set global mon_target_pg_per_osd number
IBM Storage Ceph 107
Example
[ceph: root@host01 /]# ceph config set global mon_target_pg_per_osd 150
Updating noautoscale flag
If you want to enable or disable the autoscaler for all the pools at same time, you can use the noautoscale global flag. This global flag is useful during upgradation of the
storage cluster when some OSDs are bounced or when the cluster is under maintenance. You can set the flag before any activity and unset it once the activity is complete.
By default, the noautoscale flag is set to off. When this flag is set, then all the pools have pg_autoscale_mode as off and all the pools have the autoscaler disabled.
Prerequisites
A running IBM Storage Ceph cluster.
Root-level access to all the nodes.
Procedure
1. Get the value of the noautoscale flag:
Example
[ceph: root@host01 /]# ceph osd pool get noautoscale
2. Set the noautoscale flag before any activity:
Example
[ceph: root@host01 /]# ceph osd pool set noautoscale
3. Unset the noautoscale flag on completion of the activity:
Example
[ceph: root@host01 /]# ceph osd pool unset noautoscale
Specifying target pool size
Specify target pool size.
A newly created pool consumes a small fraction of the total cluster capacity and appears to the system that it will need a small number of PGs. However, in most cases,
cluster administrators know which pools are expected to consume most of the system capacity over time. If you provide this information, known as the target size to
IBM Storage Ceph, such pools can use a more appropriate number of PGs (pg_num) from the beginning. This approach prevents subsequent changes in pg_num and the
overhead associated with moving data around when making those adjustments.
You can specify target size of a pool by either using the absolute size of the pool or the total cluster capacity.
Specifying target size using the absolute size of the pool
Specify the target size of a pool by using the absolute size of the pool.
Specifying target size using the total cluster capacity
Specify the target size of a pool by using the total cluster capacity.
Specifying target size using the absolute size of the pool
Specify the target size of a pool by using the absolute size of the pool.
1. Set the target size using the absolute size of the pool in bytes:
ceph osd pool set pool-name target_size_bytes value
For example, to instruct the system that mypool is expected to consume 100T of space:
$ ceph osd pool set mypool target_size_bytes 100T
You can also set the target size of a pool at creation time by adding the optional --target-size-bytes <bytes> argument to the ceph osd pool create command.
Specifying target size using the total cluster capacity
Specify the target size of a pool by using the total cluster capacity.
1. Set the target size using the ratio of the total cluster capacity:
Syntax
108 IBM Storage Ceph
ceph osd pool set pool-name target_size_ratio ratio
Example
[ceph: root@host01 /]# ceph osd pool set mypool target_size_ratio 1.0
tells the system that the pool mypool is expected to consume 1.0 relative to the other pools with target_size_ratio set. If mypool is the only pool in the
cluster, this means an expected use of 100% of the total capacity. If there is a second pool with target_size_ratio as 1.0, both pools would expect to use 50%
of the cluster capacity.
You can also set the target size of a pool at creation time by adding the optional --target-size-ratio <ratio> argument to the ceph osd pool
create command.
NOTE If you specify impossible target size values, for example, a capacity larger than the total cluster, or ratios that sum to more than 1.0, the cluster raises a
POOL_TARGET_SIZE_RATIO_OVERCOMMITTED or POOL_TARGET_SIZE_BYTES_OVERCOMMITTED health warning. If you specify both target_size_ratio and
target_size_bytes for a pool, the cluster considers only the ratio, and raises a POOL_HAS_TARGET_SIZE_BYTES_AND_RATIO health warning.
Placement group command line interface
The ceph CLI allows you to set and get the number of placement groups for a pool, view the PG map and retrieve PG statistics.
Setting number of placement groups in a pool
To set the number of placement groups in a pool, you must specify the number of placement groups at the time you create the pool.
Getting number of placement groups in a pool
Getting statistics for placement groups
Getting statistics for stuck placement groups
Getting placement group maps
Getting a placement group statistics
Scrubbing placement groups
Marking unfound objects
Setting number of placement groups in a pool
To set the number of placement groups in a pool, you must specify the number of placement groups at the time you create the pool.
Once you set placement groups for a pool, you can increase or decrease the number of placement groups. For more information, see Creating a pool.
To change the number of placement groups, use the following command:
Syntax
ceph osd pool set POOL_NAME pg_num PG_NUM
Example
[ceph: root@host01 /]# ceph osd pool set pool1 pg_num 60
set pool 2 pg_num to 60
Once you increase or decrease the number of placement groups, you must also adjust the number of placement groups for placement (pgp_num) before your cluster
rebalances. The pgp_num should be equal to the pg_num. To increase the number of placement groups for placement, execute the following:
Syntax
ceph osd pool set POOL_NAME pgp_num PGP_NUM
Example
[ceph: root@host01 /]# ceph osd pool set pool1 pgp_num 60
set pool 2 pgp_num to 60
Getting number of placement groups in a pool
To get the number of placement groups in a pool, run the following:
Syntax
ceph osd pool get POOL_NAME pg_num
Example
[ceph: root@host01 /]# ceph osd pool get testpool 60
Getting statistics for placement groups
To get the statistics for the placement groups in your storage cluster, execute the following:
IBM Storage Ceph 109
Syntax
ceph pg dump [--format FORMAT]
Valid formats are plain (default) and json.
Getting statistics for stuck placement groups
To get the statistics for all placement groups stuck in a specified state, execute the following:
Syntax
ceph pg dump_stuck {inactive|unclean|stale|undersized|degraded [inactive|unclean|stale|undersized|degraded...]} INTEGER
Inactive Placement groups cannot process reads or writes because they are waiting for an OSD with the most up-to-date data to come up and in.
Unclean Placement groups contain objects that are not replicated the desired number of times. They should be recovering.
Stale Placement groups are in an unknown state - the OSDs that host them have not reported to the monitor cluster in a while (configured by
mon_osd_report_timeout).
Valid formats are plain (default) and json. The threshold defines the minimum number of seconds the placement group is stuck before including it in the returned
statistics (default 300 seconds).
Getting placement group maps
To get the placement group map for a particular placement group, execute the following:
Syntax
ceph pg map PG_ID
Example
[ceph: root@host01 /]# ceph pg map 1.6c
osdmap e13 pg 1.6c (1.6c) -> up [1,0] acting [1,0]
Ceph returns the placement group map, the placement group, and the OSD status:
Getting a placement group statistics
Retrieve statistics for a particular placement group:
Syntax
ceph pg PG_ID query
Scrubbing placement groups
To scrub a placement group, execute the following:
Syntax
ceph pg scrub PG_ID
Ceph checks the primary and any replica nodes, generates a catalog of all objects in the placement group and compares them to ensure that no objects are missing or
mismatched, and their contents are consistent. Assuming the replicas all match, a final semantic sweep ensures that all of the snapshot-related object metadata is
consistent. Errors are reported via logs.
Marking unfound objects
If the cluster has lost one or more objects, and you have decided to abandon the search for the lost data, you must mark the unfound objects as lost.
If all possible locations have been queried and objects are still lost, you might have to give up on the lost objects. This is possible given unusual combinations of failures
that allow the cluster to learn about writes that were performed before the writes themselves are recovered.
Currently the only supported option is "revert", which will either roll back to a previous version of the object or (if it was a new object) forget about it entirely. To mark the
"unfound" objects as "lost", execute the following:
Syntax
ceph pg PG_ID mark_unfound_lost revert|delete
110 IBM Storage Ceph
Important: Use this feature with caution, because it might confuse applications that expect the object(s) to exist.
Pools overview
Ceph clients store data in pools. When you create pools, you are creating an I/O interface for clients to store data. From the perspective of a Ceph client (that is, block
device, gateway, and the rest), interacting with the Ceph storage cluster is remarkably simple: create a cluster handle and connect to the cluster; then, create an I/O
context for reading and writing objects and their extended attributes.
Create a Cluster Handle and Connect to the Cluster
To connect to the Ceph storage cluster, the Ceph client needs the cluster name (usually ceph by default) and an initial monitor address. Ceph clients usually retrieve these
parameters using the default path for the Ceph configuration file and then read it from the file, but a user might also specify the parameters on the command line too. The
Ceph client also provides a user name and secret key (authentication is on by default). Then, the client contacts the Ceph monitor cluster and retrieves a recent copy of
the cluster map, including its monitors, OSDs and pools.
Figure 1. Create Handle
Create a Pool I/O Context
To read and write data, the Ceph client creates an i/o context to a specific pool in the Ceph storage cluster. If the specified user has permissions for the pool, the Ceph
client can read from and write to the specified pool.
Figure 2. I/O Context
Ceph’s architecture enables the storage cluster to provide this remarkably simple interface to Ceph clients so that clients might select one of the sophisticated storage
strategies you define simply by specifying a pool name and creating an I/O context. Storage strategies are invisible to the Ceph client in all but capacity and performance.
Similarly, the complexities of Ceph clients (mapping objects into a block device representation, providing an S3/Swift RESTful service) are invisible to the Ceph storage
cluster.
A pool provides you with:
Resilience: You can set how many OSD are allowed to fail without losing data. For replicated pools, it is the desired number of copies/replicas of an object. A typical
configuration stores an object and one additional copy (that is, size = 2), but you can determine the number of copies/replicas. For erasure coded pools, it is the
number of coding chunks (that is m=2 in the erasure code profile)
Placement Groups: You can set the number of placement groups for the pool. A typical configuration uses approximately 50-100 placement groups per OSD to
provide optimal balancing without using up too many computing resources. When setting up multiple pools, be careful to ensure you set a reasonable number of
placement groups for both the pool and the cluster as a whole.
CRUSH Rules: When you store data in a pool, a CRUSH rule mapped to the pool enables CRUSH to identify the rule for the placement of each object and its replicas
or chunks for erasure coded pools in your cluster. You can create a custom CRUSH rule for your pool.
IBM Storage Ceph 111
Snapshots: When you create snapshots with ceph osd pool mksnap, you effectively take a snapshot of a particular pool.
Quotas: When you set quotas on a pool with ceph osd pool set-quota you might limit the maximum number of objects or the maximum number of bytes
stored in the specified pool.
Pools and storage strategies overview
Listing pool
Creating a pool
Learn how to create a pool.
Setting pool quota
You can set pool quotas for the maximum number of bytes or the maximum number of objects per pool or for both.
Deleting a pool
Learn how to delete a pool.
Renaming a pool
Migrating a pool
Sometimes it is necessary to migrate all objects from one pool to another. This is done in cases such as needing to change parameters that cannot be modified on a
specific pool. For example, needing to reduce the number of placement groups of a pool.
Viewing pool statistics
Setting pool values
Set pool values.
Getting pool values
Get a value from a pool.
Enabling a client application
Disabling a client application
Setting application metadata
Removing application metadata
Setting the number of object replicas
Learn how to set the number of object replicas.
Getting the number of object replicas
Pool values
This information contains key-values pairs that you can set or get.
Pools and storage strategies overview
To manage pools, you can list, create, and remove pools. You can also view the utilization statistics for each pool.
Listing pool
To list your cluster’s pools, execute:
ceph osd lspools
Creating a pool
Learn how to create a pool.
Before creating pools, see Pools, placement groups, and CRUSH configuration.
Note: The system administrators must expressly enable a pool to receive I/O operations from Ceph clients. See Enabling a client application for details. Failure to enable a
pool will result in a HEALTH_WARN status.
It is better to adjust the default value for the number of placement groups in the Ceph configuration file, as the default value does not have to suit your needs.
Example
osd pool default pg num = 100
osd pool default pgp num = 100
To create a replicated pool, execute:
Syntax
ceph osd pool create POOL_NAME PG_NUM PGP_NUM [replicated]
[CRUSH_RULE_NAME] [EXPECTED_NUMBER_OBJECTS]
To create an erasure-coded pool, execute:
Syntax
ceph osd pool create POOL_NAME PG_NUM PGP_NUM erasure
[ERASURE_CODE_PROFILE] [CRUSH_RULE_NAME] [EXPECTED_NUMBER_OBJECTS]
Create a bulk pool:
Syntax
ceph osd pool create POOL_NAME [--bulk]
112 IBM Storage Ceph
Where:
POOL_NAME
Description
The name of the pool. It must be unique.
Type
String
Required
Yes. If not specified, it is set to the value listed in the Ceph configuration file or to the default value.
Default
ceph
PG_NUM
Description
The total number of placement groups for the pool. For more information about calculating a suitable number, see Placement Groups and Ceph Placement Groups (PGs)
per Pool Calculator on the Red Hat Customer Portal. The default value 8 is not suitable for most systems.
Type
Integer
Required
Yes
Default
8
PGP_NUM
Description
The total number of placement groups for placement purposes. This value must be equal to the total number of placement groups, except for placement group splitting
scenarios.
Type
Integer
Required
Yes. If not specified it is set to the value listed in the Ceph configuration file or to the default value.
Default
8
replicated or erasure
Description
The pool type can be either replicated to recover from lost OSDs by keeping multiple copies of the objects or erasure to get a kind of generalized RAID5 capability.
The replicated pools require more raw storage but implement all Ceph operations. The erasure-coded pools require less raw storage but only implement a subset of the
available operations.
Type
String
Required
No
Default
replicated
crush-rule-name
Description
The name of the crush rule for the pool. The rule MUST exist. For replicated pools, the name is the rule specified by the osd_pool_default_crush_rule configuration
setting. For erasure-coded pools the name is erasure-code if you specify the default erasure code profile or POOL_NAME otherwise. Ceph creates this rule with the
specified name implicitly if the rule doesn’t already exist.
Type
String
Required
No
Default
Uses erasure-code for an erasure-coded pool. For replicated pools, it uses the value of the osd_pool_default_crush_rule variable from the Ceph configuration.
expected-num-objects
Description
The expected number of objects for the pool. By setting this value together with a negative filestore_merge_threshold variable, Ceph splits the placement groups at
pool creation time to avoid the latency impact to perform runtime directory splitting.
Type
Integer
Required
No
Default
0, no splitting at the pool creation time
IBM Storage Ceph 113
erasure-code-profile
Description
For erasure-coded pools only. Use the erasure code profile. It must be an existing profile as defined by the osd erasure-code-profile set variable in the Ceph
configuration file. For more information, see Erasure code profiles.
Type
String
Required
No
When you create a pool, set the number of placement groups to a reasonable value (for example to 100). Consider the total number of placement groups per OSD too.
Placement groups are computationally expensive, so performance will degrade when you have many pools with many placement groups, for example, 50 pools with 100
placement groups each. The point of diminishing returns depends upon the power of the OSD host.
Setting pool quota
You can set pool quotas for the maximum number of bytes or the maximum number of objects per pool or for both.
Syntax
ceph osd pool set-quota POOL_NAME [max_objects OBJECT_COUNT>] [max_bytes BYTES]
Example
[ceph: root@host01 /]# ceph osd pool set-quota data max_objects 10000
To remove a quota, set its value to 0.
Note: In-flight write operations might overrun pool quotas for a short time until Ceph propagates the pool usage across the cluster. This is normal behavior. Enforcing pool
quotas on in-flight write operations would impose significant performance penalties.
Deleting a pool
Learn how to delete a pool.
To delete a pool, execute:
Syntax
ceph osd pool delete POOL_NAME [POOL_NAME --yes-i-really-really-mean-it]
Important: To protect data, storage administrators cannot delete pools by default. Set the mon_allow_pool_delete configuration option before deleting pools.
If a pool has its own rule, consider removing it after deleting the pool. If a pool has users strictly for its own use, consider deleting those users after deleting the pool.
Renaming a pool
To rename a pool, execute:
Syntax
ceph osd pool rename CURRENT_POOL_NAME NEW_POOL_NAME
If you rename a pool and you have per-pool capabilities for an authenticated user, you must update the user’s capabilities (that is, caps) with the new pool name.
Migrating a pool
Sometimes it is necessary to migrate all objects from one pool to another. This is done in cases such as needing to change parameters that cannot be modified on a
specific pool. For example, needing to reduce the number of placement groups of a pool.
About this task
Important: When a workload is using only Ceph Block Device images, use the following procedures:
Moving images between pools
Migrating pools
The migration methods described for Ceph Block Device are more recommended than those documented here. using the cppool does not preserve all snapshots and
snapshot related metadata, resulting in an unfaithful copy of the data. For example, copying an RBD pool does not completely copy the image. In this case, snaps are not
present and will not work properly. The cppool also does not preserve the user_version field that some librados users may rely on.
If migrating a pool is necessary and your user workloads contain images other than Ceph Block Devices, continue with one of the procedures documented here.
Before you begin
114 IBM Storage Ceph
If using the rados cppool command:
Read-only access to the pool is required.
Only use this command if you do not have RBD images and its snaps and user_version consumed by librados.
If using the local drive RADOS commands, verify that sufficient cluster space is available. Two, three, or more copies of data will be present as per pool replication
factor.
Migrating directly
Copy all objects with the rados cppool command.
Important: Read-only access to the pool is required during copy.
ceph osd pool create NEW_POOL PG_NUM [ <other new pool parameters> ]
rados cppool SOURCE_POOL NEW_POOL
ceph osd pool rename SOURCE_POOL NEW_SOURCE_POOL_NAME
ceph osd pool rename NEW_POOL SOURCE_POOL
For example,
[ceph: root@host01 /]# ceph osd pool create pool1 250
[ceph: root@host01 /]# rados cppool pool2 pool1
[ceph: root@host01 /]# ceph osd pool rename pool2 pool3
[ceph: root@host01 /]# ceph osd pool rename pool1 pool2
Migrating through a local drive
Procedure
1. Use the rados export and rados import commands and a temporary local directory to save all exported data.
ceph osd pool create NEW_POOL PG_NUM [ <other new pool parameters> ]
rados export --create SOURCE_POOL FILE_PATH
rados import FILE_PATH NEW_POOL
For example,
[ceph: root@host01 /]# ceph osd pool create pool1 250
[ceph: root@host01 /]# rados export --create pool2 <path of export file>
[ceph: root@host01 /]# rados import <path of export file> pool1
2. Required: Stop all I/O to the source pool.
3. Required: Resynchronize all modified objects.
rados export --workers 5 SOURCE_POOL FILE_PATH
rados import --workers 5 FILE_PATH NEW_POOL
For example,
[ceph: root@host01 /]# rados export --workers 5 pool2 <path of export file>
[ceph: root@host01 /]# rados import --workers 5 <path of export file> pool1
Viewing pool statistics
To show a pool’s utilization statistics, run the following command:
Syntax
rados df
Setting pool values
Set pool values.
To set a value to a pool, execute the following command:
ceph osd pool set POOL_NAME KEY VALUE
For more information about the available key-value pairs, see Pool values.
Getting pool values
Get a value from a pool.
To get a value from a pool, execute the following command:
ceph osd pool get POOL_NAME KEY
For more information about the available key-value pairs, see Pool values.
IBM Storage Ceph 115
Enabling a client application
IBM Storage Ceph provides additional protection for pools to prevent unauthorized types of clients from writing data to the pool. This means that system administrators
must expressly enable pools to receive I/O operations from Ceph Block Device, Ceph Object Gateway, Ceph Filesystem or for a custom application.
To enable a client application to conduct I/O operations on a pool, execute the following:
Syntax
ceph osd pool application enable POOL_NAME APP {--yes-i-really-mean-it}
Where APP is:
cephfs for the Ceph Filesystem.
rbd for the Ceph Block Device
rgw for the Ceph Object Gateway
Important: A pool that is not enabled will generate a HEALTH_WARN status.
In that scenario, the output for ceph health detail -f json-pretty gives the following output:
{
"checks": {
"POOL_APP_NOT_ENABLED": {
"severity": "HEALTH_WARN",
"summary": {
"message": "application not enabled on 1 pool(s)"
},
"detail": [
{
"message": "application not enabled on pool 'POOL_NAME'"
},
{
"message": "use 'ceph osd pool application enable POOL_NAME APP', where APP is 'cephfs', 'rbd', 'rgw', or
freeform for custom applications."
}
]
}
},
"status": "HEALTH_WARN",
"overall_status": "HEALTH_WARN",
"detail": [
"'ceph health' JSON format has changed in luminous. If you see this your monitoring system is scraping the wrong
fields. Disable this with 'mon health preluminous compat warning = false'"
]
}
Note:
Specify a different APP value for a custom application.
Initialize pools for the Ceph Block Device with rbd pool init POOL_NAME.
Disabling a client application
To disable a client application from conducting I/O operations on a pool, execute the following:
Syntax
ceph osd pool application disable POOL_NAME APP {--yes-i-really-mean-it}
Where APP is:
cephfs for the Ceph Filesystem.
rbd for the Ceph Block Device
rgw for the Ceph Object Gateway
Note: Specify a different APP value for a custom application.
Setting application metadata
Provides the functionality to set key-value pairs describing attributes of the client application.
To set client application metadata on a pool, execute the following:
Syntax
ceph osd pool application set POOL_NAME APP KEY
Where APP is:
116 IBM Storage Ceph
cephfs for the Ceph Filesystem.
rbd for the Ceph Block Device
rgw for the Ceph Object Gateway
Note: Specify a different APP value for a custom application.
Removing application metadata
To remove client application metadata on a pool, execute the following:
Syntax
ceph osd pool application rm POOL_NAME APP KEY
Where APP is:
cephfs for the Ceph Filesystem.
rbd for the Ceph Block Device
rgw for the Ceph Object Gateway
Note: Specify a different APP value for a custom application.
Setting the number of object replicas
Learn how to set the number of object replicas.
To set the number of object replicas on a replicated pool, execute the following command:
Syntax
ceph osd pool set POOL_NAME size NUMBER_OF_REPLICAS
You can run this command for each pool.
The NUMBER_OF_REPLICAS parameter includes the object itself. If you want to include the object and two copies of the object for a total of three instances of the
object, specify 3.
Example
[ceph: root@host01 /]# ceph osd pool set data size 3
An object might accept I/O operations in degraded mode with fewer replicas than specified by the pool size setting. To set a minimum number of required
replicas for I/O, use the min_size setting.
Example
ceph osd pool set data min_size 2
This ensures that no object in the data pool will receive I/O with fewer replicas than specified by the min_size setting.
Getting the number of object replicas
To get the number of object replicas, execute the following command:
ceph osd dump | grep 'replicated size'
Ceph will list the pools, with the replicated size attribute highlighted. By default, Ceph creates two replicas of an object, that is a total of three copies, or a size of 3.
Pool values
This information contains key-values pairs that you can set or get.
For further information, see Setting pool values and Getting pool values.
size
Description
Specifies the number of replicas for objects in the pool. For more information, see Setting the number of object replicas. Applicable for the replicated pools only.
Type
Integer
min_size
IBM Storage Ceph 117
Description
Specifies the minimum number of replicas required for I/O. For more information, see Setting the number of object replicas. For erasure-coded pools, this should be set to
a value greater than k. If I/O is allowed at the value k, then there is no redundancy and data is lost in the event of a permanent OSD failure. For more information, see
Erasure code pools overview.
Type
Integer
crash_replay_interval
Description
Specifies the number of seconds to allow clients to replay acknowledged, but uncommitted requests.
Type
Integer
pg-num
Description The total number of placement groups for the pool. For more information, see Pools, placement groups, and CRUSH configuration section for details on
calculating a suitable number. The default value 8 is not suitable for most systems.
Type
Integer
Required
Yes.
Default
8
pgp-num
Description
The total number of placement groups for placement purposes. This should be equal to the total number of placement groups, except for placement group splitting
scenarios.
Type
Integer
Required
Yes. Picks up default or Ceph configuration value if not specified.
Default
8
Valid Range
Equal to or less than what specified by the pg_num variable.
crush_rule
Description
The rule to use for mapping object placement in the cluster.
Type
String
hashpspool
Description
Enable or disable the HASHPSPOOL flag on a given pool. With this option enabled, pool hashing and placement group mapping are changed to improve the way pools and
placement groups overlap.
Type
Integer
Valid Range
1 enables the flag, 0 disables the flag.
IMPORTANT Do not enable this option on production pools of a cluster with a large amount of OSDs and data. All placement groups in the pool would have to be
remapped causing too much data movement.
fast_read
Description
On a pool that uses erasure coding, if this flag is enabled, the read request issues subsequent reads to all shards, and waits until it receives enough shards to decode to
serve the client. In the case of the jerasure and isa
erasure plug-ins, once the first K replies return, the client’s request is served immediately using the data decoded from these replies. This helps to allocate some
resources for better performance. Currently this flag is only supported for erasure coding pools.
Type
Boolean
Defaults
0
allow_ec_overwrites
118 IBM Storage Ceph
Description
Whether writes to an erasure coded pool can update part of an object, so the Ceph Filesystem and Ceph Block Device can use it.
Type
Boolean
compression_algorithm
Description
Sets inline compression algorithm to use with the BlueStore storage backend. This setting overrides the bluestore_compression_algorithm configuration setting.
Type
String
Valid Settings
lz4, snappy, zlib, zstd
compression_mode
Description
Sets the policy for the inline compression algorithm for the BlueStore storage backend. This setting overrides the bluestore_compression_mode configuration setting.
Type
String
Valid Settings
none, passive, aggressive, force
compression_min_blob_size
Description
BlueStore will not compress chunks smaller than this size. This setting overrides the bluestore_compression_min_blob_size configuration setting.
Type
Unsigned Integer
compression_max_blob_size
Description
BlueStore will break chunks larger than this size into smaller blobs of compression_max_blob_size before compressing the data.
Type
Unsigned Integer
nodelete
Description
Set or unset the NODELETE flag on a given pool.
Type
Integer
Valid Range
1 sets flag. 0 unsets flag.
nopgchange
Description
Set or unset the NOPGCHANGE flag on a given pool.
Type
Integer
Valid Range
1 sets the flag. 0 unsets the flag.
nosizechange
Description
Set or unset the NOSIZECHANGE flag on a given pool.
Type
Integer
Valid Range
1 sets the flag. 0 unsets the flag.
write_fadvise_dontneed
Description
Set or unset the WRITE_FADVISE_DONTNEED flag on a given pool.
Type
Integer
Valid Range
1 sets the flag. 0 unsets the flag.
noscrub
IBM Storage Ceph 119
Description
Set or unset the NOSCRUB flag on a given pool.
Type
Integer
Valid Range
1 sets the flag. 0 unsets the flag.
nodeep-scrub
Description
Set or unset the NODEEP_SCRUB flag on a given pool.
Type
Integer
Valid Range
1 sets the flag. 0 unsets the flag.
scrub_min_interval
Description
The minimum interval in seconds for pool scrubbing when load is low. If it is 0, Ceph uses the osd_scrub_min_interval configuration setting.
Type
Double
Default
0
scrub_max_interval
Description
The maximum interval in seconds for pool scrubbing irrespective of cluster load. If it is 0, Ceph uses the osd_scrub_max_interval configuration setting.
Type
Double
Default
0
deep_scrub_interval
Description
The interval in seconds for pool deep scrubbing. If it is 0, Ceph uses the osd_deep_scrub_interval configuration setting.
Type
Double
Default
0
Erasure code pools overview
Ceph storage strategies involve defining data durability requirements. Data durability means the ability to sustain the loss of one or more OSDs without losing data.
Ceph stores data in pools and there are two types of the pools:
replicated
erasure-coded
Ceph uses the replicated pools by default, meaning the Ceph copies every object from a primary OSD node to one or more secondary OSDs.
The erasure-coded pools reduce the amount of disk space required to ensure data durability but it is computationally a bit more expensive than replication.
Erasure coding is a method of storing an object in the Ceph storage cluster durably where the erasure code algorithm breaks the object into data chunks (k) and coding
chunks (m), and stores those chunks in different OSDs.
In the event of the failure of an OSD, Ceph retrieves the remaining data (k) and coding (m) chunks from the other OSDs and the erasure code algorithm restores the object
from those chunks.
Note: Use min_size for erasure-coded pools to be K+1 or more to prevent loss of writes and data.
Erasure coding uses storage capacity more efficiently than replication. The n-replication approach maintains n copies of an object (3x by default in Ceph), whereas erasure
coding maintains only k + m chunks. For example, 3 data and 2 coding chunks use 1.5x the storage space of the original object.
While erasure coding uses less storage overhead than replication, the erasure code algorithm uses more RAM and CPU than replication when it accesses or recovers
objects. Erasure coding is advantageous when data storage must be durable and fault tolerant, but do not require fast read performance (for example, cold storage,
historical records, and so on).
For the mathematical and detailed explanation on how erasure code works in Ceph, see the Erasure Coding.
Ceph creates a default erasure code profile when initializing a cluster with k=2 and m=2, This mean that Ceph spreads the object data over four OSDs (k+m == 4) and
Ceph can lose one of those OSDs without losing data. To know more about erasure code profiling see Erasure code profiles section.
120 IBM Storage Ceph
Important: Configure only the .rgw.buckets pool as erasure-coded and all other Ceph Object Gateway pools as replicated, otherwise an attempt to create a new bucket
fails with the following error:
set_req_state_err err_no=95 resorting to 500
The reason for this is that erasure-coded pools do not support the omap operations and certain Ceph Object Gateway metadata pools require the omap support.
Creating a sample erasure-coded pool
The simplest erasure coded pool is equivalent to RAID5 and requires at least three hosts.
Erasure code profiles
Erasure Coding with Overwrites
Erasure Code Plugins
Creating a sample erasure-coded pool
The simplest erasure coded pool is equivalent to RAID5 and requires at least three hosts.
he ceph osd pool create command creates an erasure-coded pool with the default profile, unless another profile is specified. Profiles define the redundancy of data by
setting two parameters, k, and k. These parameters define the number of chunks a piece of data is split and the number of coding chunks are created. Use the following
example to create an erasure-coded pool, where the 32 in the pool create command is the number of placement groups.
$ ceph osd pool create ecpool 32 32 erasure
pool 'ecpool' created
$ echo ABCDEFGHI | rados --pool ecpool put NYAN $ rados --pool ecpool get NYAN ABCDEFGHI
Erasure code profiles
Ceph defines an erasure-coded pool with a profile. Ceph uses a profile when creating an erasure-coded pool and the associated CRUSH rule.
Ceph creates a default erasure code profile when initializing a cluster and it provides the same level of redundancy as two copies in a replicated pool. However, it uses
25% less storage capacity. The default profiles define k=2 and m=2, meaning Ceph will spread the object data over four OSDs (k+m=4) and Ceph can lose one of those
OSDs without losing data.
The default erasure code profile can sustain the loss of a single OSD. It is equivalent to a replicated pool with a size two, but requires 1.5 TB instead of 2 TB to store 1 TB
of data. To display the default profile use the following command:
$ ceph osd erasure-code-profile get default
k=2
m=2
plugin=jerasure
technique=reed_sol_van
You can create a new profile to improve redundancy without increasing raw storage requirements. For instance, a profile with k=8 and m=4 can sustain the loss of four
(m=4) OSDs by distributing an object on 12 (k+m=12) OSDs. Ceph divides the object into 8 chunks and computes 4 coding chunks for recovery. For example, if the object
size is 8 MB, each data chunk is 1 MB and each coding chunk has the same size as the data chunk, that is also 1 MB. The object will not be lost even if four OSDs fail
simultaneously.
The most important parameters of the profile are k, m and crush-failure-domain, because they define the storage overhead and the data durability.
Important: Choosing the correct profile is important because you cannot change the profile after you create the pool. To modify a profile, you must create a new pool with
a different profile and migrate the objects from the old pool to the new pool.
For instance, if the desired architecture must sustain the loss of two racks with a storage overhead of 40% overhead, the following profile can be defined:
$ ceph osd erasure-code-profile set myprofile \
k=4 \
m=2 \
crush-failure-domain=rack
$ ceph osd pool create ecpool 12 12 erasure *myprofile*
$ echo ABCDEFGHIJKL | rados --pool ecpool put NYAN $ rados --pool ecpool get NYAN ABCDEFGHIJKL
The primary OSD will divide the NYAN object into four (k=4) data chunks and create two additional chunks (m=2). The value of m defines how many OSDs can be lost
simultaneously without losing any data. The crush-failure-domain=rack will create a CRUSH rule that ensures no two chunks are stored in the same rack.
Figure 1. Erasure code
IBM Storage Ceph 121
Important:
The following jerasure coding values are supported for k, and m:
k=8 m=3
k=8 m=4
k=4 m=2
If the number of OSDs lost equals the number of coding chunks (m), some placement groups in the erasure coding pool will go into incomplete state. If the number
of OSDs lost is less than m, no placement groups will go into incomplete state. In either situation, no data loss will occur. If placement groups are in incomplete
state, temporarily reducing min_size of an erasure coded pool will allow recovery.
Setting OSD erasure-code-profile
Removing OSD erasure-code-profile
Getting OSD erasure-code-profile
Listing OSD erasure-code-profile
Setting OSD erasure-code-profile
To create a new erasure code profile:
Syntax
ceph osd erasure-code-profile set NAME \
[<directory=DIRECTORY>] \
[<plugin=PLUGIN>] \
[<stripe_unit=STRIPE_UNIT>] \
[<CRUSH_DEVICE_CLASS>]\
[<CRUSH_FAILURE_DOMAIN>]\
[<key=value> ...] \
[--force]
Where:
directory
Description
Set the directory name from which the erasure code plug-in is loaded.
Type
String
Required
No.
Default
/usr/lib/ceph/erasure-code
122 IBM Storage Ceph
plugin
Description
Use the erasure code plug-in to compute coding chunks and recover missing details. For more information, see Erasure Code Plugins.
Type
String
Required
No.
Default
jerasure
stripe_unit
Description
The amount of data in a data chunk, per stripe. For example, a profile with 2 data chunks and stripe_unit=4K would put the range 0-4K in chunk 0, 4K-8K in chunk 1,
then 8K-12K in chunk 0 again. This should be a multiple of 4K for best performance. The default value is taken from the monitor config option
osd_pool_erasure_code_stripe_unit when a pool is created. The stripe_width of a pool using this profile will be the number of data chunks multiplied by this
stripe_unit.
Type
String
Required
No.
Default
4K
crush-device-class
Description
The device class, such as hdd or ssd.
Type
String
Required
No
Default
none, meaning CRUSH uses all devices regardless of class.
crush-failure-domain
Description
The failure domain, such as host or rack.
Type
String
Required
No
Default
host
key
Description
The semantic of the remaining key-value pairs is defined by the erasure code plug-in.
Type
String
Required
No.
--force
Description
Override an existing profile by the same name.
Type
String
Required
No.
Removing OSD erasure-code-profile
To remove an erasure code profile:
Syntax
ceph osd erasure-code-profile rm NAME
Important: If the profile is referenced by a pool, the deletion fails.
IBM Storage Ceph 123
Warning: Removing an erasure code profile using osd erasure-code-profile rm command does not automatically delete the associated CRUSH rule that is associated with
the erasure code profile. Manually remove the associated CRUSH rule using ceph osd crush rule remove RULE_NAME command to avoid unexpected behavior.
Getting OSD erasure-code-profile
To display an erasure code profile:
Syntax
ceph osd erasure-code-profile get NAME
Listing OSD erasure-code-profile
To list the names of all erasure code profiles:
Syntax
ceph osd erasure-code-profile ls
Erasure Coding with Overwrites
By default, erasure coded pools only work with the Ceph Object Gateway, which performs full object writes and appends.
Using erasure coded pools with overwrites allows Ceph Block Devices and CephFS store their data in an erasure coded pool:
Syntax
ceph osd pool set ERASURE_CODED_POOL_NAME allow_ec_overwrites true
Example
[ceph: root@host01 /]# ceph osd pool set ec_pool allow_ec_overwrites true
Enabling erasure coded pools with overwrites can only reside in a pool using BlueStore OSDs. Since BlueStore’s checksumming is used to detect bit rot or other corruption
during deep scrubs. Using FileStore with erasure coded overwrites is unsafe, and yields lower performance when compared to BlueStore.
Erasure coded pools do not support omap. To use erasure coded pools with Ceph Block Devices and CephFS, store the data in an erasure coded pool, and the metadata in
a replicated pool.
For Ceph Block Devices, use the --data-pool option during image creation:
Syntax
rbd create --size IMAGE_SIZE_M|G|T --data-pool _ERASURE_CODED_POOL_NAME REPLICATED_POOL_NAME/IMAGE_NAME
Example
[ceph: root@host01 /]# rbd create --size 1G --data-pool ec_pool rep_pool/image01
If using erasure coded pools for CephFS, then setting the overwrites must be done in a file layout.
Erasure Code Plugins
Ceph supports erasure coding with a plug-in architecture, which means you can create erasure coded pools using different types of algorithms. Ceph supports: - Jerasure
(Default)
Creating a new erasure code profile using jerasure erasure code plugin
Controlling CRUSH Placement
Creating a new erasure code profile using jerasure erasure code plugin
The jerasure plug-in is the most generic and flexible plug-in. It is also the default for Ceph erasure coded pools.
The jerasure plug-in encapsulates the JerasureH library. For detailed information about the parameters, see the jerasure documentation.
To create a new erasure code profile using the jerasure plug-in, run the following command:
Syntax
ceph osd erasure-code-profile set NAME \
plugin=jerasure \
k=DATA_CHUNKS \
m=DATA_CHUNKS \
technique=TECHNIQUE \
124 IBM Storage Ceph
[crush-root=ROOT] \
[crush-failure-domain=BUCKET_TYPE] \
[directory=DIRECTORY] \
[--force]
Where:
k
Description
Each object is split in data-chunks parts, each stored on a different OSD.
Type
Integer
Required
Yes.
Example
4
m
Description
Compute coding chunks for each object and store them on different OSDs. The number of coding chunks is also the number of OSDs that can be down without losing data.
Type
Integer
Required
Yes.
Example
2
technique
Description
The more flexible technique is reed_sol_van; it is enough to set k and m. The cauchy_good technique can be faster but you need to choose the packetsize carefully. All of
reed_sol_r6_op, liberation, blaum_roth, liber8tion are RAID6 equivalents in the sense that they can only be configured with m=2.
Type
String
Required
No.
Valid Settings
reed_sol_van reed_sol_r6_op cauchy_orig cauchy_good liberation blaum_roth liber8tion
Default
reed_sol_van
packetsize
Description
The encoding will be done on packets of bytes size at a time. Choosing the correct packet size is difficult. The jerasure documentation contains extensive information on
this topic.
Type
Integer
Required
No.
Default
2048
crush-root
Description
The name of the CRUSH bucket used for the first step of the rule. For instance step take default.
Type
String
Required
No.
Default
default
crush-failure-domain
Description
Ensure that no two chunks are in a bucket with the same failure domain. For instance, if the failure domain is host no two chunks will be stored on the same host. It is used
to create a rule step such as step chooseleaf host.
Type
String
Required
No.
IBM Storage Ceph 125
Default
host
directory
Description
Set the directory name from which the erasure code plug-in is loaded.
Type
String
Required
No.
Default
/usr/lib/ceph/erasure-code
--force
Description
Override an existing profile by the same name.
Type
String
Required
No.
Controlling CRUSH Placement
The default CRUSH rule provides OSDs that are on different hosts. For instance:
chunk nr
01234567
step 1
step 2
step 3
_cDD_cDD
cDDD____
____cDDD
needs exactly 8 OSDs, one for each chunk. If the hosts are in two adjacent racks, the first four chunks can be placed in the first rack and the last four in the second rack.
Recovering from the loss of a single OSD does not require using bandwidth between the two racks.
For instance:
crush-steps='[ [ "choose", "rack", 2 ], [ "chooseleaf", "host", 4 ] ]'
creates a rule that selects two crush buckets of type rack and for each of them choose four OSDs, each of them located in a different bucket of type host.
The rule can also be created manually for finer control.
Installing
This information provides instructions on installing IBM Storage Ceph on Red Hat Enterprise Linux running on AMD64 and Intel 64 architectures.
Installing Pro Edition for free
Install IBM Storage Ceph Pro Edition for a free 60-day trial.
Initial installation
Managing an IBM Storage Ceph cluster using cephadm-ansible modules
Comparison between Ceph Ansible and Cephadm
Understand the differences between Cephadm and Ceph-Ansible playbooks for the containerized deployment of the storage cluster.
What to do next? Day 2
Installing Pro Edition for free
Install IBM Storage Ceph Pro Edition for a free 60-day trial.
About this task
You can install an IBM Storage Ceph trial version by following the steps provided.
Procedure
1. Acquire and install Red Hat Enterprise Linux (RHEL) 9.
a. Create and or log in to your Red Hat account.
b. Acquire a Red Hat Enterprise Linux 9 subscription from Red Hat for free here.
i. Confirm that your version meets the Operating system requirements for IBM Storage Ceph 6.
c. Install Red Hat Enterprise Linux 9 on your storage nodes or hosts.
2. Install IBM Storage Ceph Pro Edition.
126 IBM Storage Ceph
a. Register for a free 60-day IBM Storage Ceph Pro trial here.
b. Log in to your account with the username and password from the previous step and obtain your IBM entitlement key from the container software library in My
IBM.
c. Register your Red Hat Enterprise Linux nodes and enable the IBM Storage Ceph repositories by following the steps mentioned here.
d. Run dnf install cephadm -y command to install cephadm.
e. Run the preflight playbook.
f. Bootstrap the IBM Storage Ceph cluster that uses cephadm.
g. Use the dashboard to expand your cluster with OSD, Ceph Object Gateway, S3, and other daemons and services.
Initial installation
As a storage administrator, you can use the cephadm utility to deploy new IBM Storage Ceph clusters.
The cephadm utility manages the entire life cycle of a Ceph cluster. Installation and management tasks comprise two types of operations:
Day One operations involve installing and bootstrapping a bare-minimum, containerized Ceph storage cluster, running on a single node. Day One also includes
deploying the Monitor and Manager daemons and adding Ceph OSDs.
Day Two operations use the Ceph orchestration interface, cephadm orch, or the IBM Storage Ceph Dashboard to expand the storage cluster by adding other Ceph
services to the storage cluster.
cephadm utility
Registering the IBM Storage Ceph nodes
Register the IBM Storage Ceph nodes on Red Hat Enterprise Linux.
Configuring Ansible inventory location
Creating an Ansible user with sudo access
Configuring SSH
Enabling password-less SSH for Ansible
Enabling SSH login as root user on Red Hat Enterprise Linux 9
Running the preflight playbook
Bootstrapping a new storage cluster
Distributing SSH keys
You can use the cephadm-distribute-ssh-key.yml playbook to distribute the SSH keys instead of creating and distributing the keys manually.
Disconnected installation
Launching the cephadm shell
cephadm commands
The cephadm is a command line tool to manage the local host for the Cephadm Orchestrator. It provides commands to investigate and modify the state of the
current host.
Verifying the cluster installation
Adding hosts
Removing hosts
Labeling hosts
Adding Monitor service
Setting up the admin node
Adding Manager service
Adding OSDs
Purging the Ceph storage cluster
Purging the Ceph storage cluster clears any data or connections that remain from previous deployments on your server.
Deploying client nodes
Prerequisites
At least one running virtual machine (VM) or bare-metal server with an active internet connection.
Ansible 2.9 or later.
Server and Ceph repositories enabled.
Root-level access to all nodes.
An active IBM Network or service account to access the IBM Registry.
Remove troubling configurations in iptables so that refresh of iptables services does not cause issues to the cluster. For an example, see Verifying firewall rules are
configured for default Ceph ports.
For the latest supported Red Hat Enterprise Linux versions, see Compatibility matrix.
cephadm utility
The cephadm utility deploys and manages a Ceph storage cluster. It is tightly integrated with both the command-line interface (CLI) and the IBM Storage Ceph Dashboard
web interface, so that you can manage storage clusters from either environment. cephadm uses SSH to connect to hosts from the manager daemon to add, remove, or
update Ceph daemon containers. It does not rely on external configuration or orchestration tools such as Ansible or Rook.
Note: The cephadm utility is available after running the preflight playbook on a host.
The cephadm utility consists of two main components:
The cephadm shell.
IBM Storage Ceph 127
The cephadm orchestrator.
The cephadm shell
The cephadm shell launches a bash shell within a container. This enables you to perform “Day One” cluster setup tasks, such as installation and bootstrapping, and to
invoke ceph commands.
Note: If the node contains configuration and keyring files in /etc/ceph/, the container environment uses the values in those files as defaults for the cephadm shell.
However, if you execute the cephadm shell on a Ceph Monitor node, the cephadm shell inherits its default configuration from the Ceph Monitor container, instead of using
the default configuration.
The cephadm orchestrator
The cephadm orchestrator enables you to perform “Day Two” Ceph functions, such as expanding the storage cluster and provisioning Ceph daemons and services. You
can use the cephadm orchestrator through either the command-line interface (CLI) or the web-based IBM Storage Ceph Dashboard. Orchestrator commands take the
form ceph
orch.
The cephadm script interacts with the Ceph orchestration module used by the Ceph Manager.
How cephadm works
cephadm-ansible playbooks
How cephadm works
The cephadm command manages the full lifecycle of an IBM Storage Ceph cluster. The cephadm command can perform the following operations:
Bootstrap a new IBM Storage Ceph cluster.
Launch a containerized shell that works with the IBM Storage Ceph command-line interface (CLI).
Aid in debugging containerized daemons.
The cephadm command uses ssh to communicate with the nodes in the storage cluster. This allows you to add, remove, or update IBM Storage Ceph containers without
using external tools. Generate the ssh key pair during the bootstrapping process, or use your own ssh key.
The cephadm bootstrapping process creates a small storage cluster on a single node, consisting of one Ceph Monitor and one Ceph Manager, as well as any required
dependencies. You then use the orchestrator CLI or the IBM Storage Ceph Dashboard to expand the storage cluster to include nodes, and to provision all of the IBM
Storage Ceph daemons and services. You can perform management functions through the CLI or from the IBM Storage Ceph Dashboard web interface.
Figure 1. Ceph storage cluster deployment
128 IBM Storage Ceph
cephadm-ansible playbooks
The cephadm-ansible package is a collection of Ansible playbooks to simplify workflows that are not covered by cephadm. After installation, the playbooks are located
in /usr/share/cephadm-ansible/.
The cephadm-ansbile package includes the following playbooks:
cephadm-preflight.yml
cephadm-clients.yml
cephadm-purge-cluster.yml
The cephadm-preflight playbook
Use the cephadm-preflight playbook to initially setup hosts before bootstrapping the storage cluster and before adding new nodes or clients to your storage cluster.
This playbook configures the Ceph repository and installs some prerequisites such as podman, lvm2, chrony, and cephadm.
For more information, see Running the preflight playbook.
The cephadm-clients playbook
Use the cephadm-clients playbook to set up client hosts. This playbook handles the distribution of configuration and keyring files to a group of Ceph clients.
The cephadm-purge-cluster playbook
Use the cephadm-purge-cluster playbook to remove a Ceph cluster. This playbook purges a Ceph cluster managed with cephadm.
For more information, see Purging the Ceph storage cluster.
Registering the IBM Storage Ceph nodes
Register the IBM Storage Ceph nodes on Red Hat Enterprise Linux.
Prerequisites
At least one running virtual machine (VM) or bare-metal server with an active internet connection.
A valid IBM subscription with the appropriate entitlements.
Root-level access to all nodes.
For the latest supported Red Hat Enterprise Linux versions, see Compatibility matrix.
Procedure
1. Register the system, and when prompted, enter your Red Hat customer portal credentials:
Example
[root@admin ~]# subscription-manager register
2. Pull the latest subscription:
[root@admin ~]# subscription-manager refresh
3. Identify the appropriate subscription and retrieve its Pool ID.
4. Attach the pool ID to gain access to the software entitlements. Use the Pool ID you identified in the previous step.
Example
[root@admin ~]# subscription-manager attach --pool=POOL_ID
5. Disable the software repositories:
Example
[root@admin ~]# subscription-manager repos --disable=*
6. Enable the Red Hat Enterprise Linux baseos and appstream repositories:
Example
[root@admin ~]# subscription-manager repos --enable=rhel-9-for-x86_64-baseos-rpms
[root@admin ~]# subscription-manager repos --enable=rhel-9-for-x86_64-appstream-rpms
7. Update the system:
Example
[root@admin ~]# dnf update
8. Enable the ceph-tools repository:
IBM Storage Ceph 129
[root@admin ~]# curl https://public.dhe.ibm.com/ibmdl/export/pub/storage/ceph/ibm-storage-ceph-6-rhel-9.repo | sudo tee
/etc/yum.repos.d/ibm-storage-ceph-6-rhel-9.repo
Repeat the above steps on all the nodes of the storage cluster.
9. Add license to install IBM Storage Ceph and click "Accept":
Example
[root@admin ~]# dnf install ibm-storage-ceph-license
10. Accept these provisions:
Example
[root@admin ~]# sudo touch /usr/share/ibm-storage-ceph-license/accept
11. Install cephadm-ansible:
[root@admin ~]# dnf install cephadm-ansible
Configuring Ansible inventory location
You can configure inventory location files for the cephadm-ansible staging and production environments. The Ansible inventory hosts file contains all the hosts that are
part of the storage cluster. You can list nodes individually in the inventory hosts file or you can create groups such as [mons],[osds], and [rgws] to provide clarity to
your inventory and ease the usage of the --limit option to target a group or node when running a playbook.
Note: If deploying clients, client nodes must be defined in a dedicated [clients] group.
Prerequisites
An Ansible administration node.
Root-level access to the Ansible administration node.
The cephadm-ansible package is installed on the node.
Procedure
1. Navigate to the /usr/share/cephadm-ansible/ directory:
[root@admin ~]# cd /usr/share/cephadm-ansible
2. Optional: Create subdirectories for staging and production:
[root@admin cephadm-ansible]# mkdir -p inventory/staging inventory/production
3. Optional: Edit the ansible.cfg file and add the following line to assign a default inventory location:
[defaults]
inventory = ./inventory/staging
4. Optional: Create an inventory hosts file for each environment:
[root@admin cephadm-ansible]# touch inventory/staging/hosts
[root@admin cephadm-ansible]# touch inventory/production/hosts
5. Open and edit each hosts file and add the nodes and [admin] group:
Syntax
NODE_NAME_1
NODE_NAME_2
[admin]
ADMIN_NODE_NAME_1
Replace NODE_NAME_1 and NODE_NAME_2 with the Ceph nodes such as monitors, OSDs, MDSs, and gateway nodes.
Replace ADMIN_NODE_NAME_1 with the name of the node where the admin keyring is stored.
Example
host02
host03
host04
[admin]
host01
Note: If you set the inventory location in the ansible.cfg file to staging, you need to run the playbooks in the staging environment as follows:
Syntax
ansible-playbook -i inventory/staging/hosts PLAYBOOK.yml
To run the playbooks in the production environment:
130 IBM Storage Ceph
Syntax
ansible-playbook -i inventory/production/hosts PLAYBOOK.yml
Creating an Ansible user with sudo access
You can create an Ansible user with password-less root access on all nodes in the storage cluster to run the cephadm-ansible playbooks. The Ansible user must be
able to log into all the IBM Storage Ceph nodes as a user that has root privileges to install software and create configuration files without prompting for a password.
Important: If you are using Red Hat Enterprise Linux 9 only use these steps if you are a non-root user. If you are a root user, go to Bootstrapping a new storage cluster.
Prerequisites
Root-level access to all nodes.
- For Red Hat Enterprise 9, to log in as a root user, see Enabling SSH login as root user on Red Hat Enterprise Linux 9
Procedure
1. Log in to the node as the root user:
Syntax
ssh root@HOST_NAME
Replace HOST_NAME with the host name of the Ceph node.
Example
[root@admin ~]# ssh root@host01
Enter the root password when prompted.
2. Create a new Ansible user:
Syntax
adduser USER_NAME
Replace USER_NAME with the new user name for the Ansible user.
Example
[root@host01 ~]# adduser ceph-admin
Important: Do not use ceph as the user name. The ceph user name is reserved for the Ceph daemons. A uniform user name across the cluster can improve ease of
use, but avoid using obvious user names, because intruders typically use them for brute-force attacks.
3. Set a new password for this user:
Syntax
passwd USER_NAME
Replace USER_NAME with the new user name for the Ansible user.
Example
[root@host01 ~]# passwd ceph-admin
Enter the new password twice when prompted.
4. Configure sudo access for the newly created user:
Syntax
cat << EOF >/etc/sudoers.d/USER_NAME
$USER_NAME ALL = (root) NOPASSWD:ALL
EOF
Replace USER_NAME with the new user name for the Ansible user.
Example
[root@host01 ~]# cat << EOF >/etc/sudoers.d/ceph-admin
ceph-admin ALL = (root) NOPASSWD:ALL
EOF
5. Assign the correct file permissions to the new file:
Syntax
chmod 0440 /etc/sudoers.d/USER_NAME
Replace USER_NAME with the new user name for the Ansible user.
Example
IBM Storage Ceph 131
[root@host01 ~]# chmod 0440 /etc/sudoers.d/ceph-admin
6. Repeat the above steps on all nodes in the storage cluster.
Reference
For more information about creating user accounts, see Configuring basic system settings > Getting started with managing user accounts within the Red Hat
Enterprise Linux guide.
Configuring SSH
As a storage administrator, with Cephadm, you can use an SSH key to securely authenticate with remote hosts. The SSH key is stored in the monitor to connect to remote
hosts.
Configuring a different SSH user
Prerequisites
A running IBM Storage Ceph cluster.
An Ansible administration node.
Root-level access to the Ansible administration node.
The cephadm-ansible package is installed on the node.
Procedure
1. Navigate to the cephadm-ansible directory.
2. Generate a new SSH key:
Example
[ceph-admin@admin cephadm-ansible]$ ceph cephadm generate-key
3. Retrieve the public portion of the SSH key:
Example
[ceph-admin@admin cephadm-ansible]$ ceph cephadm get-pub-key
4. Delete the currently stored SSH key:
Example
[ceph-admin@admin cephadm-ansible]$ceph cephadm clear-key
5. Restart the mgr daemon to reload the configuration:
Example
[ceph-admin@admin cephadm-ansible]$ ceph mgr fail
Configuring a different SSH user
As a storage administrator, you can configure a non-root SSH user who can log in to all the Ceph cluster nodes with enough privileges to download container images, start
containers, and run commands without prompting for a password.
Important: Before configuring a non-root SSH user, the cluster SSH key needs to be added to the user's authorized_keys file and non-root users must have
passwordless sudo access.
Prerequisites
A running IBM Storage Ceph cluster.
An Ansible administration node.
Root-level access to the Ansible administration node.
The cephadm-ansible package is installed on the node.
Add the cluster SSH keys to the user's authorized_keys.
Enable passwordless sudo access for the non-root users.
Procedure
132 IBM Storage Ceph
1. Go to the cephadm-ansible directory.
2. Provide Cephadm the name of the user who is going to perform all the Cephadm operations.
Syntax
ceph cephadm set-user USER
Example
[ceph-admin@admin cephadm-ansible]$ ceph cephadm set-user user
3. Retrieve the SSH public key.
Syntax
ceph cephadm get-pub-key > ~/ceph.pub
Example
[ceph-admin@admin cephadm-ansible]$ ceph cephadm get-pub-key > ~/ceph.pub
4. Copy the SSH keys to all the hosts.
Syntax
ssh-copy-id -f -i ~/ceph.pub USER@HOST
Example
[ceph-admin@admin cephadm-ansible]$ ssh-copy-id ceph-admin@host01
Enabling password-less SSH for Ansible
Generate an SSH key pair on the Ansible administration node and distribute the public key to each node in the storage cluster so that Ansible can access the nodes
without being prompted for a password.
Important: If you are using Red Hat Enterprise Linux 9 only use these steps if you are a non-root user. If you are a root user, go to Bootstrapping a new storage cluster.
Prerequisites
Access to the Ansible administration node.
Ansible user with sudo access to all nodes in the storage cluster.
- For Red Hat Enterprise 9, to log in as a root user, see Enabling SSH login as root user on Red Hat Enterprise Linux 9
Procedure
1. Generate the SSH key pair, accept the default file name and leave the passphrase empty:
[ansible@admin ~]$ ssh-keygen
2. Copy the public key to all nodes in the storage cluster:
ssh-copy-id USER_NAME@HOST_NAME
Replace USER_NAME with the new user name for the Ansible user. Replace HOST_NAME with the host name of the Ceph node.
Example
[ansible@admin ~]$ ssh-copy-id ceph-admin@host01
3. Create the user’s SSH config file:
[ansible@admin ~]$ touch ~/.ssh/config
4. Open for editing the config file. Set values for the Hostname and User options for each node in the storage cluster:
Syntax
Host host01
Hostname HOST_NAME
User USER_NAME
Host host02
Hostname HOST_NAME
User USER_NAME
...
Replace HOST_NAME with the host name of the Ceph node. Replace USER_NAME with the new user name for the Ansible user.
Example
Host host01
Hostname host01
User ceph-admin
Host host02
Hostname host02
User ceph-admin
IBM Storage Ceph 133
Host host03
Hostname host03
User ceph-admin
Important: By configuring the ~/.ssh/config file you do not have to specify the -u _USER_NAME_ option each time you execute the ansible-playbook
command.
5. Set the correct file permissions for the ~/.ssh/config file:
[ansible@admin ~]$ chmod 600 ~/.ssh/config
Reference
The ssh_config(5) manual page.
See Using secure communications between two systems with OpenSSH.
Enabling SSH login as root user on Red Hat Enterprise Linux 9
Red Hat Enterprise Linux 9 does not support SSH login as a root user even if PermitRootLogin parameter is set to yes in the /etc/ssh/sshd_config file. You get the
following error:
Example
[root@host01 ~]# ssh root@myhostname
root@myhostname password:
Permission denied, please try again.
You can run one of the following methods to enable login as a root user:
Use "Allow root SSH login with password" flag while setting the root password during installation of Red Hat Enterprise Linux 9.
Manually set the PermitRootLogin parameter after Red Hat Enterprise Linux 9 installation.
This section describes manual setting of the PermitRootLogin parameter.
Prerequisites
Root-level access to all nodes.
Procedure
1. Open the etc/ssh/sshd_config file and set the PermitRootLogin to yes:
Example
[root@admin ~]# echo 'PermitRootLogin yes' >> /etc/ssh/sshd_config.d/01-permitrootlogin.conf
2. Restart the SSH service:
Example
[root@admin ~]# systemctl restart sshd.service
3. Login to the node as the root user:
Syntax
ssh root@HOST_NAME
Replace HOST_NAME with the host name of the Ceph node.
Example
[root@admin ~]# ssh root@host01
Enter the root password when prompted.
Reference
For more information, see Not able to login as root user via ssh in RHEL 9 server
Running the preflight playbook
This Ansible playbook configures the Ceph repository and prepares the storage cluster for bootstrapping. It also installs some prerequisites, such as podman, lvm2,
chrony, and cephadm. The default location for cephadm-ansible and cephadm-preflight.yml is /usr/share/cephadm-ansible.
The preflight playbook uses the cephadm-ansible inventory file to identify all the admin and nodes in the storage cluster.
The default location for the inventory file is /usr/share/cephadm-ansible/hosts. The following example shows the structure of a typical inventory file:
134 IBM Storage Ceph
Example
host02
host03
host04
[admin]
host01
The [admin] group in the inventory file contains the name of the node where the admin keyring is stored. On a new storage cluster, the node in the [admin] group will be
the bootstrap node. To add additional admin hosts after bootstrapping the cluster see Setting up the admin node.
Important: If you are performing a disconnected installation, see Running the preflight playbook for a disconnected installation.
Note: Run the preflight playbook before you bootstrap the initial host.
Prerequisites
Root-level access to the Ansible administration node.
Ansible user with sudo and passwordless ssh access to all nodes in the storage cluster.
NOTE: In the below example, host01 is the bootstrap node.
Procedure
1. Navigate to the the /usr/share/cephadm-ansible directory.
2. Open and edit the hosts file and add your nodes:
Example
host02
host03
host04
[admin]
host01
3. Add license to install IBM Storage Ceph and click Accept on all nodes:
Example
[root@admin ~]# dnf install ibm-storage-ceph-license
a. Accept these provisions:
Example
[root@admin ~]# sudo touch /usr/share/ibm-storage-ceph-license/accept
4. Run the preflight playbook:
Syntax
ansible-playbook -i INVENTORY_FILE cephadm-preflight.yml --extra-vars "ceph_origin=ibm"
Example
[ansible@admin cephadm-ansible]$ ansible-playbook -i hosts cephadm-preflight.yml --extra-vars "ceph_origin=ibm"
After installation is complete, cephadm resides in the /usr/sbin/ directory.
Use the --limit option to run the preflight playbook on a selected set of hosts in the storage cluster:
Syntax
ansible-playbook -i INVENTORY_FILE cephadm-preflight.yml --extra-vars "ceph_origin=ibm" --limit GROUP_NAME|NODE_NAME
Replace GROUP_NAME with a group name from your inventory file. Replace NODE_NAME with a specific node name from your inventory file.
NOTE: Optionally, you can group your nodes in your inventory file by group name such as [mons], [osds], and [mgrs]. However, admin nodes must be
added to the [admin] group and clients must be added to the [clients] group.
Example
[ansible@admin cephadm-ansible]$ ansible-playbook -i hosts cephadm-preflight.yml --extra-vars "ceph_origin=ibm" -limit clients
[ansible@admin cephadm-ansible]$ ansible-playbook -i hosts cephadm-preflight.yml --extra-vars "ceph_origin=ibm" -limit host01
When you run the preflight playbook, cephadm-ansible automatically installs chrony and ceph-common on the client nodes.
If you want to configure multiple sources or if you have a disconnected environment, see the following documentation for more information:
How to install chrony?
Best practices for NTP
Basic chrony NTP troubleshooting
Bootstrapping a new storage cluster
IBM Storage Ceph 135
The cephadm utility performs the following tasks during the bootstrap process:
Installs and starts a Ceph Monitor daemon and a Ceph Manager daemon for a new IBM Storage Ceph cluster on the local node as containers.
Creates the /etc/ceph directory.
Writes a copy of the public key to /etc/ceph/ceph.pub for the IBM Storage Ceph cluster and adds the SSH key to the root user’s
/root/.ssh/authorized_keys file.
Applies the _admin label to the bootstrap node.
Writes a minimal configuration file needed to communicate with the new cluster to /etc/ceph/ceph.conf.
Writes a copy of the client.admin administrative secret key to /etc/ceph/ceph.client.admin.keyring.
Deploys a basic monitoring stack with Prometheus, grafana, and other tools such as node-exporter and alert-manager.
Important:
If you are performing a disconnected installation, see Performing a disconnected installation.
If you are deploying a monitoring stack, see Deploying the monitoring stack using the Ceph Orchestrator.
Bootstrapping provides the default user name and password for the initial login to the Dashboard. Bootstrap requires you to change the password after you log in.
Before you begin the bootstrapping process, make sure that the container image that you want to use has the same version of IBM Storage Ceph as cephadm. If the
two versions do not match, bootstrapping fails at the Creating initial admin user stage.
Note:
If you have existing prometheus services that you want to run with the new storage cluster, or if you are running Ceph with Rook, use the --skip-monitoringstack option with the cephadm bootstrap command. This option bypasses the basic monitoring stack so that you can manually configure it later.
Before you begin the bootstrapping process, you must create a username and password for the cp.icr.io/cp container registry.
Recommended cephadm bootstrap command options
Using a JSON file to protect login information
Bootstrapping a storage cluster using a service configuration file
Bootstrapping the storage cluster as a non-root user
Bootstrap command options
The cephadm bootstrap command bootstraps a Ceph storage cluster on the local host. It deploys a MON daemon and a MGR daemon on the bootstrap node,
automatically deploys the monitoring stack on the local host, and calls ceph orch host add HOSTNAME.
Obtaining entitlement key
Entitlement keys determine whether IBM Storage Ceph can automatically pull the required container default images. During installation, image pull failures can
occur due to an invalid entitlement key or a key belonging to an account that does not have entitlement to IBM Storage Ceph.
Prerequisites
An IP address for the first Ceph Monitor container, which is also the IP address for the first node in the storage cluster.
Login access to cp.icr.io/cp. For information about obtaining credentials for cp.icr.io/cp, see Obtaining entitlement key.
A minimum of 10 GB of free space for /var/lib/containers/.
Root-level access to all nodes.
Important:
Run cephadm bootstrap on the node that you want to be the initial Monitor node in the cluster. The IP_ADDRESS option should be the IP address of the node you
are using to run cephadm bootstrap.
Note:
If the storage cluster includes multiple networks and interfaces, be sure to choose a network that is accessible by any node that uses the storage cluster.
If the local node uses fully-qualified domain names (FQDN), then add the --allow-fqdn-hostname option to cephadm bootstrap on the command line.
If you want to deploy a storage cluster using IPV6 addresses, then use the IPV6 address format for the --mon-ip IP_ADDRESS option. For example: cephadm
bootstrap --mon-ip
2620:52:0:880:225:90ff:fefc:2536 --registry-json /etc/mylogin.json
Procedure
1. Bootstrap a storage cluster:
Syntax
cephadm bootstrap --cluster-network NETWORK_CIDR --mon-ip IP_ADDRESS --registry-url cp.icr.io/cp --registry-username
USER_NAME --registry-password PASSWORD --yes-i-know
The USER_NAME is cp and the PASSWORD is the token generated while obtaining the entitlement key.
Example
[root@host01 ~]# cephadm bootstrap --cluster-network 10.10.128.0/24 --mon-ip 10.10.128.68 --registry-url cp.icr.io/cp -registry-username cp --registry-password mypassword1 --yes-i-know
Note: If you want internal cluster traffic routed over the public network, you can omit the --cluster-network NETWORK_CIDR option.
The script takes a few minutes to complete. Once the script completes, it provides the credentials to the IBM Storage Ceph Dashboard URL, a command to access
the Ceph command-line interface (CLI), and a request to enable telemetry.
136 IBM Storage Ceph
Ceph Dashboard is now available at:
URL: https://host01:8443/
User: admin
Password: i8nhu7zham
Enabling client.admin keyring and conf on hosts with "admin" label
You can access the Ceph CLI with:
sudo /usr/sbin/cephadm shell --fsid 266ee7a8-2a05-11eb-b846-5254002d4916 -c /etc/ceph/ceph.conf -k
/etc/ceph/ceph.client.admin.keyring
Please consider enabling telemetry to help improve Ceph:
ceph telemetry on
For more information see:
https://docs.ceph.com/docs/master/mgr/telemetry/
Bootstrap complete.
Reference
Recommended cephadm bootstrap command options
Bootstrap command options
Using a JSON file to protect login information
Recommended cephadm bootstrap command options
The cephadm bootstrap command has multiple options that allow you to specify file locations, configure ssh settings, set passwords, and perform other initial
configuration tasks.
IBM recommends that you use a basic set of command options for cephadm
bootstrap. You can configure additional options after your initial cluster is up and running.
The following examples show how to specify the recommended options.
Syntax
cephadm bootstrap --ssh-user USER_NAME --mon-ip IP_ADDRESS --allow-fqdn-hostname --registry-json REGISTRY_JSON
Example
[root@host01 ~]# cephadm bootstrap --ssh-user ceph --mon-ip 10.10.128.68 --allow-fqdn-hostname --registry-json
/etc/mylogin.json
For non-root users, see Creating an Ansible user with sudo access and Enabling password-less SSH for Ansible for more details.
Reference
For more information about the --registry-json option, see Using a JSON file to protect login information.
For more information about all available cephadm bootstrap options, see Bootstrap command options.
For more information about bootstrapping the storage cluster as a non-root user, see Bootstrapping the storage cluster as a non-root user.
Using a JSON file to protect login information
As a storage administrator, you might choose to add login and password information to a JSON file, and then refer to the JSON file for bootstrapping. This protects the
login credentials from exposure.
NOTE: You can also use a JSON file with the cephadm --registry-login command.
Prerequisites
An IP address for the first Ceph Monitor container, which is also the IP address for the first node in the storage cluster.
Login access to cp.icr.io/cp.
A minimum of 10 GB of free space for /var/lib/containers/.
Root-level access to all nodes.
Procedure
1. Create the JSON file. In this example, the file is named mylogin.json.
IBM Storage Ceph 137
Syntax
{
"url":"REGISTRY_URL",
"username":"USER_NAME",
"password":"PASSWORD"
}
Example
{
"url":"cp.icr.io/cp",
"username":"myuser1",
"password":"mypassword1"
}
2. Bootstrap a storage cluster:
Syntax
cephadm bootstrap --mon-ip IP_ADDRESS --registry-json /etc/mylogin.json
Example
[root@host01 ~]# cephadm bootstrap --mon-ip 10.10.128.68 --registry-json /etc/mylogin.json
Bootstrapping a storage cluster using a service configuration file
To bootstrap the storage cluster and configure additional hosts and daemons using a service configuration file, use the --apply-spec option with the cephadm
bootstrap command. The configuration file is a .yaml file that contains the service type, placement, and designated nodes for services that you want to deploy.
Note: If you want to use a non-default realm or zone for applications such as multi-site, configure your Ceph Object Gateway daemons after you bootstrap the storage
cluster, instead of adding them to the configuration file and using the --apply-spec option. This gives you the opportunity to create the realm or zone you need for the
Ceph Object Gateway daemons before deploying them.
Note: If deploying a NFS-Ganesha gateway, or Metadata Server (MDS) service, configure them after bootstrapping the storage cluster.
To deploy a Ceph NFS-Ganesha gateway, you must create a RADOS pool first.
To deploy the MDS service, you must create a CephFS volume first.
Note: If you run the bootstrap command with --apply-spec option, ensure to include the IP address of the bootstrap host in the specification file. This prevents
resolving the IP address to loopback address while re-adding the bootstrap host where active Ceph Manager is already running. If you do not use the --apply spec
option during bootstrap and instead use ceph orch apply command with another specification file which includes re-adding the host and contains an active Ceph
Manager running, then ensure to explicitly provide the addr field. This is applicable for applying any specification file after bootstrapping.
For more information, see Operations.
Prerequisites
At least one running virtual machine (VM) or server.
Root-level access to all nodes.
Login access to cp.icr.io/cp.
Passwordless ssh is set up on all hosts in the storage cluster.
cephadm is installed on the node that you want to be the initial Monitor node in the storage cluster.
For the latest supported Red Hat Enterprise Linux versions, see Compatibility matrix.
Procedure
1. Log in to the bootstrap host.
2. Create the service configuration .yaml file for your storage cluster. The example file directs cephadm bootstrap to configure the initial host and two additional
hosts, and it specifies that OSDs be created on all available disks.
Example
service_type: host
addr: host01
hostname: host01
--service_type: host
addr: host02
hostname: host02
--service_type: host
addr: host03
hostname: host03
--service_type: host
addr: host04
hostname: host04
--service_type: mon
138 IBM Storage Ceph
placement:
host_pattern: "host[0-2]"
--service_type: osd
service_id: my_osds
placement:
host_pattern: "host[1-3]"
data_devices:
all: true
3. Bootstrap the storage cluster with the --apply-spec option:
Syntax
cephadm bootstrap --apply-spec CONFIGURATION_FILE_NAME --mon-ip MONITOR_IP_ADDRESS --registry-url cp.icr.io/cp --registryusername USER_NAME --registry-password PASSWORD
Example
[root@host01 ~]# cephadm bootstrap --apply-spec initial-config.yaml --mon-ip 10.10.128.68 --registry-url cp.icr.io/cp -registry-username myuser1 --registry-password mypassword1
The script takes a few minutes to complete. Once the script completes, it provides the credentials to the IBM Storage Ceph Dashboard URL, a command to access
the Ceph command-line interface (CLI), and a request to enable telemetry.
Once your storage cluster is up and running, see Operations for more information about configuring additional daemons and services.
Reference
Bootstrap command options
Bootstrapping the storage cluster as a non-root user
You can bootstrap the storage cluster as a non-root user if you have passwordless sudo privileges.
To bootstrap the IBM Storage Ceph cluster as a non-root user on the bootstrap node, use the --ssh-user option with the cephadm bootstrap command. --ssh-user
specifies a user for SSH connections to cluster nodes.
Non-root users must have passwordless sudo access. For more information, see Creating an Ansible user with sudo access and Enabling password-less SSH for Ansible.
Prerequisites
An IP address for the first Ceph Monitor container, which is also the IP address for the initial Monitor node in the storage cluster.
Login access to cp.icr.io/cp.
A minimum of 10 GB of free space for /var/lib/containers/.
Optional: SSH public and private keys.
Passwordless sudo access to the bootstrap node.
Non-root users have passwordless sudo access on all nodes intended to be part of the cluster.
Procedure
1. Change to sudo on the bootstrap node:
Syntax
su - SSH_USER_NAME
Example
[root@host01 ~]$ su - ceph
Last login: Tue Sep 14 12:00:29 EST 2021 on pts/0
2. Check the SSH connection to the bootstrap node:
Example
[ceph@host01 ~]# ssh host01
Last login: Tue Sep 14 12:03:29 EST 2021 on pts/0
3. Optional: Invoke the cephadm bootstrap command.
Note: Using private and public keys is optional.
If SSH keys have not previously been created, these can be created during this step.
Include the --ssh-private-key and --ssh-public-key options:
Syntax
sudo cephadm bootstrap --ssh-user USER_NAME --mon-ip IP_ADDRESS --ssh-private-key PRIVATE_KEY --ssh-public-key PUBLIC_KEY
--registry-url cp.icr.io/cp --registry-username USER_NAME --registry-password PASSWORD
IBM Storage Ceph 139
Example
sudo cephadm bootstrap --ssh-user ceph --mon-ip 10.10.128.68 --ssh-private-key /home/ceph/.ssh/id_rsa --ssh-public-key
/home/ceph/.ssh/id_rsa.pub --registry-url cp.icr.io/cp --registry-username myuser1 --registry-password mypassword1
Reference
Bootstrap command options
For more information about utilizing Ansible to automate bootstrapping a rootless cluster, see the knowledge base article Red Hat Ceph Storage 5.3 rootless
deployment utilizing ansible ad-hoc commands.
For more information about sudo privileges, see Managing sudo access
Bootstrap command options
The cephadm bootstrap command bootstraps a Ceph storage cluster on the local host. It deploys a MON daemon and a MGR daemon on the bootstrap node,
automatically deploys the monitoring stack on the local host, and calls ceph orch host add HOSTNAME.
Table 1 lists the available options for cephadm bootstrap.
Table 1. cephadm bootstrap command options
cephadm bootstrap option
Description
--config CONFIG_FILE, -c CONFIG_FILE
CONFIG_FILE is the ceph.conf file to use with the bootstrap command
--cluster-network NETWORK_CIDR
Use the subnet defined by NETWORK_CIDR for internal cluster traffic. This is specified in CIDR notation. For
example: 10.10.128.0/24.
--mon-id MON_ID
Bootstraps on the host named MON_ID. Default value is the local host.
--mon-addrv MON_ADDRV
mon IPs (e.g., [v2:localipaddr:3300,v1:localipaddr:6789])
--mon-ip IP_ADDRESS
IP address of the node you are using to run cephadm bootstrap.
--mgr-id MGR_ID
Host ID where a MGR node should be installed. Default: randomly generated.
--fsid FSID
Cluster FSID.
--output-dir OUTPUT_DIR
Use this directory to write config, keyring, and pub key files.
--output-keyring OUTPUT_KEYRING
Use this location to write the keyring file with the new cluster admin and mon keys.
--output-config OUTPUT_CONFIG
Use this location to write the configuration file to connect to the new cluster.
--output-pub-ssh-key OUTPUT_PUB_SSH_KEY
Use this location to write the public SSH key for the cluster.
--skip-ssh
Skip the setup of the ssh key on the local host.
--initial-dashboard-user INITIAL_DASHBOARD_USER Initial user for the dashboard.
--initial-dashboard-password
INITIAL_DASHBOARD_PASSWORD
Initial password for the initial dashboard user.
--ssl-dashboard-port SSL_DASHBOARD_PORT
Port number used to connect with the dashboard using SSL.
--dashboard-key DASHBOARD_KEY
Dashboard key.
--dashboard-crt DASHBOARD_CRT
Dashboard certificate.
--ssh-config SSH_CONFIG
SSH config.
--ssh-private-key SSH_PRIVATE_KEY
SSH private key.
--ssh-public-key SSH_PUBLIC_KEY
SSH public key.
--ssh-user SSH_USER
Sets the user for SSH connections to cluster hosts. Passwordless sudo is needed for non-root users.
--skip-mon-network
Sets mon public_network based on the bootstrap mon ip.
--skip-dashboard
Do not enable the Ceph Dashboard.
--dashboard-password-noupdate
Disable forced dashboard password change.
--no-minimize-config
Do not assimilate and minimize the configuration file.
--skip-ping-check
Do not verify that the mon IP is pingable.
--skip-pull
Do not pull the latest image before bootstrapping.
--skip-firewalld
Do not configure firewalld.
--allow-overwrite
Allow the overwrite of existing –output-* config/keyring/ssh files.
--allow-fqdn-hostname
Allow fully qualified host name.
--skip-prepare-host
Do not prepare host.
--orphan-initial-daemons
Do not create initial mon, mgr, and crash service specs.
--skip-monitoring-stack
Do not automatically provision the monitoring stack] (prometheus, grafana, alertmanager, node-exporter).
--apply-spec APPLY_SPEC
Apply cluster spec file after bootstrap (copy ssh key, add hosts and apply services).
--registry-url REGISTRY_URL
Specifies the URL of the custom registry to log into. For example: cp.icr.io/cp.
--registry-username REGISTRY_USERNAME
User name of the login account to the custom registry.
--registry-password REGISTRY_PASSWORD
Password of the login account to the custom registry.
--registry-json REGISTRY_JSON
JSON file containing registry login information.
Reference
For more information about the --skip-monitoring-stack option, see Adding hosts.
For more information about logging into the registry with the registry-json option, see help for the registry-login command.
For more information about cephadm options, see help for cephadm.
140 IBM Storage Ceph
Obtaining entitlement key
Entitlement keys determine whether IBM Storage Ceph can automatically pull the required container default images. During installation, image pull failures can occur due
to an invalid entitlement key or a key belonging to an account that does not have entitlement to IBM Storage Ceph.
Procedure
1. Log in to the IBM container software library with the IBM ID and password that is associated with the entitled IBM Storage Ceph software.
2. In the navigation bar, click Get entitlement key.
3. On the Access your container software page, click Copy key to copy the generated entitlement key.
4. Save the key to a secure location for future use.
The user is cp while the key is the token which is the password.
5. Verify the login against the registry.
podman login -u cp -p TOKEN cp.icr.io/cp
Login Succeeded!
Distributing SSH keys
You can use the cephadm-distribute-ssh-key.yml playbook to distribute the SSH keys instead of creating and distributing the keys manually.
Before you begin
Ansible is installed on the administration node.
Access to the Ansible administration node.
Ansible user with sudo access to all nodes in the storage cluster.
Bootstrapping is completed. See Bootstrapping a new storage cluster for more details.
About this task
The playbook distributes an SSH public key over all hosts in the inventory. You can also generate an SSH key pair on the Ansible administration node and distribute the
public key to each node in the storage cluster so that Ansible can access the nodes without being prompted for a password.
Procedure
1. Navigate to the /usr/share/cephadm-ansible directory on the Ansible administration node.
[ansible@admin ~]$ cd /usr/share/cephadm-ansible
2. From the Ansible administration node, distribute the SSH keys. The optional cephadm_pubkey_path parameter is the full path name of the SSH public key file on
the ansible controller host.
Note:
If cephadm_pubkey_path is not specified, the playbook gets the key from the cephadm get-pub-key command. This implies that you have at least
bootstrapped a minimal cluster.
ansible-playbook -i INVENTORY_HOST_FILE cephadm-distribute-ssh-key.yml -e cephadm_ssh_user=USER_NAME -e
cephadm_pubkey_path= home/cephadm/ceph.key -e admin_node=ADMIN_NODE_NAME_1
[ansible@admin cephadm-ansible]$ ansible-playbook -i hosts cephadm-distribute-ssh-key.yml -e cephadm_ssh_user=ceph-admin e cephadm_pubkey_path=/home/cephadm/ceph.key -e admin_node=host01
[ansible@admin cephadm-ansible]$ ansible-playbook -i hosts cephadm-distribute-ssh-key.yml -e cephadm_ssh_user=ceph-admin e admin_node=host01
Disconnected installation
As a storage administrator, you can use a disconnected installation procedure to install cephadm and bootstrap your storage cluster on a private network. A disconnected
installation uses a private registry for installation.
Configuring a private registry for a disconnected installation
Use this procedure when the IBM Storage Ceph nodes do NOT have access to the internet during deployment.
Running the preflight playbook for a disconnected installation
Performing a disconnected installation
Changing configurations of custom container images for disconnected installations
Adding hosts in disconnected deployments
Configuring a private registry for a disconnected installation
IBM Storage Ceph 141
Use this procedure when the IBM Storage Ceph nodes do NOT have access to the internet during deployment.
Before you begin
At least one running virtual machine (VM) or server with an active internet connection.
Login access to cp.icr.io/cp.
Root-level access to all nodes.
For the latest supported Red Hat Enterprise Linux versions, see Compatibility matrix.
About this task
Follow this procedure to set up a secure private registry by using authentication and a self-signed certificate. Complete these steps on a node that has both internet
access and access to the local cluster.
Note: Do not use an insecure registry for production.
Procedure
1. Register the hosts, and when prompted, enter the appropriate Red Hat customer portal credentials:
Example
[root@admin ~]# subscription-manager register
2. Pull the latest subscription:
Example
subscription-manager refresh
3. Identify the appropriate subscription and retrieve its Pool ID.
4. Attach the pool ID to gain access to the software entitlements. Use the Pool ID that you identified in the previous step.
Example
[root@admin ~]# subscription-manager attach --pool=POOL_ID
5. Disable the software repositories:
Example
[root@admin ~]# subscription-manager repos --disable=*
6. Enable the Red Hat Enterprise Linux baseos and appstream repositories.
Example
[root@admin ~]# subscription-manager repos --enable=rhel-9-for-x86_64-baseos-rpms
[root@admin ~]# subscription-manager repos --enable=rhel-9-for-x86_64-appstream-rpms
7. Update the system:
Example
[root@admin ~]# dnf update
8. Enable the ceph-tools repository.
Example
[root@admin ~]# curl https://public.dhe.ibm.com/ibmdl/export/pub/storage/ceph/ibm-storage-ceph-6-rhel-9.repo | sudo tee
/etc/yum.repos.d/ibm-storage-ceph-6-rhel-9.repo
9. Install the podman and httpd-tools packages.
[root@admin ~]# dnf install -y podman httpd-tools
10. Create folders for the private registry.
[root@admin ~]# mkdir -p /opt/registry/{auth,certs,data}
The registry is stored in /opt/registry and the directories are mounted in the container that is running the registry.
The auth directory stores the htpasswd file that the registry uses for authentication.
The certs directory stores the certificates that the registry uses for authentication.
The data directory stores the registry images.
11. Create credentials for accessing the private registry.
htpasswd -bBc /opt/registry/auth/htpasswd PRIVATE_REGISTRY_USERNAME PRIVATE_REGISTRY_PASSWORD
The b option provides the password from the command line.
The B option stores the password using Bcrypt encryption.
The c option creates the htpasswd file.
Replace PRIVATE_REGISTRY_USERNAME with the username to create for the private registry.
Replace PRIVATE_REGISTRY_PASSWORD with the password to create for the private registry username.
For example:
[root@admin ~]# htpasswd -bBc /opt/registry/auth/htpasswd myregistryusername myregistrypassword1
12. Create a self-signed certificate.
openssl req -newkey rsa:4096 -nodes -sha256 -keyout /opt/registry/certs/domain.key -x509 -days 365 -out
/opt/registry/certs/domain.crt -addext "subjectAltName = DNS:LOCAL_NODE_FQDN"
142 IBM Storage Ceph
Replace LOCAL_NODE_FQDN with the fully qualified hostname of the private registry node.
There is a prompt for the respective options for your certificate. The CN= value is the hostname of your node and should be resolvable by DNS or the /etc/hosts
file.
For example:
[root@admin ~] # openssl req -newkey rsa:4096 -nodes -sha256 -keyout /opt/registry/certs/domain.key -x509 -days 365 -out
/opt/registry/certs/domain.crt -addext "subjectAltName = DNS:admin.lab.ibm.com"
Note: When creating a self-signed certificate, be sure to create a certificate with a proper Subject Alternative Name (SAN). Podman commands that require TLS
verification for certificates that do not include a proper SAN, return the following error:
x509: certificate relies on legacy Common Name field, use SANs or temporarily enable Common Name matching with GODEBUG=x509igno
13. Create a symbolic link to domain.cert.
This allows skopeo to locate the certificate with the file extension .cert.
For example:
[root@admin ~]# ln -s /opt/registry/certs/domain.crt /opt/registry/certs/domain.cert
14. Add the certificate to the trusted list on the private registry node.
cp /opt/registry/certs/domain.crt /etc/pki/ca-trust/source/anchors/
update-ca-trust
trust list | grep -i "LOCAL_NODE_FQDN"
Replace LOCAL_NODE_FQDN with the FQDN of the private registry node.
For example:
[root@admin ~]# cp /opt/registry/certs/domain.crt /etc/pki/ca-trust/source/anchors/
[root@admin ~]# update-ca-trust
[root@admin ~]# trust list | grep -i "admin.lab.ibm.com"
label: admin.lab.ibm.com
15. Copy the certificate to any nodes that will access the private registry for installation and update the trusted list.
For example:
[root@admin ~]# scp /opt/registry/certs/domain.crt root@host01:/etc/pki/ca-trust/source/anchors/
[root@admin ~]# ssh root@host01
[root@host01 ~]# update-ca-trust
[root@host01 ~]# trust list | grep -i "admin.lab.ibm.com"
label: admin.lab.ibm.com
16. Download and install the mirror registry.
a. Download the mirror-registry from the Red Hat Hybrid Cloud Console.
b. Install the mirror registry.
./mirror-registry install --sslKey /opt/registry/certs/domain.key --sslCert /opt/registry/certs/domain.crt --initUser
myregistryuser --initPassword myregistrypass
17. On the local registry node, verify that cp.icr.io/cp is in the container registry search path.
a. Open /etc/containers/registries.conf for editing.
b. Optional: If needed, add cp.icr.io/cp to the unqualified-search-registries list.
unqualified-search-registries = ["cp.icr.io/cp"]
18. With your IBM Customer Portal credentials, login to cp.icr.io/cp.
podman login cp.icr.io/cp
19. You need to have access to registry.redhat.io to pull custom images. Log in to the Red Hat registry.
podman login registry.redhat.io
20. Copy the following IBM Storage Ceph, Prometheus, and Dashboard images from the IBM Customer Portal to the private registry.
Table 1. Custom image details
Ceph image
Component
Image details
cp.icr.io/cp/ibm-ceph/ceph-6-rhel9:latest
Prometheus
cp.icr.io/cp/ibm-ceph/prometheus:v4.12
Grafana
cp.icr.io/cp/ibm-ceph/ceph-6-dashboard-rhel9:latest
Node-exporter
cp.icr.io/cp/ibm-ceph/prometheus-node-exporter:v4.12
AlertManager
cp.icr.io/cp/ibm-ceph/prometheus-alertmanager:v4.12
HAProxy
cp.icr.io/cp/ibm-ceph/haproxy-rhel9:latest
Keepalived
cp.icr.io/cp/ibm-ceph/keepalived-rhel9:latest
SNMP Gateway
cp.icr.io/cp/ibm-ceph/snmp-notifier-rhel9:latest
podman run -v /CERTIFICATE_DIRECTORY_PATH:/certs:Z -v /CERTIFICATE_DIRECTORY_PATH/domain.cert:/certs/domain.cert:Z --rm
registry.redhat.io/rhel9/skopeo skopeo copy --remove-signatures --src-creds
IBM_CUSTOMER_PORTAL_LOGIN:IBM_CUSTOMER_PORTAL_PASSWORD --dest-cert-dir=./certs/ --dest-creds
PRIVATE_REGISTRY_USERNAME:PRIVATE_REGISTRY_PASSWORD docker://cp.icr.io/cp/SRC_IMAGE:SRC_TAG
docker://LOCAL_NODE_FQDN:8433/DST_IMAGE:pDST_TAG
Replace CERTIFICATE_DIRECTORY_PATH with the directory path to the self-signed certificates.
Replace CERTIFICATE_DIRECTORY_PATH and IBM_CUSTOMER_PORTAL_PASSWORD with your IBM Customer Portal credentials.
Replace PRIVATE_REGISTRY_USERNAME and PRIVATE_REGISTRY_PASSWORD with the private registry credentials.
IBM Storage Ceph 143
Replace SRC_IMAGE and SRC_TAG with the name and tag of the image to copy from cp.icr.io/cp.
Replace DST_IMAGE and DST_TAG with the name and tag of the image to copy to the private registry.
Replace LOCAL_NODE_FQDN with the FQDN of the private registry.
For example:
podman run -v /opt/registry/certs:/certs:Z -v /opt/registry/certs/domain.cert:/certs/domain.cert:Z --rm --rm
registry.redhat.io/rhel9/skopeo skopeo copy --remove-signatures --src-creds myusername:mypassword1 --dest-certdir=./certs/ --dest-creds myregistryusername:myregistrypassword1 docker://cp.icr.io/cp/ibm-ceph/ceph-6-rhel9:latest
docker://admin.lab.ibm.com:8433/ibm-ceph/ceph-6-rhel9:latest
podman run -v /opt/registry/certs:/certs:Z -v /opt/registry/certs/domain.cert:/certs/domain.cert:Z --rm
registry.redhat.io/rhel9/skopeo skopeo copy --remove-signatures --src-creds myusername:mypassword1 --dest-certdir=./certs/ --dest-creds myregistryusername:myregistrypassword1 docker://cp.icr.io/cp/ibm-ceph/prometheus-nodeexporter:v4.12 docker://admin.lab.ibm.com:8433/ibm-ceph/prometheus-node-exporter:v4.12 docker
podman run -v /opt/registry/certs:/certs:Z -v /opt/registry/certs/domain.cert:/certs/domain.cert:Z --rm
registry.redhat.io/rhel9/skopeo skopeo copy --remove-signatures --src-creds myusername:mypassword1 --dest-certdir=./certs/ --dest-creds myregistryusername:myregistrypassword1 docker:cp.icr.io/cp/ibm-ceph/ceph-6-dashboardrhel9:latest docker://admin.lab.ibm.com:8433/ibm-ceph/ceph-6-dashboard-rhel9:latest
podman run -v /opt/registry/certs:/certs:Z -v /opt/registry/certs/domain.cert:/certs/domain.cert:Z --rm
registry.redhat.io/rhel9/skopeo skopeo copy --remove-signatures --src-creds myusername:mypassword1 --dest-certdir=./certs/ --dest-creds myregistryusername:myregistrypassword1 docker://cp.icr.io/cp/ibm-ceph/prometheus:v4.12
docker://admin.lab.ibm.com:8433/ibm-ceph/prometheus:v4.12
podman run -v /opt/registry/certs:/certs:Z -v /opt/registry/certs/domain.cert:/certs/domain.cert:Z --rm
registry.redhat.io/rhel9/skopeo skopeo copy --remove-signatures --src-creds myusername:mypassword1 --dest-certdir=./certs/ --dest-creds myregistryusername:myregistrypassword1 docker://cp.icr.io/cp/ibm-ceph/prometheusalertmanager:v4.12 docker://admin.lab.ibm.com:8433/ibm-ceph/prometheus-alertmanager:v4.12
podman run -v /opt/registry/certs:/certs:Z -v /opt/registry/certs/domain.cert:/certs/domain.cert:Z --rm
registry.redhat.io/rhel9/skopeo skopeo copy --remove-signatures --src-creds myusername:mypassword1 --dest-certdir=./certs/ --dest-creds myregistryusername:myregistrypassword1 docker://cp.icr.io/cp/ibm-ceph/haproxy-rhel9:latest
docker://admin.lab.ibm.com:8433/ibm-ceph/haproxy-rhel9:latest
podman run -v /opt/registry/certs:/certs:Z -v /opt/registry/certs/domain.cert:/certs/domain.cert:Z --rm
registry.redhat.io/rhel9/skopeo skopeo copy --remove-signatures --src-creds myusername:mypassword1 --dest-certdir=./certs/ --dest-creds myregistryusername:myregistrypassword1 docker://cp.icr.io/cp/ibm-ceph/keepalived-rhel9:latest
docker://admin.lab.ibm.com:8433/ibm-ceph/keepalived-rhel9:latest
Note: For more image Ceph package versions, see What are the Red Hat Ceph Storage releases and corresponding Ceph package versions? within the Red Hat
Customer Portal.
21. Using the Ceph Dashboard, verify that the images are in the local registry.
For more information, see Monitoring service of the Ceph cluster.
Running the preflight playbook for a disconnected installation
You use the cephadm-preflight.yml Ansible playbook to configure the Ceph repository and prepare the storage cluster for bootstrapping. It also installs some
prerequisites, such as podman, lvm2, chrony, and cephadm.
Important: If you are using Red Hat Enterprise Linux 9, do not use these steps, as cephadm-preflight is not supported. Continue to Performing a disconnected
installation.
The preflight playbook uses the cephadm-ansible inventory hosts file to identify all the nodes in the storage cluster. The default location for cephadm-ansible,
cephadm-preflight.yml, and the inventory hosts file is /usr/share/cephadm-ansible/.
The following example shows the structure of a typical inventory file:
Example
host02
host03
host04
[admin]
host01
The [admin] group in the inventory file contains the name of the node where the admin keyring is stored.
Note: Run the preflight playbook before you bootstrap the initial host.
Prerequisites
Nodes configured to access a local YUM repository server with the following repositories enabled on respective Red Hat Enterprise Linux versions.
rhel-9-for-x86_64-baseos-rpms
rhel-9-for-x86_64-appstream-rpms
curl https://public.dhe.ibm.com/ibmdl/export/pub/storage/ceph/ibm-storage-ceph-6-rhel-9.repo | sudo tee /etc/yum.repos.d/ibm-storage-ceph-6rhel-9.repo
NOTE: For more information about setting up a local YUM repository, see the Red Hat knowledge base article Creating a Local Repository and Sharing with
Disconnected/Offline/Air-gapped Systems
The cephadm-ansible package is installed on the Ansible administration node.
144 IBM Storage Ceph
[root@admin ~]# dnf install cephadm-ansible
Root-level access to all nodes in the storage cluster.
Passwordless ssh is set up on all hosts in the storage cluster.
Procedure
1. Navigate to the /usr/share/cephadm-ansible directory on the Ansible administration node.
2. Open and edit the hosts file and add your nodes.
3. Add license to install IBM Storage Ceph and click Accept on all nodes:
Example
[root@admin ~]# dnf install ibm-storage-ceph-license
a. Accept these provisions:
Example
[root@admin ~]# sudo touch /usr/share/ibm-storage-ceph-license/accept
4. Run the preflight playbook with the ceph_origin parameter set to custom to use a local YUM repository:
Syntax
ansible-playbook -i INVENTORY_FILE cephadm-preflight.yml --extra-vars "ceph_origin=custom" -e
"custom_repo_url=CUSTOM_REPO_URL"
Example
[ansible@admin cephadm-ansible]$ ansible-playbook -i hosts cephadm-preflight.yml --extra-vars "ceph_origin=custom" -e
"custom_repo_url=http://mycustomrepo.lab.ibm.com/x86_64/os/"
After installation is complete, cephadm resides in the /usr/sbin/ directory.
5. Alternatively, you can use the --limit option to run the preflight playbook on a selected set of hosts in the storage cluster:
Syntax
ansible-playbook -i INVENTORY_FILE cephadm-preflight.yml --extra-vars "ceph_origin=custom" -e
"custom_repo_url=CUSTOM_REPO_URL" --limit GROUP_NAME|NODE_NAME
Replace GROUP_NAME with a group name from your inventory file. Replace NODE_NAME with a specific node name from your inventory file.
Example
[ansible@admin cephadm-ansible]$ ansible-playbook -i hosts cephadm-preflight.yml --extra-vars "ceph_origin=custom" -e
"custom_repo_url=http://mycustomrepo.lab.ibm.com/x86_64/os/" --limit clients
[ansible@admin cephadm-ansible]$ ansible-playbook -i hosts cephadm-preflight.yml --extra-vars "ceph_origin=custom" -e
"custom_repo_url=http://mycustomrepo.lab.ibm.com/x86_64/os/" --limit host02
NOTE: When you run the preflight playbook, cephadm-ansible automatically installs chrony and ceph-common on the client nodes.
Performing a disconnected installation
Before you can perform the installation, you must obtain an IBM Storage Ceph container image, either from a proxy host that has access to the IBM registry or by copying
the image to your local registry.
Important: Before you begin the bootstrapping process, make sure that the container image that you want to use has the same version of IBM Storage Ceph as cephadm.
If the two versions do not match, bootstrapping fails at the Creating initial admin
user stage.
Note: If your local registry uses a self-signed certificate with a local registry, ensure you have added the trusted root certificate to the bootstrap host. For more
information, see Configuring a private registry for a disconnected installation.
Prerequisites
At least one running virtual machine (VM) or server.
Root-level access to all nodes.
Passwordless ssh is set up on all hosts in the storage cluster.
The preflight playbook has been run on the bootstrap host in the storage cluster. For more information, see Running the preflight playbook for a disconnected
installation.
A private registry has been configured and the bootstrap node has access to it. For more information, see Configuring a private registry for a disconnected
installation
An IBM Storage Ceph container image resides in the custom registry.
Procedure
IBM Storage Ceph 145
1. Log in to the bootstrap host.
2. Bootstrap the storage cluster:
Syntax
cephadm --image PRIVATE_REGISTRY_NODE_FQDN:5000/CUSTOM_IMAGE_NAME:IMAGE_TAG bootstrap --mon-ip IP_ADDRESS --registry-url
PRIVATE_REGISTRY_NODE_FQDN:5000 --registry-username PRIVATE_REGISTRY_USERNAME --registry-password
PRIVATE_REGISTRY_PASSWORD
Replace PRIVATE_REGISTRY_NODE_FQDN with the fully qualified domain name of your private registry.
Replace CUSTOM_IMAGE_NAME and IMAGE_TAG with the name and tag of the IBM Storage Ceph container image that resides in the private registry.
Replace IP_ADDRESS with the IP address of the node you are using to run cephadm
bootstrap.
Replace PRIVATE_REGISTRY_USERNAME with the username to create for the private registry.
Replace PRIVATE_REGISTRY_PASSWORD with the password to create for the private registry username.
Example
[root@host01 ~]# cephadm --image admin.lab.ibm.com:5000/ibm-ceph/ceph-6-rhel9:latest bootstrap --mon-ip 10.10.128.68
--registry-url admin.lab.ibm.com:5000 --registry-username myregistryusername --registry-password myregistrypassword1
The script takes a few minutes to complete. Once the script completes, it provides the credentials to the IBM Storage Ceph Dashboard URL, a command to
access the Ceph command-line interface (CLI), and a request to enable telemetry.
Ceph Dashboard is now available at:
URL: https://host01:8443/
User: admin
Password: i8nhu7zham
Enabling client.admin keyring and conf on hosts with "admin" label
You can access the Ceph CLI with:
sudo /usr/sbin/cephadm shell --fsid 266ee7a8-2a05-11eb-b846-5254002d4916 -c /etc/ceph/ceph.conf -k
/etc/ceph/ceph.client.admin.keyring
Please consider enabling telemetry to help improve Ceph:
ceph telemetry on
For more information see:
https://docs.ceph.com/docs/master/mgr/telemetry/
Bootstrap complete.
After the bootstrap process is complete, configure the container images, as detailed in Changing configurations of custom container images for disconnected installations.
Once your storage cluster is up and running, configure additional daemons and services. For more information, see Operations
Changing configurations of custom container images for disconnected installations
After you perform the initial bootstrap for disconnected nodes, you must specify custom container images for monitoring stack daemons. You can override the default
container images for monitoring stack daemons, since the nodes do not have access to the default container registry.
Note: Make sure that the bootstrap process on the initial host is complete before making any configuration changes.
By default, the monitoring stack components are deployed based on the primary Ceph image. For disconnected environment of the storage cluster, you can use the latest
available monitoring stack component images.
Note: When using a custom registry, be sure to log in to the custom registry on newly added nodes before adding any Ceph daemons.
Syntax
ceph cephadm registry-login --registry-url CUSTOM_REGISTRY_NAME --registry_username REGISTRY_USERNAME --registry_password
REGISTRY_PASSWORD
Example
# ceph cephadm registry-login --registry-url myregistry --registry_username myregistryusername --registry_password
myregistrypassword1
For more information, see Performing a disconnected installation.
Prerequisites
At least one running virtual machine (VM) or server.
Root-level access to all nodes.
Passwordless ssh is set up on all hosts in the storage cluster.
For the latest supported Red Hat Enterprise Linux versions, see Compatibility matrix.
146 IBM Storage Ceph
Procedure
1. Set the custom container images with the ceph config command:
Syntax
ceph config set mgr mgr/cephadm/OPTION_NAME CUSTOM_REGISTRY_NAME/IMAGE_NAME
Use the following options for OPTION_NAME:
container_image_prometheus
container_image_grafana
container_image_alertmanager
container_image_node_exporter
Example
[root@host01 ~]# ceph config set mgr mgr/cephadm/container_image_prometheus private_registry/prometheus
[root@host01 ~]# ceph config set mgr mgr/cephadm/container_image_grafana private_registry/grafana
[root@host01 ~]# ceph config set mgr mgr/cephadm/container_image_alertmanager private_registry/alertmanager
[root@host01 ~]# ceph config set mgr mgr/cephadm/container_image_node_exporter private_registry/node_exporter
2. Redeploy node-exporter:
Syntax
ceph orch redeploy node-exporter
Note:
If any of the services do not deploy, you can redeploy them with the ceph orch redeploy command.
By setting a custom image, the default values for the configuration image name and tag will be overridden, but not overwritten. The default values change when
updates become available. By setting a custom image, you will not be able to configure the component for which you have set the custom image for automatic
updates. You will need to manually update the configuration image name and tag to be able to install updates.
If you choose to revert to using the default configuration, you can reset the custom container image. Use ceph config rm to reset the configuration option:
Syntax
ceph config rm mgr mgr/cephadm/OPTION_NAME
Example
ceph config rm mgr mgr/cephadm/container_image_prometheus
Reference
Performing a disconnected installation
Adding hosts in disconnected deployments
If you are running a storage cluster on a private network and your host domain name server (DNS) cannot be reached through private IP, you must include both the host
name and the IP address for each host you want to add to the storage cluster.
Prerequisites
A running IBM Storage Ceph cluster.
Root-level access to all hosts in the storage cluster.
Procedure
1. Log into the cephadm shell.
Syntax
[root@host01 ~]# cephadm shell
2. Add the host:
Syntax
ceph orch host add HOST_NAME HOST_ADDRESS
Example
[ceph: root@host01 /]# ceph orch host add host03 10.10.128.70
Launching the cephadm shell
IBM Storage Ceph 147
The cephadm shell command launches a bash shell in a container with all of the Ceph packages installed. This enables you to perform “Day One” cluster setup tasks,
such as installation and bootstrapping, and to invoke ceph commands.
There are two ways to invoke the cephadm shell:
Enter cephadm shell at the system prompt:
Example
[root@host01 ~]# cephadm shell
[ceph: root@host01 /]# ceph -s
At the system prompt, type cephadm shell and the command you want to execute:
Example
[root@host01 ~]# cephadm shell ceph -s
Prerequisites
A running IBM Storage Ceph cluster.
Root-level access to all nodes in the storage cluster.
Procedure
There are two ways to launch the cephadm shell:
Enter cephadm shell at the system prompt. This example invokes the ceph
-s command from within the shell.
Example
[root@host01 ~]# cephadm shell
[ceph: root@host01 /]# ceph -s
At the system prompt, type cephadm shell and the command you want to execute:
Example
[root@host01 ~]# cephadm shell ceph -s
cluster:
id:
f64f341c-655d-11eb-8778-fa163e914bcc
health: HEALTH_OK
services:
mon: 3 daemons, quorum host01,host02,host03 (age 94m)
mgr: host01.lbnhug(active, since 59m), standbys: host02.rofgay, host03.ohipra
mds: 1/1 daemons up, 1 standby
osd: 18 osds: 18 up (since 10m), 18 in (since 10m)
rgw: 4 daemons active (2 hosts, 1 zones)
data:
volumes: 1/1 healthy
pools:
8 pools, 225 pgs
objects: 230 objects, 9.9 KiB
usage:
271 MiB used, 269 GiB / 270 GiB avail
pgs:
225 active+clean
io:
client:
85 B/s rd, 0 op/s rd, 0 op/s wr
NOTE: If the node contains configuration and keyring files in /etc/ceph/, the container environment uses the values in those files as defaults for the cephadm shell. If
you execute the cephadm shell on a MON node, the cephadm shell inherits its default configuration from the MON container, instead of using the default configuration.
cephadm commands
The cephadm is a command line tool to manage the local host for the Cephadm Orchestrator. It provides commands to investigate and modify the state of the current
host.
Some of the commands are generally used for debugging.
Note: cephadm is not required on all hosts, however, it is useful when investigating a particular daemon. The cephadm-ansible-preflight playbook installs cephadm
on all hosts and the cephadm-ansible purge playbook requires cephadm be installed on all hosts to work properly.
Table 1. The cephadm commands
Command
adopt
Description
Convert an upgraded storage cluster daemon to
run cephadm.
148 IBM Storage Ceph
Syntax
cephadm adopt [-h] --name DAEMON_NAME --style STYLE
[--cluster CLUSTER] --legacy-dir [LEGACY_DIR] -config-json CONFIG_JSON] [--skip-firewalld] [--skippull]
Example
[root@host01 ~]#
cephadm adopt -style=legacy --name
prometheus.host02
Command
ceph-volume
check-host
Description
Syntax
This command is used to list all the devices on the cephadm ceph-volume inventory/simple/raw/lvm [-h] [-particular host. Run the ceph-volume command fsid FSID] [--config-json CONFIG_JSON] [--config
CONFIG, -c CONFIG] [--keyring KEYRING, -k KEYRING]
inside a container Deploys OSDs with different
device technologies like lvm or physical disks
using pluggable tools and follows a predictable,
and robust way of preparing, activating, and
starting OSDs.
Check the host configuration that is suitable for a cephadm check-host [--expect-hostname HOSTNAME]
Ceph cluster.
deploy
Deploys a daemon on the local host.
enter
Run an interactive shell inside a running daemon
container.
help
View all the commands supported by cephadm.
install
Install the packages.
inspectimage
Inspect the local Ceph container image.
listnetworks
List the IP networks.
ls
List daemon instances known to cephadm on the
hosts. You can use --no-detail for the
command to run faster, which gives details of the
daemon name, fsid, style, and systemd unit per
daemon. You can use --legacy-dir option to
specify a legacy base directory to search for
daemons.
logs
Print journald logs for a daemon container. This cephadm logs [--fsid FSID] --name DAEMON_NAME
cephadm logs [--fsid FSID] --name DAEMON_NAME -- -n
is similar to the journalctl command.
NUMBER # Last N lines
cephadm logs [--fsid FSID] --name DAEMON_NAME -- -f #
Follow the logs
[root@host01 ~]#
cephadm check-host
--expect-hostname
host02
[root@host01 ~]#
cephadm shell
deploy mon --fsid
f64f341c-655d-11eb8778-fa163e914bcc
cephadm shell deploy DAEMON_TYPE [-h] [--name
DAEMON_NAME] [--fsid FSID] [--config CONFIG, -c
CONFIG] [--config-json CONFIG_JSON] [--keyring
KEYRING] [--key KEY] [--osd-fsid OSD_FSID] [--skipfirewalld] [--tcp-ports TCP_PORTS] [--reconfig] [-allow-ptrace] [--memory-request MEMORY_REQUEST] [-memory-limit MEMORY_LIMIT] [--meta-json META_JSON]
cephadm enter [-h] [--fsid FSID] --name NAME [command [root@host01 ~]#
[command …]]
cephadm enter -name 52c611f2b1d9
cephadm help
[root@host01 ~]#
cephadm help
cephadm install PACKAGES
[root@host01 ~]#
cephadm install
ceph-common cephosd
cephadm --image IMAGE_ID inspect-image
[root@host01 ~]#
cephadm --image
13ea90216d0be03003d
12d7869f72ad9de5cec
9e54a27fd308e01e467
c0d4a0a inspectimage
cephadm list-networks
[root@host01 ~]#
cephadm listnetworks
cephadm ls [--no-detail] [--legacy-dir LEGACY_DIR]
[root@host01 ~]#
cephadm ls --nodetail
prepare-host Prepare a host for cephadm.
cephadm prepare-host [--expect-hostname HOSTNAME]
pull
cephadm [-h] [--image IMAGE_ID] pull
Pull the Ceph image.
Example
[root@host01 ~]#
cephadm ceph-volume
inventory --fsid
f64f341c-655d-11eb8778-fa163e914bcc
[root@host01 ~]#
cephadm logs --fsid
57bddb48-ee04-11eb9962-001a4a000672 -name osd.8
[root@host01 ~]#
cephadm logs --fsid
57bddb48-ee04-11eb9962-001a4a000672 -name osd.8 -- -n
20
[root@host01 ~]#
cephadm logs --fsid
57bddb48-ee04-11eb9962-001a4a000672 -name osd.8 -- -f
[root@host01 ~]#
cephadm preparehost
[root@host01 ~]#
cephadm preparehost --expecthostname host01
[root@host01 ~]#
cephadm --image
13ea90216d0be03003d
12d7869f72ad9de5cec
9e54a27fd308e01e467
c0d4a0a pull
IBM Storage Ceph 149
Command
registrylogin
Description
Give cephadm login information for an
authenticated registry. Cephadm attempts to log
the calling host into that registry.
Syntax
Example
cephadm registry-login --registry-url REGISTRY_URL -- [root@host01 ~]#
registry-username USERNAME --registry-password
cephadm registryPASSWORD [--fsid FSID] [--registry-json JSON_FILE]
login --registryurl cp.icr.io/cp -You can also use a JSON registry file containing the login info formatted registry-username
myuser1 --registryas:
password
mypassword1
cat REGISTRY_FILE
{
"url":"REGISTRY_URL",
"username":"REGISTRY_USERNAME",
"password":"REGISTRY_PASSWORD"
}
[root@host01 ~]#
cat registry_file
{
"url":"cp.icr.io/cp
",
"username":"myuser"
,
"password":"mypass"
}
rm-daemon
rm-cluster
cephadm rm-daemon --fsid FSID --name DAEMON_NAME [-Remove a specific daemon instance. If you run
the cephadm rm-daemon command on the host force ] [--force-delete-data]
directly, although the command removes the
daemon, the cephadm mgr module notices that
the daemon is missing and redeploys it. This
command is problematic and should be used only
for experimental purposes and debugging.
Remove all the daemons from a storage cluster on cephadm rm-cluster --fsid FSID [--force]
that specific host where it is run. Similar to rmdaemon, if you remove a few daemons this way
and the Ceph Orchestrator is not paused and
some of those daemons belong to services that
are not unmanaged, the cephadm orchestrator
just redeploys them there.
[root@host01 ~]#
cephadm registrylogin -i
registry_file
[root@host01 ~]#
cephadm rm-daemon -fsid f64f341c655d-11eb-8778fa163e914bcc --name
osd.8
[root@host01 ~]#
cephadm rm-cluster
--fsid f64f341c655d-11eb-8778fa163e914bcc
rm-repo
Remove a package repository configuration. This
is mainly used for the disconnected installation of
IBM Storage Ceph.
cephadm rm-repo [-h]
[root@host01 ~]#
cephadm rm-repo
run
Run a Ceph daemon, in a container, in the
foreground.
cephadm run [--fsid FSID] --name DAEMON_NAME
shell
cephadm shell [--fsid FSID] [--name DAEMON_NAME, -n
Run an interactive shell with access to Ceph
DAEMON_NAME] [--config CONFIG, -c CONFIG] [--mount
commands over the inferred or specified Ceph
MOUNT, -m MOUNT] [--keyring KEYRING, -k KEYRING] [-cluster. You can enter the shell using the cephadm env ENV, -e ENV]
shell command and run all the orchestrator
commands within the shell.
cephadm unit [--fsid FSID] --name DAEMON_NAME
Start, stop, restart, enable, and disable the
daemons with this operation. This operates on the start/stop/restart/enable/disable
[root@host01 ~]#
cephadm run --fsid
f64f341c-655d-11eb8778-fa163e914bcc -name osd.8
[root@host01 ~]#
cephadm shell -ceph orch ls
[root@host01 ~]#
cephadm shell
unit
daemon’s systemd unit.
version
Provides the version of the storage cluster.
cephadm version
Verifying the cluster installation
Once the cluster installation is complete, you can verify that the IBM Storage Ceph installation is running properly.
There are two ways of verifying the storage cluster installation as a root user:
Run the podman ps command.
Run the cephadm shell ceph -s.
Prerequisites
Root-level access to all nodes in the storage cluster.
Procedure
Run the podman ps command:
Example
150 IBM Storage Ceph
[root@host01 ~]#
cephadm unit --fsid
f64f341c-655d-11eb8778-fa163e914bcc -name osd.8 start
[root@host01 ~]#
cephadm version
[root@host01 ~]# podman ps
NOTE: In the NAMES column, the unit files now include the FSID.
Run the cephadm shell ceph -s command:
Example
[root@host01 ~]# cephadm shell ceph -s
cluster:
id:
f64f341c-655d-11eb-8778-fa163e914bcc
health: HEALTH_OK
services:
mon: 3 daemons, quorum host01,host02,host03 (age 94m)
mgr: host01.lbnhug(active, since 59m), standbys: host02.rofgay, host03.ohipra
mds: 1/1 daemons up, 1 standby
osd: 18 osds: 18 up (since 10m), 18 in (since 10m)
rgw: 4 daemons active (2 hosts, 1 zones)
data:
volumes: 1/1 healthy
pools:
8 pools, 225 pgs
objects: 230 objects, 9.9 KiB
usage:
271 MiB used, 269 GiB / 270 GiB avail
pgs:
225 active+clean
io:
client:
85 B/s rd, 0 op/s rd, 0 op/s wr
NOTE: The health of the storage cluster is in HEALTH_WARN status as the hosts and the daemons are not added.
Adding hosts
Bootstrapping the IBM Storage Ceph installation creates a working storage cluster, consisting of one Monitor daemon and one Manager daemon within the same container.
As a storage administrator, you can add additional hosts to the storage cluster and configure them.
Note:
For Red Hat Enterprise Linux 8, running the preflight playbook installs podman, lvm2, chrony, and cephadm on all hosts listed in the Ansible inventory file.
For Red Hat Enterprise Linux 9, you need to manually install podman, lvm2, chrony, and cephadm on all hosts and skip steps for running ansible playbooks as the
preflight playbook is not supported.
When using a custom registry, be sure to log in to the custom registry on newly added nodes before adding any Ceph daemons.
Syntax
ceph cephadm registry-login --registry-url CUSTOM_REGISTRY_NAME --registry_username REGISTRY_USERNAME --registry_password
REGISTRY_PASSWORD
Example
# ceph cephadm registry-login --registry-url myregistry --registry_username myregistryusername --registry_password
myregistrypassword1
Using the addr option to identify hosts
Adding multiple hosts
Prerequisites
A running IBM Storage Ceph cluster.
Root-level or user with sudo access to all nodes in the storage cluster.
Register the nodes to IBM subscription.
Ansible user with sudo and passwordless ssh access to all nodes in the storage cluster.
Procedure
Note: In the following procedure, use either root, as indicated, or the username with which the user is bootstrapped.
1. From the node that contains the admin keyring, install the storage cluster’s public SSH key in the root user’s authorized_keys file on the new host:
Syntax
ssh-copy-id -f -i /etc/ceph/ceph.pub user@NEWHOST
Example
[root@host01 ~]# ssh-copy-id -f -i /etc/ceph/ceph.pub root@host02
[root@host01 ~]# ssh-copy-id -f -i /etc/ceph/ceph.pub root@host03
2. Navigate to the /usr/share/cephadm-ansible directory on the Ansible administration node.
Example
IBM Storage Ceph 151
[ansible@admin ~]$ cd /usr/share/cephadm-ansible
3. From the Ansible administration node, add the new host to the Ansible inventory file. The default location for the file is /usr/share/cephadm-ansible/hosts.
The following example shows the structure of a typical inventory file:
Note: If you have previously added the new host to the Ansible inventory file and run the preflight playbook on the host, skip to step 4.
Example
[ansible@admin ~]$ cat hosts
host02
host03
host04
[admin]
host01
4. Run the preflight playbook with the --limit option:
Syntax
ansible-playbook -i INVENTORY_FILE cephadm-preflight.yml --extra-vars "ceph_origin=ibm" --limit NEWHOST
Example
[ansible@admin cephadm-ansible]$ ansible-playbook -i hosts cephadm-preflight.yml --extra-vars "ceph_origin=ibm" --limit
host02
The preflight playbook installs podman, lvm2, chrony, and cephadm on the new host. After installation is complete, cephadm resides in the /usr/sbin/
directory.
For Red Hat Enterprise Linux 9, install podman, lvm2, chrony, and cephadm manually:
Example
[root@host01 ~]# dnf install podman lvm2 chrony cephadm
5. From the bootstrap node, use the cephadm orchestrator to add the new host to the storage cluster:
Syntax
ceph orch host add NEWHOST
Example
[ceph: root@host01 /]# ceph orch host add host02
Added host 'host02' with addr '10.10.128.69'
[ceph: root@host01 /]# ceph orch host add host03
Added host 'host03' with addr '10.10.128.70'
6. Optional: You can also add nodes by IP address, before and after you run the preflight playbook. If you do not have DNS configured in your storage cluster
environment, you can add the hosts by IP address, along with the host names.
Syntax
ceph orch host add HOSTNAME IP_ADDRESS
Example
[ceph: root@host01 /]# ceph orch host add host02 10.10.128.69
Added host 'host02' with addr '10.10.128.69'
View the status of the storage cluster and verify that the new host has been added. The STATUS of the hosts is blank, in the output of the ceph orch host
ls command.
Example
[ceph: root@host01 /]# ceph orch host ls
Reference
Registering the IBM Storage Ceph nodes
Creating an Ansible user with sudo access
Using the addr option to identify hosts
The addr option offers an additional way to contact a host. Add the IP address of the host to the addr option. If ssh cannot connect to the host by its hostname, then it
uses the value stored in addr to reach the host by its IP address.
Prerequisites
A storage cluster that has been installed and bootstrapped.
Root-level access to all nodes in the storage cluster.
Procedure
152 IBM Storage Ceph
1. Log in to the cephadm shell:
Example
[root@host01 ~]# cephadm shell
2. Add the IP address:
Syntax
ceph orch host add HOSTNAME IP_ADDR
Example
[ceph: root@host01 /]# ceph orch host add host02 10.10.128.68
NOTE: If adding a host by hostname results in that host being added with an IPv6 address instead of an IPv4 address, use ceph orch host to specify the IP
address of that host:
Syntax
ceph orch host set-addr HOSTNAME IP_ADDR
To convert the IP address from IPv6 format to IPv4 format for a host you have added, use the following command:
ceph orch host set-addr HOSTNAME IPV4_ADDRESS
Adding multiple hosts
Use a YAML file to add multiple hosts to the storage cluster at the same time.
Note: Be sure to create the hosts.yaml file within a host container, or create the file on the local host and then use the cephadm shell to mount the file within the container.
The cephadm shell automatically places mounted files in /mnt. If you create the file directly on the local host and then apply the hosts.yaml file instead of mounting it, you
might see a File does not exist error.
Prerequisites
A storage cluster that has been installed and bootstrapped.
Root-level access to all nodes in the storage cluster.
Procedure
1. Copy over the public ssh key to each of the hosts that you want to add.
2. Use a text editor to create a hosts.yaml file.
3. Add the host descriptions to the hosts.yaml file, as shown in the following example. Include the labels to identify placements for the daemons that you want to
deploy on each host. Separate each host description with three dashes (---).
Example
service_type: host
addr:
hostname: host02
labels:
- mon
- osd
- mgr
--service_type: host
addr:
hostname: host03
labels:
- mon
- osd
- mgr
--service_type: host
addr:
hostname: host04
labels:
- mon
- osd
4. If you created the hosts.yaml file directly on the local host, use the cephadm shell to mount the file:
Example
[root@host01 ~]# cephadm shell --mount hosts.yaml -- ceph orch apply -i /mnt/hosts.yaml
5. If you created the hosts.yaml file within the host container, invoke the ceph orch apply command:
Example
[ceph: root@host01 /]# ceph orch apply -i hosts.yaml
Added host 'host02' with addr '10.10.128.69'
IBM Storage Ceph 153
Added host 'host03' with addr '10.10.128.70'
Added host 'host04' with addr '10.10.128.71'
6. View the list of hosts and their labels:
Example
[ceph: root@host01 /]# ceph orch host ls
HOST
ADDR
LABELS
STATUS
host02
host02
mon osd mgr
host03
host03
mon osd mgr
host04
host04
mon osd
Note: If a host is online and operating normally, its status is blank. An offline host shows a status of OFFLINE, and a host in maintenance mode shows a status of
MAINTENANCE.
Removing hosts
You can remove hosts of a Ceph cluster with the Ceph Orchestrators. All the daemons are removed with the drain option which adds the _no_schedule label to ensure
that you cannot deploy any daemons or a cluster till the operation is complete.
Important: If you are removing the bootstrap host, be sure to copy the admin keyring and the configuration file to another host in the storage cluster before you remove
the host.
Prerequisites
A running IBM Storage Ceph cluster.
Root-level access to all the nodes.
Hosts are added to the storage cluster.
All the services are deployed.
Cephadm is deployed on the nodes where the services have to be removed.
Procedure
1. Log into the Cephadm shell:
Example
[root@host01 ~]# cephadm shell
2. Fetch the host details:
Example
[ceph: root@host01 /]# ceph orch host ls
3. Drain all the daemons from the host:
Syntax
ceph orch host drain HOSTNAME
Example
[ceph: root@host01 /]# ceph orch host drain host02
The _no_schedule label is automatically applied to the host which blocks deployment.
4. Check the status of OSD removal:
Example
[ceph: root@host01 /]# ceph orch osd rm status
When no placement groups (PG) are left on the OSD, the OSD is decommissioned and removed from the storage cluster.
5. Check if all the daemons are removed from the storage cluster:
Syntax
ceph orch ps HOSTNAME
Example
[ceph: root@host01 /]# ceph orch ps host02
6. Remove the host:
Syntax
ceph orch host rm HOSTNAME
Example
154 IBM Storage Ceph
[ceph: root@host01 /]# ceph orch host rm host02
Reference
Adding hosts using the Ceph Orchestrator
Listing hosts using the Ceph Orchestrator
Labeling hosts
The Ceph orchestrator supports assigning labels to hosts. Labels are free-form and have no specific meanings. This means that you can use mon, monitor,
mycluster_monitor, or any other text string. Each host can have multiple labels.
For example, apply the mon label to all hosts on which you want to deploy Ceph Monitor daemons, mgr for all hosts on which you want to deploy Ceph Manager daemons,
rgw for Ceph Object Gateway daemons, and so on.
Labeling all the hosts in the storage cluster helps to simplify system management tasks by allowing you to quickly identify the daemons running on each host. In addition,
you can use the Ceph orchestrator or a YAML file to deploy or remove daemons on hosts that have specific host labels.
Adding a label to a host
Use the Ceph Orchestrator to add a label to a host. Labels can be used to specify placement of daemons.
Removing a label from a host
Use the Ceph Orchestrator to remove a label from a host.
Using host labels to deploy daemons on specific hosts
Prerequisites
A storage cluster that has been installed and bootstrapped.
Adding a label to a host
Use the Ceph Orchestrator to add a label to a host. Labels can be used to specify placement of daemons.
A few examples of labels are mgr, mon, and osd based on the service deployed on the hosts. Each host can have multiple labels.
You can also add the following host labels that have special meaning to cephadm and they begin with _:
_no_schedule: This label prevents cephadm from scheduling or deploying daemons on the host. If it is added to an existing host that already contains Ceph
daemons, it causes cephadm to move those daemons elsewhere, except OSDs which are not removed automatically. When a host is added with the _no_schedule
label, no daemons are deployed on it. When the daemons are drained before the host is removed, the _no_schedule label is set on that host.
_no_autotune_memory: This label does not autotune memory on the host. It prevents the daemon memory from being tuned even when the
osd_memory_target_autotune option or other similar options are enabled for one or more daemons on that host.
_admin: By default, the _admin label is applied to the bootstrapped host in the storage cluster and the client.admin key is set to be distributed to that host with
the ceph orch client-keyring {ls|set|rm} function. Adding this label to additional hosts normally causes cephadm to deploy configuration and keyring
files in the /etc/ceph directory.
Prerequisites
A storage cluster that has been installed and bootstrapped.
Root-level access to all nodes in the storage cluster.
Hosts are added to the storage cluster.
Procedure
1. Launch the cephadm shell:
[root@host01 ~]# cephadm shell
[ceph: root@host01 /]#
2. Add a label to a host:
Syntax
ceph orch host label add HOSTNAME LABEL
Example
[ceph: root@host01 /]# ceph orch host label add host02 mon
Verification
List the hosts:
IBM Storage Ceph 155
Example
[ceph: root@host01 /]# ceph orch host ls
Removing a label from a host
Use the Ceph Orchestrator to remove a label from a host.
Before you begin
A IBM Storage Ceph cluster that has been installed and boostrapped.
Root-level access to all nodes in the storage cluster.
Hosts are added to the storage cluster.
Procedure
1. Log into the Cephadm shell.
[root@host01 ~]# cephadm shell
2. Remove the label.
ceph orch host label rm HOST_NAME LABEL_NAME
For example,
[ceph: root@host01 /]# ceph orch host label rm host02 mon
What to do next
Verify that the label has been moved from the host, by using the ceph orch host ls command.
Using host labels to deploy daemons on specific hosts
You can use host labels to deploy daemons to specific hosts. There are two ways to use host labels to deploy daemons on specific hosts:
By using the --placement option from the command line.
By using a YAML file.
Prerequisites
A storage cluster that has been installed and bootstrapped.
Root-level access to all nodes in the storage cluster.
Procedure
1. Log into the Cephadm shell:
Example
[root@host01 ~]# cephadm shell
2. List current hosts and labels:
Example
[ceph: root@host01 /]# ceph orch host ls
HOST
host01
host02
ADDR
LABELS
STATUS
_admin mon osd mgr
mon osd mgr mylabel
Method 1: Use the --placement option to deploy a daemon from the command line:
Syntax
ceph orch apply DAEMON --placement="label:LABEL"
Example
[ceph: root@host01 /]# ceph orch apply prometheus --placement="label:mylabel"
Method 2: To assign the daemon to a specific host label in a YAML file, specify the service type and label in the YAML file:
Create the placement.yml file:
Example
156 IBM Storage Ceph
[ceph: root@host01 /]# vi placement.yml
Specify the service type and label in the placement.yml file:
Example
service_type: prometheus
placement:
label: "mylabel"
Apply the daemon placement file:
Syntax
ceph orch apply -i FILENAME
Example
[ceph: root@host01 /]# ceph orch apply -i placement.yml
Scheduled prometheus update…
Verification
List the status of the daemons:
Syntax
ceph orch ps --daemon_type=DAEMON_NAME
Example
[ceph: root@host01 /]# ceph orch ps --daemon_type=prometheus
NAME
HOST
PORTS
STATUS
REFRESHED AGE MEM USE MEM LIM VERSION
prometheus.host02 host02 *:9095 running (2h)
8m ago
2h
85.3M
- 2.22.2
IMAGE ID
CONTAINER ID
ac25aac5d567 ad8c7593d7c0
Adding Monitor service
A typical IBM Storage Ceph storage cluster has three or five monitor daemons deployed on different hosts. If your storage cluster has five or more hosts, IBM recommends
that you deploy five Monitor nodes.
Note:
In the case of a firewall, see ../configuration/proc_net_firewall-settings-for-ceph-monitors.html.
The bootstrap node is the initial monitor of the storage cluster. Be sure to include the bootstrap node in the list of hosts to which you want to deploy.
If you want to apply Monitor service to more than one specific host, be sure to specify all of the host names within the same ceph orch apply command. If you
specify ceph
orch apply mon --placement host1 and then specify ceph orch apply mon --placement
host2, the second command removes the Monitor service on host1 and applies a Monitor service to host2.
If your Monitor nodes or your entire cluster are located on a single subnet, then cephadm automatically adds up to five Monitor daemons as you add new hosts to the
cluster. cephadm automatically configures the Monitor daemons on the new hosts. The new hosts reside on the same subnet as the first (bootstrap) host in the storage
cluster. cephadm can also deploy and scale monitors to correspond to changes in the size of the storage cluster.
Deploying Ceph monitor nodes using host labels
Adding Ceph Monitor nodes by IP address or network name
Prerequisites
Root-level access to all hosts in the storage cluster.
A running storage cluster.
Procedure
1. Apply the five Monitor daemons to five random hosts in the storage cluster:
ceph orch apply mon 5
2. Disable automatic Monitor deployment:
ceph orch apply mon --unmanaged
Adding Monitor nodes to specific hosts
Use host labels to identify the hosts that contain Monitor nodes.
Root-level access to all nodes in the storage cluster.
A running storage cluster.
1. Assign the mon label to the host:
Syntax
IBM Storage Ceph 157
ceph orch host label add HOSTNAME mon
Example
[ceph: root@host01 /]# ceph orch host label add host01 mon
2. View the current hosts and labels:
Syntax
ceph orch host ls
Example
[ceph: root@host01 /]# ceph orch host label add host02 mon
[ceph: root@host01 /]# ceph orch host label add host03 mon
[ceph: root@host01 /]# ceph orch host ls
HOST
ADDR
LABELS STATUS
host01
mon
host02
mon
host03
mon
host04
host05
host06
3. Deploy monitors based on the host label:
Syntax
ceph orch apply mon label:mon
4. Deploy monitors on a specific set of hosts:
Syntax
ceph orch apply mon HOSTNAME1,HOSTNAME2,HOSTNAME3
Example
[root@host01 ~]# ceph orch apply mon host01,host02,host03
NOTE: Be sure to include the bootstrap node in the list of hosts to which you want to deploy.
Deploying Ceph monitor nodes using host labels
A typical IBM Storage Ceph storage cluster has three or five Ceph Monitor daemons deployed on different hosts. If your storage cluster has five or more hosts, IBM
recommends that you deploy five Ceph Monitor nodes.
If your Ceph Monitor nodes or your entire cluster are located on a single subnet, then cephadm automatically adds up to five Ceph Monitor daemons as you add new
nodes to the cluster. cephadm automatically configures the Ceph Monitor daemons on the new nodes. The new nodes reside on the same subnet as the first (bootstrap)
node in the storage cluster. cephadm can also deploy and scale monitors to correspond to changes in the size of the storage cluster.
NOTE: Use host labels to identify the hosts that contain Ceph Monitor nodes.
Prerequisites
Root-level access to all nodes in the storage cluster.
A running storage cluster.
Procedure
1. Assign the mon label to the host:
Syntax
ceph orch host label add HOSTNAME mon
Example
[ceph: root@host01 /]# ceph orch host label add host02 mon
[ceph: root@host01 /]# ceph orch host label add host03 mon
2. View the current hosts and labels:
Syntax
ceph orch host ls
Example
[ceph: root@host01 /]# ceph orch host ls
HOST
ADDR
LABELS STATUS
host01
mon,mgr,_admin
host02
mon
host03
mon
host04
158 IBM Storage Ceph
host05
host06
Deploy Ceph Monitor daemons based on the host label:
Syntax
ceph orch apply mon label:mon
Deploy Ceph Monitor daemons on a specific set of hosts:
Syntax
ceph orch apply mon HOSTNAME1,HOSTNAME2,HOSTNAME3
Example
[ceph: root@host01 /]# ceph orch apply mon host01,host02,host03
NOTE: Be sure to include the bootstrap node in the list of hosts to which you want to deploy.
Adding Ceph Monitor nodes by IP address or network name
A typical IBM Storage Ceph storage cluster has three or five monitor daemons deployed on different hosts. If your storage cluster has five or more hosts, IBM recommends
that you deploy five Monitor nodes.
If your Monitor nodes or your entire cluster are located on a single subnet, then cephadm automatically adds up to five Monitor daemons as you add new nodes to the
cluster. You do not need to configure the Monitor daemons on the new nodes. The new nodes reside on the same subnet as the first node in the storage cluster. The first
node in the storage cluster is the bootstrap node. cephadm can also deploy and scale monitors to correspond to changes in the size of the storage cluster.
Prerequisites
Root-level access to all nodes in the storage cluster.
A running storage cluster.
Procedure
1. To deploy each additional Ceph Monitor node:
Syntax
ceph orch apply mon NODE:IP_ADDRESS_OR_NETWORK_NAME [NODE:IP_ADDRESS_OR_NETWORK_NAME...]
Example
[ceph: root@host01 /]# ceph orch apply mon host02:10.10.128.69 host03:mynetwork
Setting up the admin node
Use an admin node to administer the storage cluster.
An admin node contains both the cluster configuration file and the admin keyring. Both of these files are stored in the directory /etc/ceph and use the name of the
storage cluster as a prefix.
For example, the default ceph cluster name is ceph. In a cluster using the default name, the admin keyring is named /etc/ceph/ceph.client.admin.keyring. The
corresponding cluster configuration file is named /etc/ceph/ceph.conf.
To set up additional hosts in the storage cluster as admin nodes, apply the _admin label to the host you want to designate as an administrator node.
NOTE: By default, after applying the _admin label to a node, cephadm copies the ceph.conf and client.admin keyring files to that node. The _admin label is
automatically applied to the bootstrap node unless the --skip-admin-label option was specified with the cephadm bootstrap command.
Removing the admin label from a host
Prerequisites
A running storage cluster with cephadm installed.
The storage cluster has running Monitor and Manager nodes.
Root-level access to all nodes in the cluster.
Procedure
1. Use ceph orch host ls to view the hosts in your storage cluster:
Example
IBM Storage Ceph 159
[root@host01 ~]# ceph orch host ls
HOST
ADDR
LABELS STATUS
host01
mon,mgr,_admin
host02
mon
host03
mon,mgr
host04
host05
host06
2. Use the _admin label to designate the admin host in your storage cluster. For best results, this host should have both Monitor and Manager daemons running.
Syntax
ceph orch host label add HOSTNAME _admin
Example
[root@host01 ~]#
ceph orch host label add host03 _admin
3. Verify that the admin host has the _admin label.
Example
[root@host01 ~]# ceph orch host ls
HOST
ADDR
LABELS STATUS
host01
mon,mgr,_admin
host02
mon
host03
mon,mgr,_admin
host04
host05
host06
4. Log in to the admin node to manage the storage cluster.
Removing the admin label from a host
You can use the Ceph orchestrator to remove the admin label from a host.
Prerequisites
A running storage cluster with cephadm installed and bootstrapped.
The storage cluster has running Monitor and Manager nodes.
Root-level access to all nodes in the cluster.
Procedure
1. Use ceph orch host ls to view the hosts and identify the admin host in your storage cluster:
Example
[root@host01 ~]# ceph orch host ls
HOST
ADDR
LABELS STATUS
host01
mon,mgr,_admin
host02
mon
host03
mon,mgr,_admin
host04
host05
host06
2. Log into the Cephadm shell:
Example
[root@host01 ~]# cephadm shell
3. Use the ceph orchestrator to remove the admin label from a host:
Syntax
ceph orch host label rm HOSTNAME LABEL
Example
[ceph: root@host01 /]# ceph orch host label rm host03 _admin
4. Verify that the admin host has the _admin label.
Example
[root@host01 ~]# ceph orch host ls
HOST
ADDR
LABELS STATUS
host01
mon,mgr,_admin
host02
mon
host03
mon,mgr
host04
160 IBM Storage Ceph
host05
host06
Important: After removing the admin label from a node, ensure you remove the ceph.conf and client.admin keyring files from that node. Also, the node must be
removed from the admin Ansible inventory file.
Adding Manager service
cephadm automatically installs a Manager daemon on the bootstrap node during the bootstrapping process. Use the Ceph orchestrator to deploy additional Manager
daemons.
The Ceph orchestrator deploys two Manager daemons by default. To deploy a different number of Manager daemons, specify a different number. If you do not specify the
hosts where the Manager daemons should be deployed, the Ceph orchestrator randomly selects the hosts and deploys the Manager daemons to them.
NOTE: If you want to apply Manager daemons to more than one specific host, be sure to specify all of the host names within the same ceph orch apply command. If
you specify ceph orch apply mgr --placement host1 and then specify ceph orch
apply mgr --placement host2, the second command removes the Manager daemon on host1 and applies a Manager daemon to host2.
Use the --placement option to deploy to specific hosts.
Prerequisites
A running storage cluster.
Procedure
To specify that you want to apply a certain number of Manager daemons to randomly selected hosts:
Syntax
ceph orch apply mgr NUMBER_OF_DAEMONS
Example
[ceph: root@host01 /]# ceph orch apply mgr 3
To add Manager daemons to specific hosts in your storage cluster:
Syntax
ceph orch apply mgr --placement "HOSTNAME1 HOSTNAME2 HOSTNAME3"
Example
[ceph: root@host01 /]# ceph orch apply mgr --placement "host02 host03 host04"
Adding OSDs
Cephadm does not provision an OSD on a device that is not available. A storage device is considered available if it meets all of the following conditions:
The device must have no partitions.
The device must not be mounted.
The device must not contain a file system.
The device must not contain a Ceph BlueStore OSD.
The device must be larger than 5 GB.
Important: By default, the osd_memory_target_autotune parameter is set to true in IBM Storage Ceph.
See Automatically tuning OSD memory
for tuning OSD memory automatically.
Prerequisites
A running IBM Storage Ceph cluster.
Procedure
1. List the available devices to deploy OSDs:
Syntax
ceph orch device ls [--hostname=HOSTNAME1 HOSTNAME2] [--wide] [--refresh]
Example
IBM Storage Ceph 161
[ceph: root@host01 /]# ceph orch device ls --wide --refresh
2. You can either deploy the OSDs on specific hosts or on all the available devices:
To create an OSD from a specific device on a specific host:
Syntax
ceph orch daemon add osd HOSTNAME:DEVICE_PATH
Example
[ceph: root@host01 /]# ceph orch daemon add osd host02:/dev/sdb
To deploy OSDs on any available and unused devices, use the --all-available-devices option.
Example
[ceph: root@host01 /]# ceph orch apply osd --all-available-devices
Note: This command creates colocated WAL and DB daemons. If you want to create non-colocated daemons, do not use this command.
Reference
For more information about drive specifications for OSDs, see Advanced service specifications and filters for deploying OSDs
For more information about zapping devices to clear data on devices, see Zapping devices for Ceph OSD deployment
Purging the Ceph storage cluster
Purging the Ceph storage cluster clears any data or connections that remain from previous deployments on your server.
About this task
For Red Hat Enterprise Linux 8, this Ansible script removes all daemons, logs, and data that belong to the FSID passed to the script from all hosts in the storage cluster.
For Red Hat Enterprise Linux 9, use the cephadm rm-cluster command, since Ansible is not supported.
Red Hat Enterprise Linux 8
About this task
Important: This process works only if the cephadm binary is installed on all hosts in the storage cluster.
The Ansible inventory file lists all the hosts in your cluster and what roles each host plays in your Ceph storage cluster. The default location for an inventory file is
/usr/share/cephadm-ansible/hosts, but this file can be placed anywhere.
The following example shows the structure of an inventory file:
Example
host02
host03
host04
[admin]
host01
[clients]
client01
client02
client03
Before you begin
Before purging the Ceph storage cluster on a Red Hat Enterprise Linux 8 operating system, be sure that you have the following:
A running IBM Storage Ceph.
Ansible 2.12 or later is installed on the bootstrap node.
Root-level access to the Ansible administration node.
Ansible user with sudo and passwordless ssh access to all nodes in the storage cluster.
The [admin] group is defined in the inventory file with a node where the admin keyring is present at /etc/ceph/ceph.client.admin.keyring.
Procedure
As an Ansible user on the bootstrap node, run the purge script.
Syntax
ansible-playbook -i hosts cephadm-purge-cluster.yml -e fsid=FSID -vvv
Example
162 IBM Storage Ceph
[ansible@host01 cephadm-ansible]$ ansible-playbook -i hosts cephadm-purge-cluster.yml -e fsid=a6ca415a-cde7-11eb-a41a002590fc2544 -vvv
Note: An additional extra-var (-e ceph_origin=ibm) is required to zap the disk devices during the purge.
When the script has completed, the entire storage cluster, including all OSD disks, will have been removed from all hosts in the cluster.
Red Hat Enterprise Linux 9
Before you begin
Before purging the Ceph storage cluster on a Red Hat Enterprise Linux 8 operating system, be sure that you have a running IBM Storage Ceph.
Procedure
1. Disable cephadm to stop all the orchestration operations to avoid deploying new daemons.
Example
[ceph: root#host01 /]# ceph mgr module disable cephadm
2. Get the FSID of the cluster.
Example
[ceph: root#host01 /]# ceph fsid
3. Exit the shell.
[ceph: root#host01 /]# cexit
4. Purge the Ceph daemons from all hosts in the cluster.
Syntax
cephadm rm-cluster --force --zap-osds --fsid FSID
Example
[root@host01 ~]# cephadm rm-cluster --force --zap-osds --fsid a6ca415a-cde7-11eb-a41a-002590fc2544
Deploying client nodes
As a storage administrator, you can deploy client nodes by running the cephadm-preflight.yml and cephadm-clients.yml playbooks. The cephadmpreflight.yml playbook configures the Ceph repository and prepares the storage cluster for bootstrapping. It also installs some prerequisites, such as podman, lvm2,
chronyd, and cephadm.
Important: If using Red Hat Enterprise Linux 9, do not use these steps. The cephadm-preflight playbook is not supported.
The cephadm-clients.yml playbook handles the distribution of configuration and keyring files to a group of Ceph clients.
Note: If you are not using the cephadm-ansible playbooks, after upgrading your Ceph cluster, you must upgrade the ceph-common package and client libraries on your
client nodes. For more information, see Upgrading the IBM Storage Ceph cluster.
Prerequisites
Root-level access to the Ansible administration node.
Ansible user with sudo and passwordless SSH access to all nodes in the storage cluster.
Installation of the cephadm-ansible package.
The [admin] group is defined in the inventory file with a node where the admin keyring is present at /etc/ceph/ceph.client.admin.keyring.
Procedure
1. As an Ansible user, navigate to the /usr/share/cephadm-ansible directory on the Ansible administration node.
Example
[ceph-admin@admin ~]$ cd /usr/share/cephadm-ansible
2. Open and edit the hosts inventory file and add the [clients] group and clients to your inventory:
Example
host02
host03
host04
[clients]
client01
client02
client03
[admin]
host01
3. Run the cephadm-preflight.yml playbook to install the prerequisites on the clients:
IBM Storage Ceph 163
Syntax
ansible-playbook -i INVENTORY_FILE cephadm-preflight.yml --limit CLIENT_GROUP_NAME|CLIENT_NODE_NAME
Example
[ceph-admin@admin cephadm-ansible]$ ansible-playbook -i hosts cephadm-preflight.yml --limit clients
4. Run the cephadm-clients.yml playbook to distribute the keyring and Ceph configuration files to a set of clients.
a. To copy the keyring with a custom destination keyring name:
Syntax
ansible-playbook -i INVENTORY_FILE cephadm-clients.yml --extra-vars
'{"fsid":"FSID","keyring":"KEYRING_PATH","client_group":"CLIENT_GROUP_NAME","conf":"CEPH_CONFIGURATION_PATH","keyring
_dest":"KEYRING_DESTINATION_PATH"}'
Replace INVENTORY_FILE with the Ansible inventory file name.
Replace FSID with the FSID of the cluster.
Replace KEYRING_PATH with the full path name to the keyring on the admin host that you want to copy to the client.
Optional: Replace CLIENT_GROUP_NAME with the Ansible group name for the clients to set up.
Optional: Replace CEPH_CONFIGURATION_PATH with the full path to the Ceph configuration file on the admin node.
Optional: Replace KEYRING_DESTINATION_PATH with the full path name of the destination where the keyring will be copied.
NOTE: If you do not specify a configuration file with the conf option when you run the playbook, the playbook generates and distributes a minimal
configuration file. By default, the generated file is located at /etc/ceph/ceph.conf.
Example
[ceph-admin@host01 cephadm-ansible]$ ansible-playbook -i hosts cephadm-clients.yml --extra-vars '{"fsid":"266ee7a82a05-11eb-b8465254002d4916","keyring":"/etc/ceph/ceph.client.admin.keyring","client_group":"clients","conf":"/etc/ceph/ceph.conf","
keyring_dest":"/etc/ceph/custom.name.ceph.keyring"}'
b. To copy a keyring with the default destination keyring name of ceph.keyring and using the default group of clients:
Syntax
ansible-playbook -i INVENTORY_FILE cephadm-clients.yml --extra-vars
'{"fsid":"FSID","keyring":"KEYRING_PATH","conf":"CONF_PATH"}'
Example
[ceph-admin@host01 cephadm-ansible]$ ansible-playbook -i hosts cephadm-clients.yml --extra-vars '{"fsid":"266ee7a82a05-11eb-b846-5254002d4916","keyring":"/etc/ceph/ceph.client.admin.keyring","conf":"/etc/ceph/ceph.conf"}'
Verification
Log into the client nodes and verify that the keyring and configuration files exist.
Example
[user@client01 ~]# ls -l /etc/ceph/
-rw-------. 1 ceph ceph 151 Jul 11 12:23 custom.name.ceph.keyring
-rw-------. 1 ceph ceph 151 Jul 11 12:23 ceph.keyring
-rw-------. 1 ceph ceph 269 Jul 11 12:23 ceph.conf
Reference
For more information about admin keys, see Ceph User Management.
For more information about the cephadm-preflight playbook, see Running the preflight playbook.
Managing an IBM Storage Ceph cluster using cephadm-ansible modules
As a storage administrator, you can use cephadm-ansible modules in Ansible playbooks to administer your IBM Storage Ceph cluster. The cephadm-ansible package
provides several modules that wrap cephadm calls to let you write your own unique Ansible playbooks to administer your cluster.
Note: At this time, cephadm-ansible modules only support the most important tasks. Any operation not covered by cephadm-ansible modules must be completed
using either the command or shell Ansible modules in your playbooks.
cephadm-ansible modules
cephadm-ansible modules options
This section lists the available options for the cephadm-ansible modules.
Bootstrapping a storage cluster using the cephadm_ansible modules
Adding or removing hosts using the ceph_orch_host module
Setting configuration options using the ceph_config module
Applying a service specification using the ceph_orch_apply module
164 IBM Storage Ceph
Managing Ceph daemon states using the ceph_orch_daemon module
cephadm-ansible modules
The cephadm-ansible modules are a collection of modules that simplify writing Ansible playbooks by providing a wrapper around cephadm and ceph
orch commands. You can use the modules to write your own unique Ansible playbooks to administer your cluster using one or more of the modules.
The cephadm-ansible package includes the following modules:
cephadm_bootstrap
ceph_orch_host
ceph_config
ceph_orch_apply
ceph_orch_daemon
cephadm_registry_login
cephadm-ansible modules options
This section lists the available options for the cephadm-ansible modules.
Options listed as required need to be set when using the modules in your Ansible playbooks. Options listed with a default value of true indicate that the option is
automatically set when using the modules and you do not need to specify it in your playbook. For example, for the cephadm_bootstrap module, the Ceph Dashboard is
installed unless you set dashboard: false.
Table 1. Available options for the cephadm_bootstrap module
Description
cephadm_bootstrap
Required
Default
mon_ip
Ceph Monitor IP address.
true
image
Ceph container image.
false
docker
Use docker instead of podman.
false
fsid
Define the Ceph FSID.
false
pull
Pull the Ceph container image.
false
true
dashboard
Deploy the Ceph Dashboard.
false
true
dashboard_user
Specify a specific Ceph Dashboard user.
false
dashboard_password
Ceph Dashboard password.
false
monitoring
Deploy the monitoring stack.
false
true
firewalld
Manage firewall rules with firewalld.
false
true
allow_overwrite
Allow overwrite of existing --output-config, --output-keyring, or --output-pub-ssh-key files. false
false
registry_url
URL for custom registry.
false
registry_username
Username for custom registry.
false
registry_password
Password for custom registry.
false
registry_json
JSON file with custom registry login information.
false
ssh_user
SSH user to use for cephadm ssh to hosts.
false
ssh_config
SSH config file path for cephadm SSH client.
false
allow_fqdn_hostname Allow hostname that is a fully-qualified domain name (FQDN).
false
cluster_network
false
Subnet to use for cluster replication, recovery and heartbeats.
false
Table 2. Available options for the ceph_orch_host module
ceph_orch_ho
st
fsid
The FSID of the Ceph cluster to interact with.
image
The Ceph container image to use.
Description
Required
Default
false
false
name
Name of the host to add, remove, or update.
true
address
IP address of the host.
true when
state is
present.
set_admin_la Set the _admin label on the specified host.
bel
labels
The list of labels to apply to the host.
state
If set to present, it ensures the name specified in name is present. If set to absent, it removes the host
false
false
false
[]
false
present
specified in name. If set to drain, it schedules to remove all daemons from the host specified in name.
Table 3. Available options for the ceph_config module
ceph_config
Description
Required
fsid
The FSID of the Ceph cluster to interact with.
false
image
The Ceph container image to use.
false
action
Whether to set or get the parameter specified in option. false
who
Which daemon to set the configuration to.
Default
set
true
IBM Storage Ceph 165
Description
ceph_config
Required
option
Name of the parameter to set or get.
true
value
Value of the parameter to set.
true if action is set
Default
Table 4. Available options for the ceph_orch_apply module
Description
ceph_orch_apply
Required
fsid
The FSID of the Ceph cluster to interact with. false
image
The Ceph container image to use.
false
spec
The service specification to apply.
true
Table 5. Available options for the ceph_orch_daemon module
Description
ceph_orch_daemon
Required
fsid
The FSID of the Ceph cluster to interact with.
false
image
The Ceph container image to use.
false
state
The desired state of the service specified in name. true
If started, it ensures the service is started.
If stopped, it ensures the service is stopped.
If restarted, it will restart the service.
daemon_id
The ID of the service.
daemon_type
The type of service.
true
true
Table 6. Available options for the cephadm_registry_login module
cephadm_registry_
login
state
Login or logout of a registry.
docker
Use docker instead of podman.
registry_url
Description
The URL for custom registry.
Required
false
false
false
registry_username Username for custom registry.
true when state is
login.
registry_password Password for custom registry.
true when state is
login.
registry_json
Default
login
The path to a JSON file. This file must be present on remote hosts prior to running this task. This
option is currently not supported.
Bootstrapping a storage cluster using the cephadm_ansible modules
As a storage administrator, you can bootstrap a storage cluster using the cephadm-ansible modules such as cephadm_bootstrap and cephadm_registry_login
playbook.
Prerequisites
An IP address for the first Ceph Monitor container, which is also the IP address for the first node in the storage cluster.
Login access to cp.icr.io/cp.
A minimum of 10 GB of free space for /var/lib/containers/.
Installation of the cephadm-ansible package on the Ansible administration node.
Passwordless SSH is set up on all hosts in the storage cluster.
Hosts are registered with CDN.
For the latest supported Red Hat Enterprise Linux versions, see Compatibility matrix.
Procedure
1. Log in to the Ansible administration node.
2. Navigate to the /usr/share/cephadm-ansible directory on the Ansible administration node:
Example
[ansible@admin ~]$ cd /usr/share/cephadm-ansible
3. Create the hosts file and add hosts, labels, and monitor IP address of the first host in the storage cluster:
Syntax
sudo vi INVENTORY_FILE
HOST1 labels="[LABEL1, LABEL2]"
HOST2 labels="[LABEL1, LABEL2]"
HOST3 labels="[LABEL1]"
166 IBM Storage Ceph
[admin]
ADMIN_HOST monitor_address=MONITOR_IP_ADDRESS labels="[ADMIN_LABEL, LABEL1, LABEL2]"
Example
[ansible@admin cephadm-ansible]$ sudo vi hosts
host02 labels="['mon', 'mgr']"
host03 labels="['mon', 'mgr']"
host04 labels="['osd']"
host05 labels="['osd']"
host06 labels="['osd']"
[admin]
host01 monitor_address=10.10.128.68 labels="['_admin', 'mon', 'mgr']"
4. Run the preflight playbook:
Syntax
ansible-playbook -i INVENTORY_FILE cephadm-preflight.yml --extra-vars "ceph_origin=ibm"
Example
[ansible@admin cephadm-ansible]$ ansible-playbook -i hosts cephadm-preflight.yml --extra-vars "ceph_origin=ibm"
5. Create a playbook to bootstrap your cluster:
Syntax
sudo vi PLAYBOOK_FILENAME.yml
--- name: NAME_OF_PLAY
hosts: BOOTSTRAP_HOST
become: USE_ELEVATED_PRIVILEGES
gather_facts: GATHER_FACTS_ABOUT_REMOTE_HOSTS
tasks:
-name: NAME_OF_TASK
cephadm_registry_login:
state: STATE
registry_url: REGISTRY_URL
registry_username: REGISTRY_USER_NAME
registry_password: REGISTRY_PASSWORD
- name: NAME_OF_TASK
cephadm_bootstrap:
mon_ip: "{{ monitor_address }}"
dashboard_user: DASHBOARD_USER
dashboard_password: DASHBOARD_PASSWORD
allow_fqdn_hostname: ALLOW_FQDN_HOSTNAME
cluster_network: NETWORK_CIDR
Example
[ansible@admin cephadm-ansible]$ sudo vi bootstrap.yml
--- name: bootstrap the cluster
hosts: host01
become: true
gather_facts: false
tasks:
- name: login to registry
cephadm_registry_login:
state: login
registry_url: cp.icr.io/cp
registry_username: user1
registry_password: mypassword1
- name: bootstrap initial cluster
cephadm_bootstrap:
mon_ip: "{{ monitor_address }}"
dashboard_user: mydashboarduser
dashboard_password: mydashboardpassword
allow_fqdn_hostname: true
cluster_network: 10.10.128.0/28
6. Run the playbook:
Syntax
ansible-playbook -i INVENTORY_FILE PLAYBOOK_FILENAME.yml -vvv
Example
ansible@admin cephadm-ansible]$ ansible-playbook -i hosts bootstrap.yml -vvv
Review the Ansible output after running the playbook.
Adding or removing hosts using the ceph_orch_host module
IBM Storage Ceph 167
Add and remove hosts in your storage cluster by using the ceph_orch_host module in your Ansible playbook.
Prerequisites
A running IBM Storage Ceph cluster.
Register the nodes to the CDN and attach subscriptions.
Ansible user with sudo and passwordless SSH access to all nodes in the storage cluster.
Installation of the cephadm-ansible package on the Ansible administration node.
New hosts have the storage cluster’s public SSH key. For more information about copying the storage cluster’s public SSH keys to new hosts, see Adding hosts.
Procedure
1. Use the following procedure to add new hosts to the cluster:
a. Log in to the Ansible administration node.
b. Navigate to the /usr/share/cephadm-ansible directory on the Ansible administration node:
Example
[ansible@admin ~]$ cd /usr/share/cephadm-ansible
c. Add the new hosts and labels to the Ansible inventory file.
Syntax
sudo vi INVENTORY_FILE
NEW_HOST1 labels="[LABEL1, LABEL2]"
NEW_HOST2 labels="[LABEL1, LABEL2]"
NEW_HOST3 labels="[LABEL1]"
[admin]
ADMIN_HOST monitor_address=MONITOR_IP_ADDRESS labels="[ADMIN_LABEL, LABEL1, LABEL2]"
Example
[ansible@admin cephadm-ansible]$ sudo vi hosts
host02 labels="['mon', 'mgr']"
host03 labels="['mon', 'mgr']"
host04 labels="['osd']"
host05 labels="['osd']"
host06 labels="['osd']"
[admin]
host01 monitor_address= 10.10.128.68 labels="['_admin', 'mon', 'mgr']"
NOTE: If you have previously added the new hosts to the Ansible inventory file and ran the preflight playbook on the hosts, skip to step 3.
d. Run the preflight playbook with the --limit option:
Syntax
ansible-playbook -i INVENTORY_FILE cephadm-preflight.yml --extra-vars "ceph_origin=ibm" --limit NEWHOST
Example
[ansible@admin cephadm-ansible]$ ansible-playbook -i hosts cephadm-preflight.yml --extra-vars "ceph_origin=ibm" -limit host02
The preflight playbook installs podman, lvm2, chrony, and cephadm on the new host. After installation is complete, cephadm resides in the /usr/sbin/
directory.
e. Create a playbook to add the new hosts to the cluster:
Syntax
sudo vi PLAYBOOK_FILENAME.yml
--- name: PLAY_NAME
hosts: HOSTS_OR_HOST_GROUPS
become: USE_ELEVATED_PRIVILEGES
gather_facts: GATHER_FACTS_ABOUT_REMOTE_HOSTS
tasks:
- name: NAME_OF_TASK
ceph_orch_host:
name: "{{ ansible_facts[hostname] }}"
address: "{{ ansible_facts[default_ipv4][address] }}"
labels: "{{ labels }}"
delegate_to: HOST_TO_DELEGATE_TASK_TO
- name: NAME_OF_TASK
when: inventory_hostname in groups[admin]
ansible.builtin.shell:
cmd: CEPH_COMMAND_TO_RUN
168 IBM Storage Ceph
register: REGISTER_NAME
- name: NAME_OF_TASK
when: inventory_hostname in groups[admin]
debug:
msg: "{{ REGISTER_NAME.stdout }}"
NOTE: By default, Ansible executes all tasks on the host that matches the hosts line of your playbook. The ceph orch commands must run on the host
that contains the admin keyring and the Ceph configuration file. Use the delegate_to keyword to specify the admin host in your cluster.
Example
[ansible@admin cephadm-ansible]$ sudo vi add-hosts.yml
--- name: add additional hosts to the cluster
hosts: all
become: true
gather_facts: true
tasks:
- name: add hosts to the cluster
ceph_orch_host:
name: "{{ ansible_facts['hostname'] }}"
address: "{{ ansible_facts['default_ipv4']['address'] }}"
labels: "{{ labels }}"
delegate_to: host01
- name: list hosts in the cluster
when: inventory_hostname in groups['admin']
ansible.builtin.shell:
cmd: ceph orch host ls
register: host_list
- name: print current list of hosts
when: inventory_hostname in groups['admin']
debug:
msg: "{{ host_list.stdout }}"
In this example, the playbook adds the new hosts to the cluster and displays a current list of hosts.
f. Run the playbook to add additional hosts to the cluster:
Syntax
ansible-playbook -i INVENTORY_FILE PLAYBOOK_FILENAME.yml
Example
[ansible@admin cephadm-ansible]$ ansible-playbook -i hosts add-hosts.yml
2. Use the following procedure to remove hosts from the cluster:
a. Log in to the Ansible administration node.
b. Navigate to the /usr/share/cephadm-ansible directory on the Ansible administration node:
Example
[ansible@admin ~]$ cd /usr/share/cephadm-ansible
c. Create a playbook to remove a host or hosts from the cluster:
Syntax
sudo vi PLAYBOOK_FILENAME.yml
--- name: NAME_OF_PLAY
hosts: ADMIN_HOST
become: USE_ELEVATED_PRIVILEGES
gather_facts: GATHER_FACTS_ABOUT_REMOTE_HOSTS
tasks:
- name: NAME_OF_TASK
ceph_orch_host:
name: HOST_TO_REMOVE
state: STATE
- name: NAME_OF_TASK
ceph_orch_host:
name: HOST_TO_REMOVE
state: STATE
retries: NUMBER_OF_RETRIES
delay: DELAY
until: CONTINUE_UNTIL
register: REGISTER_NAME
- name: NAME_OF_TASK
ansible.builtin.shell:
cmd: ceph orch host ls
register: REGISTER_NAME
- name: NAME_OF_TASK
debug:
msg: "{{ REGISTER_NAME.stdout }}"
IBM Storage Ceph 169
Example
[ansible@admin cephadm-ansible]$ sudo vi remove-hosts.yml
--- name: remove host
hosts: host01
become: true
gather_facts: true
tasks:
- name: drain host07
ceph_orch_host:
name: host07
state: drain
- name: remove host from the cluster
ceph_orch_host:
name: host07
state: absent
retries: 20
delay: 1
until: result is succeeded
register: result
- name: list hosts in the cluster
ansible.builtin.shell:
cmd: ceph orch host ls
register: host_list
- name: print current list of hosts
debug:
msg: "{{ host_list.stdout }}"
In this example, the playbook tasks drain all daemons on host07, removes the host from the cluster, and displays a current list of hosts.
d. Run the playbook to remove host from the cluster:
Syntax
ansible-playbook -i INVENTORY_FILE PLAYBOOK_FILENAME.yml
Example
[ansible@admin cephadm-ansible]$ ansible-playbook -i hosts remove-hosts.yml
Verification
Review the Ansible task output displaying the current list of hosts in the cluster:
Example
TASK [print current hosts]
******************************************************************************************************
Friday 24 June 2022 14:52:40 -0400 (0:00:03.365)
0:02:31.702 ***********
ok: [host01] =>
msg: |HOST
ADDR
LABELS
STATUS
host01 10.10.128.68
_admin mon mgr
host02 10.10.128.69
mon mgr
host03 10.10.128.70
mon mgr
host04 10.10.128.71
osd
host05 10.10.128.72
osd
host06 10.10.128.73
osd
Setting configuration options using the ceph_config module
As a storage administrator, you can set or get IBM Storage Ceph configuration options using the ceph_config module.
Prerequisites
A running IBM Storage Ceph cluster.
Ansible user with sudo and passwordless SSH access to all nodes in the storage cluster.
Installation of the cephadm-ansible package on the Ansible administration node.
The Ansible inventory file contains the cluster and admin hosts. For more information about adding hosts to your storage cluster, see Adding or removing hosts
using the ceph_orch_host module.
Procedure
1. Log in to the Ansible administration node.
2. Navigate to the /usr/share/cephadm-ansible directory on the Ansible administration node:
Example
170 IBM Storage Ceph
[ansible@admin ~]$ cd /usr/share/cephadm-ansible
3. Create a playbook with configuration changes:
Syntax
sudo vi PLAYBOOK_FILENAME.yml
--- name: PLAY_NAME
hosts: ADMIN_HOST
become: USE_ELEVATED_PRIVILEGES
gather_facts: GATHER_FACTS_ABOUT_REMOTE_HOSTS
tasks:
- name: NAME_OF_TASK
ceph_config:
action: GET_OR_SET
who: DAEMON_TO_SET_CONFIGURATION_TO
option: CEPH_CONFIGURATION_OPTION
value: VALUE_OF_PARAMETER_TO_SET
- name: NAME_OF_TASK
ceph_config:
action: GET_OR_SET
who: DAEMON_TO_SET_CONFIGURATION_TO
option: CEPH_CONFIGURATION_OPTION
register: REGISTER_NAME
- name: NAME_OF_TASK
debug:
msg: "MESSAGE_TO_DISPLAY {{ REGISTER_NAME.stdout }}"
Example
[ansible@admin cephadm-ansible]$ sudo vi change_configuration.yml
--- name: set pool delete
hosts: host01
become: true
gather_facts: false
tasks:
- name: set the allow pool delete option
ceph_config:
action: set
who: mon
option: mon_allow_pool_delete
value: true
- name: get the allow pool delete setting
ceph_config:
action: get
who: mon
option: mon_allow_pool_delete
register: verify_mon_allow_pool_delete
- name: print current mon_allow_pool_delete setting
debug:
msg: "the value of 'mon_allow_pool_delete' is {{ verify_mon_allow_pool_delete.stdout }}"
In this example, the playbook first sets the mon_allow_pool_delete option to false. The playbook then gets the current mon_allow_pool_delete setting
and displays the value in the Ansible output.
4. Run the playbook:
Syntax
ansible-playbook -i INVENTORY_FILE _PLAYBOOK_FILENAME.yml
Example
[ansible@admin cephadm-ansible]$ ansible-playbook -i hosts change_configuration.yml
Verification
Review the output from the playbook tasks.
Example
TASK [print current mon_allow_pool_delete setting] *************************************************************
Wednesday 29 June 2022 13:51:41 -0400 (0:00:05.523)
0:00:17.953 ********
ok: [host01] =>
msg: the value of 'mon_allow_pool_delete' is true
Reference
For more information, see Configuring.
Applying a service specification using the ceph_orch_apply module
IBM Storage Ceph 171
As a storage administrator, you can apply service specifications to your storage cluster using the ceph_orch_apply module in your Ansible playbooks. A service
specification is a data structure to specify the service attributes and configuration settings that is used to deploy the Ceph service. You can use a service specification to
deploy Ceph service types like mon, crash, mds, mgr, osd, rdb, or rbd-mirror.
Prerequisites
A running IBM Storage Ceph cluster.
Ansible user with sudo and passwordless SSH access to all nodes in the storage cluster.
Installation of the cephadm-ansible package on the Ansible administration node.
The Ansible inventory file contains the cluster and admin hosts. For more information about adding hosts to your storage cluster, see Adding or removing hosts
using the ceph_orch_host module.
Procedure
1. Log in to the Ansible administration node.
2. Navigate to the /usr/share/cephadm-ansible directory on the Ansible administration node:
Example
[ansible@admin ~]$ cd /usr/share/cephadm-ansible
3. Create a playbook with the service specifications:
Syntax
sudo vi PLAYBOOK_FILENAME.yml
--- name: PLAY_NAME
hosts: HOSTS_OR_HOST_GROUPS
become: USE_ELEVATED_PRIVILEGES
gather_facts: GATHER_FACTS_ABOUT_REMOTE_HOSTS
tasks:
- name: NAME_OF_TASK
ceph_orch_apply:
spec: |
service_type: SERVICE_TYPE
service_id: UNIQUE_NAME_OF_SERVICE
placement:
host_pattern: HOST_PATTERN_TO_SELECT_HOSTS
label: LABEL
spec:
SPECIFICATION_OPTIONS:
Example
[ansible@admin cephadm-ansible]$ sudo vi deploy_osd_service.yml
--- name: deploy osd service
hosts: host01
become: true
gather_facts: true
tasks:
- name: apply osd spec
ceph_orch_apply:
spec: |
service_type: osd
service_id: osd
placement:
host_pattern: '*'
label: osd
spec:
data_devices:
all: true
In this example, the playbook deploys the Ceph OSD service on all hosts with the label osd.
4. Run the playbook:
Syntax
ansible-playbook -i INVENTORY_FILE _PLAYBOOK_FILENAME.yml
Example
[ansible@admin cephadm-ansible]$ ansible-playbook -i hosts deploy_osd_service.yml
Verification
Review the output from the playbook tasks.
Managing Ceph daemon states using the ceph_orch_daemon module
172 IBM Storage Ceph
Start, stop, and restart Ceph daemons on hosts using the ceph_orch_daemon module in your Ansible playbooks.
Prerequisites
A running IBM Storage Ceph cluster.
Ansible user with sudo and passwordless SSH access to all nodes in the storage cluster.
Installation of the cephadm-ansible package on the Ansible administration node.
The Ansible inventory file contains the cluster and admin hosts. For more information about adding hosts to your storage cluster, see Adding or removing hosts
using the ceph_orch_host module.
Procedure
1. Log in to the Ansible administration node.
2. Navigate to the /usr/share/cephadm-ansible directory on the Ansible administration node:
Example
[ansible@admin ~]$ cd /usr/share/cephadm-ansible
3. Create a playbook with daemon state changes:
Syntax
sudo vi PLAYBOOK_FILENAME.yml
--- name: PLAY_NAME
hosts: ADMIN_HOST
become: USE_ELEVATED_PRIVILEGES
gather_facts: GATHER_FACTS_ABOUT_REMOTE_HOSTS
tasks:
- name: NAME_OF_TASK
ceph_orch_daemon:
state: STATE_OF_SERVICE
daemon_id: DAEMON_ID
daemon_type: TYPE_OF_SERVICE
Example
[ansible@admin cephadm-ansible]$ sudo vi restart_services.yml
--- name: start and stop services
hosts: host01
become: true
gather_facts: false
tasks:
- name: start osd.0
ceph_orch_daemon:
state: started
daemon_id: 0
daemon_type: osd
- name: stop mon.host02
ceph_orch_daemon:
state: stopped
daemon_id: host02
daemon_type: mon
In this example, the playbook starts the OSD with an ID of 0 and stops a Ceph Monitor with an id of host02.
4. Run the playbook:
Syntax
ansible-playbook -i INVENTORY_FILE _PLAYBOOK_FILENAME.yml
Example
[ansible@admin cephadm-ansible]$ ansible-playbook -i hosts restart_services.yml
Verification
Review the output from the playbook tasks.
Comparison between Ceph Ansible and Cephadm
Understand the differences between Cephadm and Ceph-Ansible playbooks for the containerized deployment of the storage cluster.
The tables compare Cephadm with Ceph-Ansible playbooks for managing the containerized deployment of a Ceph cluster for day one and day two operations.
IBM Storage Ceph 173
Table 1. Day one operations
Description
Ceph-Ansible
Cephadm
Installation of the IBM
Storage Ceph cluster
Run the site-container.yml playbook.
Run cephadm bootstrap command to bootstrap the cluster on the admin node.
Addition of hosts
Use the Ceph Ansible inventory.
Run ceph orch add host HOST_NAME to add hosts to the cluster.
Gathering Ceph logs
Run the gather ceph logs playbook.
Run the journalctl command.
Addition of monitors
Run the add-mon.yml playbook.
Run the ceph orch apply mon command.
Addition of managers
Run the site-container.yml playbook.
Run the ceph orch apply mgr command.
Addition of OSDs
Run the add-osd.yml playbook.
Run the ceph orch apply osd command to add OSDs on all available devices or on
specific hosts.
Addition of OSDs on
specific devices
Select the devices in the osd.yml file and then
run the add-osd.yml playbook.
Select the paths filter under the data_devices in the osd.yml file and then run
ceph orch apply -i FILE_NAME.yml command.
Addition of MDS
Run the site-container.yml playbook.
Run the ceph orch apply FILESYSTEM_NAME command to add MDS.
Addition of Ceph Object
Gateway
Run the site-container.yml playbook.
Run the ceph orch apply rgw commands to add Ceph Object Gateway.
Table 2. Day two operations
Description
Ceph-Ansible
Cephadm
Removing hosts
Use the Ansible inventory.
Run ceph orch host rm HOST_NAME to remove the hosts.
Removing monitors
Run the shrink-mon.yml playbook.
Run ceph orch apply mon to redeploy other monitors.
Removing managers
Run the shrink-mon.yml playbook.
Run ceph orch apply mgr to redeploy other managers.
Removing OSDs
Run the shrink-osd.yml playbook.
Run ceph orch osd rm OSD_ID to remove the OSDs.
Removing MDS
Run the shrink-mds.yml playbook.
Deployment of Ceph Object Gateway
Run the site-container.yml playbook.
Run ceph orch apply rgw _SERVICE_NAME_ to deploy Ceph
Object Gateway service.
Removing Ceph Object Gateway
Run the shrink-rgw.yml playbook.
Run ceph orch rm SERVICE_NAME to remove the specific
service.
Block device mirroring
Run the site-container.yml playbook.
Run ceph orch apply rbd-mirror command.
Minor version upgrade of IBM Storage
Ceph
Run the infrastructureplaybooks/rolling_update.yml playbook.
Run ceph orch upgrade start command.
Upgrading from IBM Storage Ceph 4 to
IBM Storage Ceph 6
Run infrastructureplaybooks/rolling_update.yml playbook.
Upgrade using Cephadm is not supported.
Deployment of monitoring stack
Edit the all.yml file during installation.
Run the ceph orch apply -i FILE.yml after specifying the
services.
What to do next? Day 2
As a storage administrator, once you have installed and configured IBM Storage Ceph, you are ready to perform "Day Two" operations for your storage cluster. These
operations include adding metadata servers (MDS) and object gateways (RGW), and configuring services.
For more information about how to use the cephadm orchestrator to perform "Day Two" operations, see Operations.
To deploy, configure, and administer the Ceph Object Gateway on "Day Two" operations, see Object Gateway.
Upgrading
Upgrade to an IBM Storage Ceph cluster running Red Hat Enterprise Linux on AMD64 and Intel 64 architectures.
Note: For customers currently using Red Hat Ceph Storage with iSCSI gateway who want to migrate to IBM Storage Ceph with NVMe-oF gateway support, contact IBM
Support.
Upgrading an IBM Storage Ceph cluster using cephadm
Upgrading a host operating system from RHEL 8 to RHEL 9
You can perform a IBM Storage Ceph host operating system upgrade from Red Hat Enterprise Linux (RHEL) 8 to Red Hat Enterprise Linux 9 using the Leapp utility.
Upgrading IBM Storage Ceph 5 to IBM Storage Ceph 6 involving RHEL 8 to RHEL 9 upgrades
Upgrading IBM Storage Ceph 5 to 6 involving RHEL 8 to RHEL 9 upgrades with stretch mode enabled
Staggered upgrade
Monitoring and managing upgrade
Monitor and manage the upgrade of the storage cluster, after upgrading.
Upgrading an IBM Storage Ceph cluster using cephadm
As a storage administrator, you can use the cephadm Orchestrator to upgrade from IBM Storage Ceph 5.3 to 6.1 with the ceph orch
upgrade command. .
The automated upgrade process follows Ceph best practices. For example:
The upgrade order starts with Ceph Managers, Ceph Monitors, then other daemons.
Each daemon is restarted only after Ceph indicates that the cluster will remain available.
174 IBM Storage Ceph
The storage cluster health status is likely to switch to HEALTH_WARNING during the upgrade. When the upgrade is complete, the health status should switch back to
HEALTH_OK.
Warning: If you have an IBM Storage Ceph 6 cluster with multi-site configured, do not upgrade to the latest version of 6.1 as there are issues with data corruption on
encrypted objects when objects replicate to the disaster recovery (DR) site.
Note: You do not get a message once the upgrade is successful. Run ceph versions and ceph orch ps commands to verify the new image ID and the version of the storage
cluster.
Compatibility considerations between Ceph and podman versions
Upgrading the IBM Storage Ceph cluster
Crossgrading from Red Hat Ceph Storage 6.1 to IBM Storage Ceph 6.1
Upgrade from a Red Hat Ceph Storage 6.1. to IBM Storage Ceph 6.1 cluster with the ceph orch upgrade command.
Upgrading cluster in a disconnected environment
Compatibility considerations between Ceph and podman versions
podman and IBM Storage Ceph have different end-of-life strategies that might make it challenging to find compatible versions.
IBM recommends to use the podman version shipped with the corresponding Red Hat Enterprise Linux version for IBM Storage Ceph 6. See the Red Hat Ceph Storage:
Supported configurations knowledge base article for more details.
Table:Compatibility with podman
Podman
Ceph 6
1.9
false
2.0
true
2.1
true
2.2
false
3.0
true
>3.0
true
Upgrading the IBM Storage Ceph cluster
You can use ceph orch upgrade command to upgrade an IBM Storage Ceph cluster.
Prerequisites
Latest version of IBM Storage Ceph cluster 5.3.
Root-level access to all the nodes.
Ansible user with sudo and passwordless ssh access to all nodes in the storage cluster.
At least two Ceph Manager nodes in the storage cluster: one active and one standby.
Note: IBM Storage Ceph includes a health check function that returns a DAEMON_OLD_VERSION warning if it detects that any of the daemons in the storage cluster are
running multiple versions of IBM Storage Ceph. The warning is triggered when the daemons continue to run multiple versions of IBM Storage Ceph beyond the time value
set in the mon_warn_older_version_delay option. By default, the mon_warn_older_version_delay option is set to 1 week. This setting allows most upgrades to
proceed without falsely seeing the warning.
If the upgrade process is paused for an extended time period, you can mute the health warning:
ceph health mute DAEMON_OLD_VERSION --sticky
After the upgrade has finished, unmute the health warning:
ceph health unmute DAEMON_OLD_VERSION
Procedure
1. Register the system, and when prompted, enter your Red Hat customer portal credentials:
Example
[root@admin ~]# subscription-manager register
2. Pull the latest subscription:
subscription-manager refresh
3. List all available subscriptions for IBM Storage Ceph:
subscription-manager list --available --matches 'IBM Storage Ceph'
4. Identify the appropriate subscription and retrieve its Pool ID.
5. Attach the pool ID to gain access to the software entitlements. Use the Pool ID you identified in the previous step.
Example
[root@admin ~]# subscription-manager attach --pool=POOL_ID
6. Disable the software repositories:
Example
[root@admin ~]# subscription-manager repos --disable=*
7. Enable the Red Hat Enterprise Linux baseos and appstream repositories:
IBM Storage Ceph 175
Example
[root@admin ~]# subscription-manager repos --enable=rhel-9-for-x86_64-baseos-rpms
[root@admin ~]# subscription-manager repos --enable=rhel-9-for-x86_64-appstream-rpms
8. Update the system:
Example
[root@admin ~]# dnf update
9. Enable the ceph-tools repository for Red Hat Enterprise Linux 9:
Example
[root@admin ~]# curl https://public.dhe.ibm.com/ibmdl/export/pub/storage/ceph/ibm-storage-ceph-6-rhel-9.repo | sudo tee
/etc/yum.repos.d/ibm-storage-ceph-6-rhel-9.repo
10. Repeat the above steps on all the nodes of the storage cluster
11. Add license to install IBM Storage Ceph and click Accept on all nodes:
Example
[root@admin ~]# dnf install ibm-storage-ceph-license
a. Accept these provisions:
Example
[root@admin ~]# sudo touch /usr/share/ibm-storage-ceph-license/accept
12. Navigate to the /usr/share/cephadm-ansible/ directory:
Example
[root@admin ~]# cd /usr/share/cephadm-ansible
13. Update the cephadm and cephadm-ansible packages.
Example
[root@admin cephadm-ansible] dnf update cephadm
[root@admin cephadm-ansible] dnf update cephadm-ansible
14. Run the preflight playbook with the upgrade_ceph_packages parameter set to true on the bootstrapped host in the storage cluster:
Syntax
ansible-playbook -i INVENTORY_FILE cephadm-preflight.yml --extra-vars "ceph_origin=ibm upgrade_ceph_packages=true"
Example
[ansible@admin cephadm-ansible]$ ansible-playbook -i /etc/ansible/hosts cephadm-preflight.yml --extra-vars
"ceph_origin=ibm upgrade_ceph_packages=true"
This package upgrades cephadm on all the nodes.
15. Log into the cephadm shell:
Example
[root@host01 ~]# cephadm shell
16. Ensure all the hosts are online and that the storage cluster is healthy:
Example
[ceph: root@host01 /]# ceph -s
17. Set the OSD noout, noscrub, and nodeep-scrubflags to prevent OSDs from getting marked out during upgrade and to avoid unnecessary load on the cluster:
Example
[ceph: root@host01 /]# ceph osd set noout
[ceph: root@host01 /]# ceph osd set noscrub
[ceph: root@host01 /]# ceph osd set nodeep-scrub
18. Login to registry and check service versions and the available target containers:
Syntax
ceph cephadm registry-login cp.icr.io USERNAME PASSWORD
or (using json file)
cat mylogin.json
{ "url":"REGISTRY_URL",
"username":"USER_NAME",
"password":"PASSWORD" }
ceph cephadm registry-login -i mylogin.json
19. Check service versions and the available target containers:
Syntax
176 IBM Storage Ceph
ceph orch upgrade check IMAGE_NAME
Example
[ceph: root@host01 /]# ceph orch upgrade check cp.icr.io/cp/ibm-ceph/ceph-6-rhel9:latest
20. Upgrade the storage cluster:
Syntax
ceph orch upgrade start IMAGE_NAME
Example
[ceph: root@host01 /]# ceph orch upgrade start cp.icr.io/cp/ibm-ceph/ceph-6-rhel9:latest
While the upgrade is underway, a progress bar appears in the ceph status output.
Example
[ceph: root@host01 /]# ceph status
[...]
progress:
Upgrade to 17.2.6-70.el9c (1s)
[............................]
21. Verify the new IMAGE_ID and VERSION of the Ceph cluster:
Example
[ceph: root@host01 /]# ceph versions
[ceph: root@host01 /]# ceph orch ps
Verify you have the latest version:
Example
[root@client01 ~] ceph --version
22. When the upgrade is complete, unset the noout, noscrub, and nodeep-scrub flags:
Example
[ceph: root@host01 /]# ceph osd unset noout
[ceph: root@host01 /]# ceph osd unset noscrub
[ceph: root@host01 /]# ceph osd unset nodeep-scrub
Reference
Troubleshooting upgrade error messages
Crossgrading from Red Hat Ceph Storage 6.1 to IBM Storage Ceph 6.1
Upgrade from a Red Hat Ceph Storage 6.1. to IBM Storage Ceph 6.1 cluster with the ceph orch upgrade command.
Before you begin
Latest version of Red Hat Ceph Storage cluster 6.1.
Root-level access to all the nodes.
Ansible user with sudo and passwordless ssh access to all nodes in the storage cluster.
At least two Ceph Manager nodes in the storage cluster, one active and one on standby.
About this task
You can use ceph orch upgrade command to upgrade to an IBM Storage Ceph 6.1 cluster.
Important: Red Hat Enterprise Linux 9 and later does not support the cephadm-ansible playbook.
Important: IBM Storage Ceph 6.1 also includes a health check function that returns a DAEMON_OLD_VERSION warning if it detects that any of the daemons in the storage
cluster are running multiple versions of IBM Storage Ceph. The warning is triggered when the daemons continue to run multiple versions of IBM Storage Ceph beyond the
time value set in the mon_warn_older_version_delay option. By default, the mon_warn_older_version_delay option is set to 1 week. This setting allows most upgrades to
proceed without falsely seeing the warning.
If the upgrade process is paused for an extended time period, you can mute the health warning:
ceph health mute DAEMON_OLD_VERSION --sticky
After the upgrade, unmute the health warning:
ceph health unmute DAEMON_OLD_VERSION
Procedure
1. Enable the Red Hat Enterprise Linux baseos and appstream repositories.
IBM Storage Ceph 177
subscription-manager repos --enable=rhel-9-for-x86_64-baseos-rpms
subscription-manager repos --enable=rhel-9-for-x86_64-appstream-rpms
For example,
[root@admin ~]# subscription-manager repos --enable=rhel-9-for-x86_64-baseos-rpms
[root@admin ~]# subscription-manager repos --enable=rhel-9-for-x86_64-appstream-rpms
Repeat this step on all the nodes of the storage cluster.
2. Disable the Red Hat Ceph Storage ceph-tools repository (/etc/yum.repos.d/). Perform this step if you are upgrading from Red Hat Ceph Storage 6.1 to IBM
Storage Ceph 6.1.
Important: This step must be performed to avoid upgrade failures of cephadm and ceph-ansible packages.
subscription-manager repos --disable=rhceph-6-tools-for-rhel-9-x86_64-rpms
For example,
[root@host01 yum.repos.d]# subscription-manager repos --disable=rhceph-6-tools-for-rhel-9-x86_64-rpms
Repeat this step on all the nodes of the storage cluster.
3. Enable the IBM Storage Ceph ceph-tools repository (/etc/yum.repos.d/) for Red Hat Enterprise Linux 9.
curl https://public.dhe.ibm.com/ibmdl/export/pub/storage/ceph/ibm-storage-ceph-6-rhel-9.repo | sudo tee
/etc/yum.repos.d/ibm-storage-ceph-6-rhel-9.repo
Repeat this step on all the nodes of the storage cluster.
4. Add license to install IBM Storage Ceph and click Accept on all nodes.
dnf -y install ibm-storage-ceph-license
5. Accept these provisions.
sudo touch /usr/share/ibm-storage-ceph-license/accept
6. Reinstall Cephadm and Cephadm Ansible.
dnf -y reinstall cephadm
dnf -y reinstall cephadm-ansible
7. Navigate to the /usr/share/cephadm-ansible/ directory.
cd /usr/share/cephadm-ansible
8. Run the preflight playbook with the upgrade_ceph_packages parameter set to true on the bootstrapped host in the storage cluster.
ansible-playbook -i INVENTORY_FILE cephadm-preflight.yml --extra-vars "ceph_origin=ibm upgrade_ceph_packages=true"
[ansible@admin cephadm-ansible]$ ansible-playbook -i /etc/ansible/hosts cephadm-preflight.yml --extra-vars
"ceph_origin=ibm upgrade_ceph_packages=true"
This package upgrades cephadm on all the nodes.
9. Log in to the cephadm shell.
cephadm shell
10. Ensure all the hosts are online and that the storage cluster is healthy.
ceph -s
11. Set the OSD noout, noscrub, and nodeep-scrub flags to prevent OSDs from getting marked out during upgrade and to avoid unnecessary load on the cluster.
ceph osd set noout
ceph osd set noscrub
ceph osd set nodeep-scrub
12. Log in to registry.
ceph cephadm registry-login cp.icr.io USERNAME PASSWORD
or log in with the json file.
cat mylogin.json
{ "url":"REGISTRY_URL",
"username":"USER_NAME",
"password":"PASSWORD" }
ceph cephadm registry-login -i mylogin.json
13. Check service versions and the available target containers.
ceph orch upgrade check IMAGE_NAME
[ceph: root@host01 /]# ceph orch upgrade check cp.icr.io/cp/ibm-ceph/ceph-6-rhel9:latest
14. Upgrade the storage cluster.
ceph orch upgrade start IMAGE_NAME
[ceph: root@host01 /]# ceph orch upgrade start cp.icr.io/cp/ibm-ceph/ceph-6-rhel9:latest
15. While the upgrade is underway, a progress bar appears in the ceph status command output.
[ceph: root@host01 /]# ceph status
[...]
progress:
178 IBM Storage Ceph
Upgrade to 17.2.6-246.el9cp (1s)
[............................]
16. Verify the new IMAGE_ID and VERSION of the Ceph cluster.
ceph versions
ceph orch ps
17. If you are not using the cephadm-ansible playbooks, after upgrading your Ceph cluster, you must upgrade the ceph-common package and client libraries on your
client nodes.
dnf -y reinstall ceph-common
18. Verify that you have the latest version.
ceph --version
19. When the upgrade is complete, unset the noout, noscrub, and nodeep-scrub flags.
ceph osd unset noout
ceph osd unset noscrub
ceph osd unset nodeep-scrub
Upgrading cluster in a disconnected environment
You can upgrade the storage cluster in a disconnected environment by using the --image tag.
Prerequisites
Latest version of IBM Storage Ceph cluster 5.3.
Root-level access to all the nodes.
Ansible user with sudo and passwordless ssh access to all nodes in the storage cluster.
At least two Ceph Manager nodes in the storage cluster: one active and one standby.
Register the nodes to CDN and attach subscriptions.
Check for the customer container images in a disconnected environment and change the configuration, if required. For more information, see Changing
configurations of custom container images for disconnected installations.
By default, the monitoring stack components are deployed based on the primary Ceph image. For disconnected environment of the storage cluster, you must use the latest
available monitoring stack component images.
Table 1. Custom image details for monitoring stack components for all
versions
Monitoring stack component
Image details
Prometheus
cp.icr.io/cp/ibm-ceph/prometheus:v4.12
Grafana
cp.icr.io/cp/ibm-ceph/ceph-6-dashboard-rhel9:latest
Node-exporter
cp.icr.io/cp/ibm-ceph/prometheus-node-exporter:v4.12
AlertManager
cp.icr.io/cp/ibm-ceph/prometheus-alertmanager:v4.12
HAProxy
cp.icr.io/cp/ibm-ceph/haproxy-rhel9:latest
Keepalived
cp.icr.io/cp/ibm-ceph/keepalived-rhel9:latest
SNMP Gateway
cp.icr.io/cp/ibm-ceph/snmp-notifier-rhel9:latest
Procedure
1. Add license to install IBM Storage Ceph and click Accept on all nodes:
Example
[root@admin ~]# dnf install ibm-storage-ceph-license
a. Accept these provisions:
Example
[root@admin ~]# sudo touch /usr/share/ibm-storage-ceph-license/accept
2. Navigate to the /usr/share/cephadm-ansible/ directory:
Example
[root@admin ~]# cd /usr/share/cephadm-ansible
3. Update the cephadm and cephadm-ansible package.
Example
[root@admin cephadm-ansible]# dnf update cephadm
[root@admin cephadm-ansible]# dnf update cephadm-ansible
IBM Storage Ceph 179
4. Run the preflight playbook with the upgrade_ceph_packages parameter set to true and the ceph_origin parameter set to custom on the bootstrapped host in the
storage cluster:
Syntax
ansible-playbook -i INVENTORY_FILE cephadm-preflight.yml --extra-vars "ceph_origin=custom upgrade_ceph_packages=true"
Example
[ansible@admin cephadm-ansible]$ ansible-playbook -i /etc/ansible/hosts cephadm-preflight.yml --extra-vars
"ceph_origin=custom upgrade_ceph_packages=true"
This package upgrades cephadm on all the nodes.
5. Log in to the cephadm shell:
Example
[root@host01 ~]# cephadm shell
6. Ensure all the hosts are online and that the storage cluster is healthy:
Example
[ceph: root@host01 /]# ceph -s
7. Check service versions and the available target containers:
Syntax
ceph orch upgrade check IMAGE_NAME
Example
[ceph: root@host01 /]# ceph orch upgrade check LOCAL_NODE_FQDN:5000/ibm-ceph/ceph-6-rhel9:latest
8. Set the OSD noout, noscrub, and nodeep-scrub flags to prevent OSDs from getting marked out during upgrade and to avoid unnecessary load on the cluster:
Example
[ceph: root@host01 /]# ceph osd set noout
[ceph: root@host01 /]# ceph osd set noscrub
[ceph: root@host01 /]# ceph osd set nodeep-scrub
9. Upgrade the storage cluster:
Syntax
ceph orch upgrade start IMAGE_NAME
Example
[ceph: root@host01 /]# ceph orch upgrade start LOCAL_NODE_FQDN:5000/ibm-ceph/ceph-6-rhel9:latest
While the upgrade is underway, a progress bar appears in the ceph status output.
Example
[ceph: root@host01 /]# ceph status
[...]
progress:
Upgrade to 17.2.6-70.el9cp (1s)
[............................]
10. Verify the new IMAGE_ID and VERSION of the Ceph cluster:
Example
[ceph: root@host01 /]# ceph version
[ceph: root@host01 /]# ceph versions
[ceph: root@host01 /]# ceph orch ps
11. When the upgrade is complete, unset the noout, noscrub, and nodeep-scrub flags:
Example
[ceph: root@host01 /]# ceph osd unset noout
[ceph: root@host01 /]# ceph osd unset noscrub
[ceph: root@host01 /]# ceph osd unset nodeep-scrub
Reference
For more information, see Registering IBM Storage Ceph nodes to the CDN and attaching subscriptions and Configuring a private registry for a disconnected installation.
Upgrading a host operating system from RHEL 8 to RHEL 9
You can perform a IBM Storage Ceph host operating system upgrade from Red Hat Enterprise Linux (RHEL) 8 to Red Hat Enterprise Linux 9 using the Leapp utility.
180 IBM Storage Ceph
About this task
Important: This host operating system upgrade must be performed before upgrading the IBM Storage Ceph cluster.
The following are the supported combinations of containerized Ceph daemons. For more information, see Colocation.
Ceph Metadata Server (ceph-mds), Ceph OSD (ceph-osd), and Ceph Object Gateway (radosgw)
Ceph Monitor (ceph-mon) or Ceph Manager (ceph-mgr), Ceph OSD (ceph-osd), and Ceph Object Gateway (radosgw)
Ceph Monitor (ceph-mon), Ceph Manager (ceph-mgr), Ceph OSD (ceph-osd), and Ceph Object Gateway (radosgw)
Before you begin
Be sure that you have a running IBM Storage Ceph 5 cluster.
Procedure
1. Deploy IBM Storage Ceph 5 on RHEL 8 with service.
Note: Verify that the cluster contains two admin nodes, so that while performing host upgrade in one admin node (with _admin label), the second admin can be
used for managing clusters.
For full instructions, see IBM Storage Ceph installation and Deploying the Ceph daemons using the service specification.
2. Set the noout flag on the Ceph OSD.
For example:
[ceph: root@host01 /]# ceph osd set noout
3. Perform host upgrade one node at a time using the Leapp utility.
a. Put respective node maintenance mode before performing host upgrade using Leapp.
ceph orch host maintenance enter HOST
For example:
ceph orch host maintenance enter host01
b. Enable the Ceph tools repo manually when executing the Leapp command with the --enablerepo parameter.
For example:
leapp upgrade --enablerepo https://public.dhe.ibm.com/ibmdl/export/pub/storage/ceph/ibm-storage-ceph-6-rhel-9.repo
c. Follow the instructions provided within the Upgrading RHEL 8 to RHEL 9 guide within the Red Hat Enterprise Linux product documentation on the Red Hat
Customer Portal.
4. Verify the new IMAGE_ID and VERSION of the Ceph cluster:
[ceph: root@node0 /]# ceph version
[ceph: root@node0 /]# ceph orch ps
What to do next
Continue with the IBM Storage Ceph 5 to IBM Storage Ceph 6 upgrade, by following Upgrading an IBM Storage Ceph cluster using cephadm.
Upgrading IBM Storage Ceph 5 to IBM Storage Ceph 6 involving RHEL 8 to RHEL 9
upgrades
Before you begin
A running IBM Storage Ceph 5 cluster on RHEL 8.
Backup of Ceph binary (/usr/sbin/cephadm), ceph.pub (/etc/ceph), and the Ceph cluster's public SSH keys from the admin node.
About this task
Upgrade from IBM Storage Ceph 5 to IBM Storage Ceph 6. The upgrade includes an upgrade from RHEL 8 to RHEL 9.
Procedure
1. Upgrade the RHEL version on all hosts of the cluster. For a detailed procedure, see Upgrading a host operating system from RHEL 8 to RHEL 9.
2. After the RHEL upgrade, register the IBM Storage Ceph nodes to the CDN and add the necessary repositories. For a detailed procedure, see Registering the Red Hat
Ceph Storage nodes to the CDN and attaching subscriptions.
3. Update the IBM Storage Ceph. For a detailed procedure, see Upgrade a Red Hat Ceph Storage cluster using cephadm.
Upgrading IBM Storage Ceph 5 to 6 involving RHEL 8 to RHEL 9 upgrades with stretch
mode enabled
IBM Storage Ceph 181
You can perform an upgrade from IBM Storage Ceph 5 to IBM Storage Ceph 6 involving Red Hat Enterprise Linux 8 to Red Hat Enterprise Linux 9 with the stretch mode
enabled.
Important: Upgrade to the latest version of IBM Storage Ceph 5.3 prior to upgrading to the latest version of IBM Storage Ceph 6.1.
Prerequisties
IBM Storage Ceph on Red Hat Enterprise Linux 8 with necessary hosts and daemons running with stretch mode enabled.
Backup of Ceph binary (/usr/sbin/cephadm), ceph.pub (/etc/ceph), and the Ceph cluster’s public SSH keys from the admin node.
Note: Arbiter monitor cannot be drained or removed from the host. Hence, the arbiter mon needs to be re-provisioned to another tie-breaker node, and then drained or
removed from host as described inReplacing tiebreaker with a new monitor.
Procedure
1. Log into the Cephadm shell:
Example
[ceph: root@host01 /]# cephadm shell
2. Label a second node as the admin in the cluster to manage the cluster when the admin node is re-provisioned.
Syntax
ceph orch host label add HOSTNAME_admin
Example
[ceph: root@host01 /]# ceph orch host label add host02_admin
3. Set the noout flag.
Example
[ceph: root@host01 /]# ceph osd set noout
4. Drain all the daemons from the host:
Syntax
ceph orch host drain HOSTNAME --force
Example
[ceph: root@host01 /]# ceph orch host drain host02 --force
The _no_schedule label is automatically applied to the host which blocks deployment.
5. Check if all the daemons are removed from the storage cluster:
Syntax
ceph orch ps HOSTNAME
Example
[ceph: root@host01 /]# ceph orch ps host02
6. Check the status of OSD removal:
Example
[ceph: root@host01 /]# ceph orch osd rm status
When no placement groups (PG) are left on the OSD, the OSD is decommissioned and removed from the storage cluster.
7. Zap the devices so that if the hosts being drained have OSDs present, then they can be used to re-deploy OSDs when the host is added back.
Syntax
ceph orch device zap HOSTNAME DISK --force
Example
[ceph: root@host01 /]# ceph orch device zap ceph-host02 /dev/vdb --force
zap successful for /dev/vdb on ceph-host02
8. Remove the host from the cluster:
Syntax
ceph orch host rm HOSTNAME [--force]
Example
[ceph: root@host01 /]# ceph orch host rm host02 [--force]
9. Re-provision the respective hosts from RHEL 8 to RHEL 9 as described in Upgrading from RHEL 8 to RHEL 9 .
182 IBM Storage Ceph
10. Navigate to the /usr/share/cephadm-ansible/ directory:
Example
[root@admin ~]# cd /usr/share/cephadm-ansible
11. Run the preflight playbook with the --limit option:
Syntax
ansible-playbook -i INVENTORY_FILE cephadm-preflight.yml --limit NEWHOST_NAME
Example
[ceph: root@host01 cephadm ansible]# ansible-playbook -i hosts cephadm-preflight.yml --extra-vars "ceph_origin={storageproduct}" --limit host02
The preflight playbook installs podman, lvm2, chronyd, and cephadm on the new host. After installation is complete, cephadm resides in the /usr/sbin/
directory.
12. Extract the cluster’s public SSH keys to a folder:
Syntax
ceph cephadm get-pub-key ~/PATH
Example
[ceph: root@host01 /]# ceph cephadm get-pub-key ~/ceph.pub
13. Copy Ceph cluster’s public SSH keys to the re-provisioned node:
Syntax
ssh-copy-id -f -i ~/PATH root@HOST_NAME_2
Example
[ceph: root@host01 /]# ssh-copy-id -f -i ~/ceph.pub root@host02
a. Optional: If the removed host has a monitor daemon, then, before adding the host to the cluster, add the --unmanaged flag to monitor deployment.
Syntax
ceph orch apply mon PLACEMENT --unmanaged
14. Add the host again to the cluster and add the labels present earlier:
Syntax
ceph orch host add HOSTNAME IP_ADDRESS --labels=LABELS
a. Syntax
ceph mon add HOSTNAME IP_LOCATION
Example
[ceph: root@host01 /]# ceph mon add ceph-host02 10.0.211.62 datacenter=DC2
Syntax
ceph orch daemon add mon HOSTNAME
Example
[ceph: root@host01 /]# ceph orch daemon add mon ceph-host02
15. Verify the daemons on the re-provisioned host running successfully with the same ceph version:
Syntax
ceph orch ps
16. Set back the monitor daemon placement to managed.
Note: This step needs to be done one by one.
Syntax
ceph orch apply mon PLACEMENT
17. Repeat the above steps for all hosts.
18. Follow the same approach to re-provision admin nodes and use a second admin node to manage clusters.
19. Add the backup files again to the node.
20. Add admin nodes again to cluster using the second admin node. Set the mon deployment to unmanaged.
21. Follow Replacing tiebreaker with a new monitor to add back the old arbiter mon and remove the temporary monitor created earlier.
22. Unset the noout flag.
Example
IBM Storage Ceph 183
ceph osd unset noout
23. Verify the Ceph version and the cluster status to ensure that all demons are working as expected after the OS upgrade.
24. Now that the RHEL OS is successfully upgraded, follow the Upgrading an IBM Storage Ceph cluster using cephadm to perform upgrade from IBM Storage Ceph 5 to
6.
Staggered upgrade
As a storage administrator, you can upgrade IBM Storage Ceph components in phases rather than all at once. The ceph orch upgrade command enables you to specify
options to limit which daemons are upgraded by a single upgrade command.
Note: If you want to upgrade from a version that does not support staggered upgrades, you must first manually upgrade the Ceph Manager (ceph-mgr) daemons.
Staggered upgrade options
Performing a staggered upgrade
Staggered upgrade options
The ceph orch upgrade command supports several options to upgrade cluster components in phases. The staggered upgrade options include:
--daemon_types: The --daemon_types option takes a comma-separated list of daemon types and will only upgrade daemons of those types. Valid daemon types
for this option include mgr, mon, crash, osd, mds, rgw, rbd-mirror, and cephfs-mirror.
--services: The --services option is mutually exclusive with --daemon-types, only takes services of one type at a time, and will only upgrade daemons
belonging to those services. For example, you cannot provide an OSD and RGW service simultaneously.
--hosts: You can combine the --hosts option with --daemon_types, --services, or use it on its own. The --hosts option parameter follows the same format
as the command line options for orchestrator CLI placement specification.
--limit: The --limit option takes an integer greater than zero and provides a numerical limit on the number of daemons cephadm will upgrade. You can combine
the --limit option with --daemon_types, --services, or --hosts. For example, if you specify to upgrade daemons of type osd on host01 with a limit set to
3, cephadm will upgrade up to three OSD daemons on host01.
Performing a staggered upgrade
As a storage administrator, you can use the ceph orch upgrade options to limit which daemons are upgraded by a single upgrade command.
Cephadm strictly enforces an order for the upgrade of daemons that is still present in staggered upgrade scenarios. The current upgrade order is:
1. Ceph Manager nodes
2. Ceph Monitor nodes
3. Ceph-crash daemons
4. Ceph OSD nodes
5. Ceph Metadata Server (MDS) nodes
6. Ceph Object Gateway (RGW) nodes
7. Ceph RBD-mirror node
8. CephFS-mirror node
Note:
If you specify parameters that upgrade daemons out of order, the upgrade command blocks and notes which daemons you need to upgrade before you proceed.
There is no required order for restarting the instances. It is recommended to restart the instance pointing to the pool with primary images followed by the instance
pointing to the mirrored pool.
Example
[ceph: root@host01 /]# ceph orch upgrade start --image cp.icr.io/cp/ibm-ceph/ceph-6-rhel9:latest --hosts host02
Error EINVAL: Cannot start upgrade. Daemons with types earlier in upgrade order than daemons on given host need upgrading.
Please first upgrade mon.ceph-host01
NOTE: Enforced upgrade order is: mgr -> mon -> crash -> osd -> mds -> rgw -> rbd-mirror -> cephfs-mirror
Prerequisites
Latest version of IBM Storage Ceph cluster 5.3.
Root-level access to all the nodes.
At least two Ceph Manager nodes in the storage cluster: one active and one standby.
Procedure
1. Log into the cephadm shell:
Example
184 IBM Storage Ceph
[root@host01 ~]# cephadm shell
2. Ensure all the hosts are online and that the storage cluster is healthy:
Example
[ceph: root@host01 /]# ceph -s
3. Set the OSD noout, noscrub, and nodeep-scrubflags to prevent OSDs from getting marked out during upgrade and to avoid unnecessary load on the cluster:
Example
[ceph: root@host01 /]# ceph osd set noout
[ceph: root@host01 /]# ceph osd set noscrub
[ceph: root@host01 /]# ceph osd set nodeep-scrub
4. Check service versions and the available target containers:
Syntax
ceph orch upgrade check IMAGE_NAME
Example
[ceph: root@host01 /]# ceph orch upgrade check cp.icr.io/cp/ibm-ceph/ceph-6-rhel9:latest
5. Upgrade the storage cluster:
a. To upgrade specific daemon types on specific hosts:
Syntax
ceph orch upgrade start --image IMAGE_NAME --daemon-types DAEMON_TYPE1,DAEMON_TYPE2 --hosts HOST1,HOST2
Example
[ceph: root@host01 /]# ceph orch upgrade start --image cp.icr.io/cp/ibm-ceph/ceph-6-rhel9:latest --daemon-types
mgr,mon --hosts host02,host03
b. To specify specific services and limit the number of daemons to upgrade:
Syntax
ceph orch upgrade start --image IMAGE_NAME --services SERVICE1,SERVICE2 --limit LIMIT_NUMBER
Example
[ceph: root@host01 /]# ceph orch upgrade start --image cp.icr.io/cp/ibm-ceph/ceph-6-rhel9:latest --services
rgw.example1,rgw1.example2 --limit 2
Note:
In staggered upgrade scenarios, if using a limiting parameter, the monitoring stack daemons, including Prometheus and node-exporter, are
refreshed after the upgrade of the Ceph Manager daemons. As a result of the limiting parameter, Ceph Manager upgrades take longer to complete. The
versions of monitoring stack daemons might not change between Ceph releases, in which case, they are only redeployed.
Upgrade commands with limiting parameters validates the options before beginning the upgrade, which can require pulling the new container image.
As a result, the upgrade start command might take a while to return when you provide limiting parameters.
6. To see which daemons you still need to upgrade, run the ceph orch upgrade check or ceph versions command:
Example
[ceph: root@host01 /]# ceph orch upgrade check --image cp.icr.io/cp/ibm-ceph/ceph-6-rhel9:latest
7. To complete the staggered upgrade, verify the upgrade of all remaining services:
Syntax
ceph orch upgrade start --image IMAGE_NAME
Example
[ceph: root@host01 /]# ceph orch upgrade start --image cp.icr.io/cp/ibm-ceph/ceph-6-rhel9:latest
8. Verify the new IMAGE_ID and VERSION of the Ceph cluster:
Example
[ceph: root@host01 /]# ceph versions
[ceph: root@host01 /]# ceph orch ps
9. When the upgrade is complete, unset the noout, noscrub, and nodeep-scrub flags:
Example
[ceph: root@host01 /]# ceph osd unset noout
[ceph: root@host01 /]# ceph osd unset noscrub
[ceph: root@host01 /]# ceph osd unset nodeep-scrub
Reference
For more information about performing a staggered upgrade and staggered upgrade options, see Performing a staggered upgrade.
IBM Storage Ceph 185
Monitoring and managing upgrade
Monitor and manage the upgrade of the storage cluster, after upgrading.
After running the ceph orch upgrade start command to upgrade the IBM Storage Ceph cluster, you can check the status, pause, resume, or stop the upgrade process. The
health of the cluster changes to HEALTH_WARNING during an upgrade. If the host of the cluster is offline, the upgrade is paused.
Note: You have to upgrade one daemon type after the other. If a daemon cannot be upgraded, the upgrade is paused.
Prerequisites
A running IBM Storage Ceph 5.3.
Root-level access to all the nodes.
At least two Ceph Manager nodes in the storage cluster: one active and one standby.
Upgrade for the storage cluster initiated.
Procedure
1. Determine whether an upgrade is in process and the version to which the cluster is upgrading:
Example
[ceph: root@node0 /]# ceph orch upgrade status
Note: You do not get a message once the upgrade is successful. Run ceph versions and ceph orch ps commands to verify the new image ID and the version of the
storage cluster.
2. Optional: Pause the upgrade process:
Example
[ceph: root@node0 /]# ceph orch upgrade pause
3. Optional: Resume a paused upgrade process:
Example
[ceph: root@node0 /]# ceph orch upgrade resume
4. Optional: Stop the upgrade process:
Example
[ceph: root@node0 /]# ceph orch upgrade stop
Configuring
Use this information to help configure IBM Storage Ceph at start and run time.
As a storage administrator, you need to have a basic understanding of how to view the Ceph configuration, and how to set the Ceph configuration options for the IBM
Storage Ceph cluster. You can view and set the Ceph configuration options at runtime.
As part of the Ceph authentication configuration, consider key rotation for your Ceph and gateway daemons for increased security. Key rotation can be done either through
the command-line or the Ceph dashboard. For more information see Creating user capabilities and Enabling key rotation.
Be sure that the IBM Storage Ceph software is installed before configuring.
Ceph configuration
IBM Storage Ceph clusters typically have a configuration file that is created and defined through a deployment tool. Configuration files can also be created
manually.
Ceph network configuration
It is important to understand the network environment that the IBM Storage Ceph cluster will operate in, in order to configure the IBM Storage Ceph.
Ceph Monitor configuration
Use the default configuration values for the Ceph Monitor or customize them according to the intended workload.
Ceph authentication configuration
Authenticating users and services is important to the security of the IBM Storage Ceph cluster. IBM Storage Ceph includes the Cephx protocol, as the default, for
cryptographic authentication, and the tools to manage authentication in the storage cluster.
Pools, placement groups, and CRUSH configuration
Ceph Object Storage Daemon (OSD) configuration
Configure the Ceph Object Storage Daemon (OSD) to be redundant and optimized based on the intended workload.
Ceph Monitor and OSD interaction configuration
Be sure to properly configure the interactions between the Ceph Monitors and OSDs to ensure a stable working environment.
Debugging and logging configuration
Increase the amount of debugging and logging information in cephadm to help diagnose problems with IBM Storage Ceph.
General configuration options
Understand the general configuration options for Ceph.
Network configuration options
Understand the various network configuration options for Ceph.
186 IBM Storage Ceph
Ceph firewall ports
Understand the various firewall ports used in IBM Storage Ceph.
Ceph Monitor configuration options
Understand the various Ceph monitor configuration options that can be set up during deployment.
Cephx configuration options
Understand the various Cephx configuration options that can be set up during deployment.
Pools, placement groups, and CRUSH configuration options
Understand the various Ceph options that govern pools, placement groups, and the CRUSH algorithm.
Object Storage Daemon (OSD) configuration options
Understand the various Ceph Object Storage Daemon (OSD) configuration options that can be set during deployment.
Ceph Monitor and OSD configuration options
Understand the various Ceph Monitor and OSD configuration options.
Debugging and logging configuration options
Logging and debugging settings are not required in a Ceph configuration file, but you can override default settings as needed.
Scrubbing options
Ceph ensures data integrity by scrubbing placement groups.
BlueStore configuration options
Ceph BlueStore configuration options can be configured during deployment.
Ceph configuration
IBM Storage Ceph clusters typically have a configuration file that is created and defined through a deployment tool. Configuration files can also be created manually.
Deployment tools, such as cephadm, typically create an initial Ceph configuration file for you. However, you can manually create one, if you prefer to bootstrap an IBM
Storage Ceph cluster without using a deployment tool.
IBM Storage Ceph clusters have a configuration, which defines:
Cluster Identity
Authentication settings
Ceph daemons
Network configuration
Node names and addresses
Paths to keyrings
Paths to OSD log files
Other runtime options
For more information about cephadm and the Ceph orchestrator, see Operations.
Configuration database
The Ceph Monitor manages a configuration database of Ceph options that centralize configuration management by storing configuration options for the entire
storage cluster. By centralizing the Ceph configuration in a database, this simplifies storage cluster administration.
Using the Ceph metavariables
Viewing the Ceph configuration at runtime
Viewing a specific configuration at runtime
Setting a specific configuration at runtime
OSD Memory Target
Automatically tuning OSD memory
MDS Memory Cache Limit
MDS servers keep their metadata in a separate storage pool, named cephfs_metadata and are the users of Ceph OSDs.
Configuration database
The Ceph Monitor manages a configuration database of Ceph options that centralize configuration management by storing configuration options for the entire storage
cluster. By centralizing the Ceph configuration in a database, this simplifies storage cluster administration.
The priority order that Ceph uses to set options is:
1. Compiled-in default values
2. Ceph cluster configuration database
3. Local ceph.conf file
4. Runtime override, using the ceph daemon DAEMON-NAME config set or ceph tell DAEMON-NAME injectargs commands
There are still a few Ceph options that can be defined in the local Ceph configuration file, which is /etc/ceph/ceph.conf by default.
cephadm uses a basic ceph.conf file that only contains a minimal set of options for connecting to Ceph Monitors, authenticating, and fetching configuration information. In
most cases, cephadm uses only the mon_host option. To avoid using ceph.conf only for the mon_host option, use DNS SRV records to perform operations with Monitors.
Important: Use the assimilate-conf administrative command to move valid options into the configuration database from the ceph.conf file. For more information about
assimilate-conf, see Administrative commands.
IBM Storage Ceph 187
Ceph allows you to make changes to the configuration of a daemon at runtime. This capability can be useful for increasing or decreasing the logging output, by enabling or
disabling debug settings, and can even be used for runtime optimization.
Note: When the same option exists in the configuration database and the Ceph configuration file, the configuration database option has a lower priority than what is set in
the Ceph configuration file.
Sections and masks
Just as you can configure Ceph options globally, per daemon type, or by a specific daemon in the Ceph configuration file, you can also configure the Ceph options in the
configuration database according to these sections. These sections are described in Table 1.
Table 1. Configuration database sections
Section
global
Affects all daemons and clients.
Description
mon
Affects all Ceph Monitors.
mgr
Affects all Ceph Managers.
osd
Affects all Ceph OSDs.
mds
Affects all Ceph Metadata Servers.
client
Affects all Ceph Clients, including mounted file systems, block devices, and RADOS Gateways.
Ceph configuration options can have a mask associated with them. These masks can further restrict which daemons or clients the options apply to.
Masks have two forms:
type:location
The type is a CRUSH property, for example, rack or host. The location is a value for the property type. For example, host:foo limits the option only to
daemons or clients running on the foo host.
For example:
ceph config set osd/host:magna045 debug_osd 20
class:device-class
The device-class is the name of the CRUSH device class, such as hdd or ssd. For example, class:ssd limits the option only to Ceph OSDs backed by solid
state drives (SSD). This mask has no effect on non-OSD daemons of clients.
For example:
ceph config set osd/class:hdd osd_max_backfills 8
Administrative commands
The Ceph configuration database can be administered with the subcommand ceph config ACTION. These are the actions you can do:
ls
Lists the available configuration options.
dump
Dumps the entire configuration database of options for the storage cluster.
get WHO
Dumps the configuration for a specific daemon or client. For example, WHO can be a daemon, like mds.a.
set WHO OPTION VALUE
Sets a configuration option in the Ceph configuration database, where WHO is the target daemon, OPTION is the option to set, and VALUE is the desired value.
show WHO
Shows the reported running configuration for a running daemon. These options might be different from those stored by the Ceph Monitors if there is a local
configuration file in use or options have been overridden on the command line or at run time. Also, the source of the option values is reported as part of the output.
assimilate-conf -i INPUT_FILE -o OUTPUT_FILE
Assimilate a configuration file from the INPUT_FILE and move any valid options into the Ceph Monitors’ configuration database. Any options that are unrecognized,
invalid, or cannot be controlled by the Ceph Monitor return in an abbreviated configuration file stored in the OUTPUT_FILE. This command can be useful for
transitioning from legacy configuration files to a centralized configuration database. Note that when you assimilate a configuration and the Monitors or other
daemons have different configuration values set for the same set of options, the end result depends on the order in which the files are assimilated.
help OPTION -f json-pretty
Displays help for a particular OPTION using a JSON-formatted output.
For more information, see Setting a specific configuration at runtime.
Using the Ceph metavariables
Metavariables simplify Ceph storage cluster configuration dramatically. When a metavariable is set in a configuration value, Ceph expands the metavariable into a concrete
value.
Metavariables are very powerful when used within the [global], [osd], [mon], or [client] sections of the Ceph configuration file. However, you can also use them
with the administration socket. Ceph metavariables are similar to Bash shell expansion.
Ceph supports the following metavariables:
$cluster
188 IBM Storage Ceph
Description
Expands to the Ceph storage cluster name. Useful when running multiple Ceph storage clusters on the same hardware.
Example
/etc/ceph/$cluster.keyring
Default ceph
$type
Description
Expands to one of osd or mon, depending on the type of instant daemon.
Example
/var/lib/ceph/$type
$id
Description
Expands to the daemon identifier. For osd.0, this would be 0.
Example
/var/lib/ceph/$type/$cluster-$id
$host
Description
Expands to the host name of the instant daemon.
$name
Description
Expands to $type.$id.
Example
/var/run/ceph/$cluster-$name.asok
Viewing the Ceph configuration at runtime
The Ceph configuration files can be viewed at boot time and run time.
Prerequisites
Root-level access to the Ceph node.
Access to admin keyring.
Procedure
1. To view a runtime configuration, log in to a Ceph node running the daemon and execute:
Syntax
ceph daemon DAEMON_TYPE.ID config show
To see the configuration for osd.0, you can log into the node containing osd.0 and execute this command:
Example
[root@osd ~]# ceph daemon osd.0 config show
2. For additional options, specify a daemon and help.
Example
[root@osd ~]# ceph daemon osd.0 help
Viewing a specific configuration at runtime
Configuration settings for IBM Storage Ceph can be viewed at runtime from the Ceph Monitor node.
Prerequisites
A running IBM Storage Ceph cluster.
Root-level access to the Ceph Monitor node.
Procedure
1. Log into a Ceph node and execute:
Syntax
ceph daemon DAEMON_TYPE.ID config get PARAMETER
Example
[root@mon ~]# ceph daemon osd.0 config get public_addr
IBM Storage Ceph 189
Setting a specific configuration at runtime
To set a specific Ceph configuration at runtime, use the ceph config set command.
Prerequisites
A running IBM Storage Ceph cluster.
Root-level access to the Ceph Monitor or OSD nodes.
Procedure
1. Set the configuration on all Monitor or OSD daemons :
Syntax
ceph config set DAEMON CONFIG-OPTION VALUE
Example
[root@mon ~]# ceph config set osd debug_osd 10
2. Validate that the option and value are set:
Example
[root@mon ~]# ceph config dump
osd
advanced debug_osd 10/10
To remove the configuration option from all daemons:
Syntax
ceph config rm DAEMON CONFIG-OPTION VALUE
Example
[root@mon ~]# ceph config rm osd debug_osd
To set the configuration for a specific daemon:
Syntax
ceph config set DAEMON.DAEMON-NUMBER CONFIG-OPTION VALUE
Example
[root@mon ~]# ceph config set osd.0 debug_osd 10
To validate that the configuration is set for the specified daemon:
Example
[root@mon ~]# ceph config dump
osd.0
advanced debug_osd
10/10
To remove the configuration for a specific daemon:
Syntax
ceph config rm DAEMON.DAEMON-NUMBER CONFIG-OPTION
Example
[root@mon ~]# ceph config rm osd.0 debug_osd
Note: If you use a client that does not support reading options from the configuration database, or if you still need to use ceph.conf to change your cluster configuration for
other reasons, run the following command:
ceph config set mgr mgr/cephadm/manage_etc_ceph_ceph_conf false
Be sure to maintain and distribute the ceph.conf file across the storage cluster.
OSD Memory Target
BlueStore keeps OSD heap memory usage under a designated target size with the osd_memory_target configuration option.
The option osd_memory_target sets OSD memory based upon the available RAM in the system. Use this option when TCMalloc is configured as the memory allocator,
and when the bluestore_cache_autotune option in BlueStore is set to true.
Ceph OSD memory caching is more important when the block device is slow; for example, traditional hard drives, because the benefit of a cache hit is much higher than it
would be with a solid state drive. However, this must be weighed into a decision to collocate OSDs with other services, such as in a hyper-converged infrastructure (HCI) or
other applications.
Setting the OSD memory target
Use the osd_memory_target option to set the maximum memory threshold for all OSDs in the storage cluster, or for specific OSDs.
190 IBM Storage Ceph
Setting the OSD memory target
Use the osd_memory_target option to set the maximum memory threshold for all OSDs in the storage cluster, or for specific OSDs.
An OSD with an osd_memory_target option set to 16 GB might use up to 16 GB of memory.
Note: Configuration options for individual OSDs take precedence over the settings for all OSDs.
Prerequisites
A running IBM Storage Ceph cluster.
Root-level access to all hosts in the storage cluster.
Procedure
1. To set osd_memory_target for all OSDs in the storage cluster:
Syntax
ceph config set osd osd_memory_target VALUE
VALUE is the number of GBytes of memory to be allocated to each OSD in the storage cluster.
2. To set osd_memory_target for a specific OSD in the storage cluster:
Syntax
ceph config set osd.id osd_memory_target VALUE
.id is the ID of the OSD and VALUE is the number of GB of memory to be allocated to the specified OSD. For example, to configure the OSD with ID 8 to use up to
16 GBytes of memory:
Example
[ceph: root@host01 /]# ceph config set osd.8 osd_memory_target 16G
3. To set an individual OSD to use one maximum amount of memory and configure the rest of the OSDs to use another amount, specify the individual OSD first:
Example
[ceph: root@host01 /]# ceph config set osd osd_memory_target 16G
[ceph: root@host01 /]# ceph config set osd.8 osd_memory_target 8G
Reference
To configure IBM Storage Ceph to autotune OSD memory usage, see Automatically tuning OSD memory.
Automatically tuning OSD memory
The OSD daemons adjust the memory consumption based on the osd_memory_target configuration option. The option osd_memory_target sets OSD memory based
upon the available RAM in the system.
If IBM Storage Ceph is deployed on dedicated nodes that do not share memory with other services, cephadm automatically adjusts the per-OSD consumption based on
the total amount of RAM and the number of deployed OSDs.
Important: By default, the osd_memory_target_autotune parameter is set to true in IBM Storage Ceph cluster.
ceph config set osd osd_memory_target_autotune true
Cephadm starts with a fraction mgr/cephadm/autotune_memory_target_ratio, which defaults to 0.7 of the total RAM in the system, subtract off any memory
consumed by non-autotuned daemons such as non-OSDS and for OSDs for which osd_memory_target_autotune is false, and then divide by the remaining OSDs.
The osd_memory_target parameter is calculated as follows:
osd_memory_target = TOTAL_RAM_OF_THE_OSD * (1048576) * (autotune_memory_target_ratio) / NUMBER_OF_OSDS_IN_THE_OSD_NODE - (SPACE_AL
SPACE_ALLOCATED_FOR_OTHER_DAEMONS can optionally include the following daemon space allocations:
Alertmanager: 1 GB
Grafana: 1 GB
Ceph Manager: 4 GB
Ceph Monitor: 2 GB
Node-exporter: 1 GB
Prometheus: 1 GB
For example, if a node has 24 OSDs and has 251 GB RAM space, then osd_memory_target is 7860684936.
The final targets are reflected in the configuration database with options. You can view the limits and the current memory consumed by each daemon from the ceph
orch ps output under MEM LIMIT column.
Note:
IBM Storage Ceph 191
In IBM Storage Ceph 5.1, the default setting of osd_memory_target_autotune true is unsuitable for hyperconverged infrastructures where compute and Ceph
storage services are colocated. In a hyperconverged infrastructure, the autotune_memory_target_ratio can be set to 0.2 to reduce the memory consumption of
Ceph.
[ceph: root@host01 /]# ceph config set mgr mgr/cephadm/autotune_memory_target_ratio 0.2
You can manually set a specific memory target for an OSD in the storage cluster.
[ceph: root@host01 /]# ceph config set osd.123 osd_memory_target 7860684936
You can manually set a specific memory target for an OSD host in the storage cluster.
ceph config set osd/host:HOSTNAME osd_memory_target TARGET_BYTES
[ceph: root@host01 /]# ceph config set osd/host:host01 osd_memory_target 1000000000
Note:
Enabling osd_memory_target_autotune overwrites existing manual OSD memory target settings. To prevent daemon memory from being tuned even when the
osd_memory_target_autotune option or other similar options are enabled, set the _no_autotune_memory label on the host.
ceph orch host label add HOSTNAME _no_autotune_memory
You can exclude an OSD from memory autotuning by disabling the autotune option and setting a specific memory target.
[ceph: root@host01 /]# ceph config set osd.123 osd_memory_target_autotune false
[ceph: root@host01 /]# ceph config set osd.123 osd_memory_target 16G
MDS Memory Cache Limit
MDS servers keep their metadata in a separate storage pool, named cephfs_metadata and are the users of Ceph OSDs.
For Ceph File Systems, MDS servers support an entire IBM Storage Ceph cluster, not just a single storage device within the storage cluster. As a result, their memory
requirements can be significant. This is the case particularly if the workload consists of small-to-medium-size files, where the ratio of metadata to data is higher.
For example, to set the mds_cache_memory_limit to 2000000000 bytes, run:
ceph_conf_overrides:
mds:
mds_cache_memory_limit=2000000000
Note: For a large IBM Storage Ceph cluster with a metadata-intensive workload, do not put an MDS server on the same node as other memory-intensive services. Doing so
gives you the option to allocate more memory to MDS, for example, sizes greater than 100 GB.
For more information, see Metadata Server cache size limits.
Ceph network configuration
It is important to understand the network environment that the IBM Storage Ceph cluster will operate in, in order to configure the IBM Storage Ceph.
Understanding and configuring the Ceph network options ensures optimal performance and reliability of the overall storage cluster.
Be sure to have active network connectivity and an IBM Storage Ceph cluster installed before configuring your Ceph network.
For more information, see Network configuration options and Ceph on-wire encryption.
Network configuration for Ceph
The Ceph storage cluster does not perform request routing or dispatching on behalf of the Ceph client. Instead, Ceph clients make requests directly to Ceph OSD
daemons. Ceph OSDs perform data replication on behalf of Ceph clients, which means replication and other factors impose additional loads on the networks of
Ceph storage clusters.
Network messenger
Messenger is the Ceph network layer implementation. Both simple and async messenger types are supported.
Configuring a public network
While Ceph functions well with only a public network, you can establish more specific criteria, including multiple IP networks, for your public network.
Configuring a private network
Configuring multiple public networks to the cluster
Verifying firewall rules are configured for default Ceph ports
Verify that the host’s firewall allows connection on the default TCP ports.
Firewall settings for Ceph Monitor node
Firewall settings for Ceph OSDs
Network configuration for Ceph
The Ceph storage cluster does not perform request routing or dispatching on behalf of the Ceph client. Instead, Ceph clients make requests directly to Ceph OSD
daemons. Ceph OSDs perform data replication on behalf of Ceph clients, which means replication and other factors impose additional loads on the networks of Ceph
storage clusters.
Ceph has one network configuration requirement that applies to all daemons. The Ceph configuration file must specify the host for each daemon.
Some deployment utilities, such as cephadm creates a configuration file for you. Do not set these values if the deployment utility does it for you.
192 IBM Storage Ceph
Important:
The host option is the short name of the node, not its FQDN. It is not an IP address.
All Ceph clusters must use a public network. However, unless you specify an internal cluster network, Ceph assumes a single public network. Ceph can function
with a public network only, but for large storage clusters, you will see significant performance improvement with a second private network for carrying only clusterrelated traffic.
It is recommended to run a Ceph storage cluster with two networks. One public network and one private network.
To support two networks, each Ceph Node needs to have more than one network interface card (NIC).
Figure 1. Network architecture
Consider operating two separate networks for better performance and security.
Performance: Ceph OSDs handle data replication for the Ceph clients. When Ceph OSDs replicate data more than once, the network load between Ceph OSDs
easily dwarfs the network load between Ceph clients and the Ceph storage cluster. This can introduce latency and create a performance problem. Recovery and
rebalancing can also introduce significant latency on the public network.
Security: While most people are generally civil, some actors will engage in what is known as a Denial of Service (DoS) attack. When traffic between Ceph OSDs gets
disrupted, peering may fail and placement groups may no longer reflect an active + clean state, which may prevent users from reading and writing data. A great
way to defeat this type of attack is to maintain a completely separate cluster network that does not connect directly to the internet.
Network configuration settings are not required. Ceph can function with a public network only, assuming a public network is configured on all hosts running a Ceph
daemon. However, Ceph allows you to establish much more specific criteria, including multiple IP networks and subnet masks for your public network. You can also
establish a separate cluster network to handle OSD heartbeat, object replication, and recovery traffic.
Do not confuse the IP addresses you set in the configuration with the public-facing IP addresses network clients might use to access your service. Typical internal IP
networks are often 192.168.0.0 or 10.0.0.0.
Important: If you specify more than one IP address and subnet mask for either the public or the private network, the subnets within the network must be capable of
routing to each other. Additionally, make sure you include each IP address and subnet in your IP tables and open ports for them as necessary.
Note: Ceph uses CIDR notation for subnets, for example, 10.0.0.0/24.
When you configured the networks, you can restart the cluster or restart each daemon. Ceph daemons bind dynamically, so you do not have to restart the entire cluster at
once if you change the network configuration.
For common option descriptions and usage information, see Network configuration options.
Network messenger
IBM Storage Ceph 193
Messenger is the Ceph network layer implementation. Both simple and async messenger types are supported.
The default messenger type is async. To change the messenger type, specify the ms_type configuration setting in the [global] section of the Ceph configuration file.
Note: For the async messenger, IBM supports the posix transport type, but does not currently support rdma or dpdk. By default, the ms_type setting in IBM Storage
Ceph 5.3 or higher reflects async+posix, where async is the messenger type and posix is the transport type.
SimpleMessenger
The SimpleMessenger implementation uses TCP sockets with two threads per socket. Ceph associates each logical session with a connection. A pipe handles the
connection, including the input and output of each message. While SimpleMessenger is effective for the posix transport type, it is not effective for other
transport types such as rdma or dpdk.
AsyncMessenger
Consequently, AsyncMessenger is the default messenger type for IBM Storage Ceph 5.3 or higher. For IBM Storage Ceph 5.3 or higher, the AsyncMessenger
implementation uses TCP sockets with a fixed-size thread pool for connections, which should be equal to the highest number of replicas or erasure-code chunks.
The thread count can be set to a lower value if performance degrades due to a low CPU count or a high number of OSDs per server.
Note: Other transport types, such as rdma and dpdk are not currently supported.
For more information about using on-wire encryption with the Ceph messenger version 2 protocol, see Ceph on-wire encryption.
For more information about asynchronous messenger options, see Network configuration options.
Configuring a public network
While Ceph functions well with only a public network, you can establish more specific criteria, including multiple IP networks, for your public network.
To configure Ceph networks, use the config set command within the cephadm shell. Note that the IP addresses you set in your network configuration are different from
the public-facing IP addresses that network clients might use to access your service.
Ceph functions perfectly well with only a public network. However, Ceph allows you to establish much more specific criteria, including multiple IP networks for your public
network.
You can also establish a separate, private cluster network to handle OSD heartbeat, object replication, and recovery traffic. For more information about the private
network, see Configuring a private network.
Note:
Ceph uses CIDR notation for subnets, for example, 10.0.0.0/24. Typical internal IP networks are often 192.168.0.0/24 or 10.0.0.0/24.
If you specify more than one IP address for either the public or the cluster network, the subnets within the network must be capable of routing to each other. In
addition, make sure you include each IP address in your IP tables, and open ports for them as necessary.
The public network configuration allows you specifically define IP addresses and subnets for the public network.
Prerequisites
Installation of the IBM Storage Ceph software.
Procedure
1. Log in to the cephadm shell:
Example
[root@host01 ~]# cephadm shell
2. Configure the public network with the subnet:
Syntax
ceph config set mon public_network IP_ADDRESS_WITH_SUBNET
Example
[ceph: root@host01 /]# ceph config set mon public_network 192.168.0.0/24
3. Get the list of services in the storage cluster:
Example
[ceph: root@host01 /]# ceph orch ls
4. Restart the daemons. Ceph daemons bind dynamically, so you do not have to restart the entire cluster at once if you change the network configuration for a specific
daemon.
Example
[ceph: root@host01 /]# ceph orch restart mon
5. Optional: If you want to restart the cluster, on the admin node as a root user, run systemctl command:
Syntax
systemctl restart ceph-FSID_OF_CLUSTER.target
Example
[root@host01 ~]# systemctl restart ceph-1ca9f6a8-d036-11ec-8263-fa163ee967ad.target
Reference
194 IBM Storage Ceph
For common option descriptions and usage information, see Network configuration options.
Configuring a private network
Network configuration settings are not required. Ceph assumes a public network with all hosts operating on it, unless you specifically configure a cluster network, also
known as a private network.
If you create a cluster network, OSDs routes heartbeat, object replication, and recovery traffic over the cluster network. This can improve performance, compared to using
a single network.
Important: For added security, the cluster network should not be reachable from the public network or the Internet.
To assign a cluster network, use the --cluster-network option with the cephadm bootstrap command. The cluster network that you specify must define a subnet in
CIDR notation (for example, 10.90.90.0/24 or fe80::/64).
You can also configure the cluster_network after boostrap.
Prerequisites
Access to the Ceph software repository.
Root-level access to all nodes in the storage cluster.
Procedure
1. Run the cephadm bootstrap command from the initial node that you want to use as the Monitor node in the storage cluster. Include the --cluster-network
option in the command.
Syntax
cephadm bootstrap --mon-ip IP-ADDRESS --registry-url registry.redhat.io --registry-username USER_NAME --registry-password
PASSWORD --cluster-network NETWORK-IP-ADDRESS
Example
[root@host01 ~]# cephadm bootstrap --mon-ip 10.10.128.68 --registry-url registry.redhat.io --registry-username myuser1 -registry-password mypassword1 --cluster-network 10.10.0.0/24
2. To configure the cluster_network after bootstrap, run the config
set command and redeploy the daemons:
a. Log in to the cephadm shell:
Example
[root@host01 ~]# cephadm shell
b. Configure the cluster network with the subnet:
Syntax
ceph config set global cluster_network IP_ADDRESS_WITH_SUBNET
Example
[ceph: root@host01 /]# ceph config set global cluster_network 10.10.0.0/24
c. Get the list of services in the storage cluster:
Example
[ceph: root@host01 /]# ceph orch ls
d. Restart the daemons. Ceph daemons bind dynamically, so you do not have to restart the entire cluster at once if you change the network configuration for a
specific daemon.
Example
[ceph: root@host01 /]# ceph orch restart mon
e. Optional: If you want to restart the cluster, on the admin node as a root user, run systemctl command:
Syntax
systemctl restart ceph-FSID_OF_CLUSTER.target
Example
[root@host01 ~]# systemctl restart ceph-1ca9f6a8-d036-11ec-8263-fa163ee967ad.target
Reference
For more information about invoking cephadm bootstrap, see Bootstrapping a new storage cluster.
Configuring multiple public networks to the cluster
When the user wants to place the Ceph Monitor daemons on hosts belonging to multiple network subnets, configuring multiple public networks to the cluster is necessary.
IBM Storage Ceph 195
An example of usage is a stretch cluster mode used for Advanced Cluster Management (ACM) in Metro DR for OpenShift Data Foundation.
You can configure multiple public networks to the cluster during bootstrap and once bootstrap is complete.
Prerequisites
Before adding a host be sure that you have a running IBM Storage Ceph cluster.
Procedure
1. Bootstrap a Ceph cluster configured with multiple public networks.
a. Prepare a ceph.conf file containing a mon public network section.
Important: At least one of the provided public networks must be configured on the current host used for bootstrap.
Syntax
[mon]
public_network = PUBLIC_NETWORK1, PUBLIC_NETWORK2
Example
[mon]
public_network = 10.40.0.0/24, 10.41.0.0/24, 10.42.0.0/24
This is an example with three public networks to be provided for bootstrap.
b. Bootstrap the cluster by providing the ceph.conf file as input.
Note: During the bootstrap you can include any other arguments that you want to provide.
Syntax
cephadm --image IMAGE_URL bootstrap --mon-ip MONITOR_IP -c PATH_TO_CEPH_CONF
Note: Alternatively, an IMAGE_ID (such as, 13ea90216d0be03003d12d7869f72ad9de5cec9e54a27fd308e01e467c0d4a0a) can be used instead of
IMAGE_URL.
Example
[root@host01 ~]# cephadm –image cp.icr.io/cp/ibm-ceph/ceph-6-rhel9:latest bootstrap –mon-ip 10.40.0.0/24 -c
/etc/ceph/ceph.conf
2. Add new hosts to the subnets.
Note: The host being added must be reachable from the host that the active manager is running on.
a. Install the cluster’s public SSH key in the new host’s root user’s authorized_keys file:
Syntax
# ssh-copy-id -f -i /etc/ceph/ceph.pub root@NEW-HOST
Example
[root@host01 ~]# ssh-copy-id -f -i /etc/ceph/ceph.pub root@host02
[root@host01 ~]# ssh-copy-id -f -i /etc/ceph/ceph.pub root@host03
b. Add the new node to the Ceph cluster:
Syntax
ceph orch host add NEW_HOST IP [LABEL1 ...]
Example
ceph orch host add host02 10.10.0.102 label1
ceph orch host add host03 10.10.0.103 label2
Note:
It is best to explicitly provide the host IP address. If an IP is not provided, then the host name will be immediately resolved via DNS and that IP will be
used.
One or more labels can also be included to immediately label the new host. For example, by default the _admin label will make cephadm maintain a
copy of the ceph.conf file and a client.admin keyring file in /etc/ceph.
3. Add the networks configurations for the public network parameters to a running cluster. Be sure that the subnets are separated by commas and that the subnets
are listed in subnet/mask format.
Syntax
ceph config set mon public_network "SUBNET_1,SUBNET_2, ..."
Example
[root@host01 ~]# ceph config set mon public_network "192.168.0.0/24, 10.42.0.0/24, ..."
If necessary, update the mon specifications to place the mon daemons on hosts within the specified subnets.
Reference
For more information about adding hosts, see Adding hosts.
For more information about stretch clusters, see Stretch clusters for Ceph storage.
196 IBM Storage Ceph
Verifying firewall rules are configured for default Ceph ports
Verify that the host’s firewall allows connection on the default TCP ports.
By default, IBM Storage Ceph daemons use TCP ports 6800—7100 to communicate with other hosts in the cluster.
Note: If your network has a dedicated firewall, you might need to verify its configuration in addition to following this procedure. See the firewall’s documentation for more
information.
See the firewall’s documentation for more information.
Prerequisites
Root-level access to the host.
Procedure
1. Verify the host’s iptables configuration:
a. List active rules:
[root@host1 ~]# iptables -L
b. Verify the absence of rules that restrict connectivity on TCP ports 6800—7100.
Example
REJECT all -- anywhere anywhere reject-with icmp-host-prohibited
2. Verify the host’s firewalld configuration:
a. List ports open on the host:
Syntax
firewall-cmd --zone ZONE --list-ports
Example
[root@host1 ~]# firewall-cmd --zone default --list-ports
b. Verify the range is inclusive of TCP ports 6800—7100.
Firewall settings for Ceph Monitor node
You can enable encryption for all Ceph traffic over the network with the introduction of the messenger version 2 protocol. The secure mode setting for messenger v2
encrypts communication between Ceph daemons and Ceph clients, giving you end-to-end encryption.
Messenger v2 Protocol
The second version of Ceph’s on-wire protocol, msgr2, includes several new features:
A secure mode encrypts all data moving through the network.
Encapsulation improvement of authentication payloads.
Improvements to feature advertisement and negotiation.
The Ceph daemons bind to multiple ports allowing both the legacy, v1-compatible, and the new, v2-compatible, Ceph clients to connect to the same storage cluster. Ceph
clients or other Ceph daemons connecting to the Ceph Monitor daemon will try to use the v2 protocol first, if possible, but if not, then the legacy v1 protocol will be used.
By default, both messenger protocols, v1 and v2, are enabled. The new v2 port is 3300, and the legacy v1 port is 6789, by default.
Prerequisites
A running IBM Storage Ceph cluster.
Access to the Ceph software repository.
Root-level access to the Ceph Monitor node.
Procedure
1. Add rules using the following example:
[root@mon ]# sudo iptables -A INPUT -i IFACE -p tcp -s IP-ADDRESS/NETMASK --dport 6789 -j ACCEPT
[root@mon ]# sudo iptables -A INPUT -i IFACE -p tcp -s IP-ADDRESS/NETMASK --dport 3300 -j ACCEPT
a. Replace IFACE with the public network interface (for example, eth0, eth1, and so on).
b. Replace IP-ADDRESS with the IP address of the public network and NETMASK with the netmask for the public network.
2. For the firewalld daemon, execute the following commands:
[root@mon ~]# firewall-cmd --zone=public --add-port=6789/tcp
[root@mon ~]# firewall-cmd --zone=public --add-port=6789/tcp --permanent
[root@mon ~]# firewall-cmd --zone=public --add-port=3300/tcp
[root@mon ~]# firewall-cmd --zone=public --add-port=3300/tcp --permanent
IBM Storage Ceph 197
Firewall settings for Ceph OSDs
By default, Ceph OSDs bind to the first available ports on a Ceph node beginning at port 6800. Ensure to open at least four ports beginning at port 6800 for each OSD that
runs on the node:
One for talking to clients and monitors on the public network.
One for sending data to other OSDs on the cluster network.
Two for sending heartbeat packets on the cluster network.
Figure 1. OSD firewall
Ports are node-specific. However, you might need to open more ports than the number of ports needed by Ceph daemons running on that Ceph node in the event that
processes get restarted and the bound ports do not get released. Consider opening a few additional ports in case a daemon fails and restarts without releasing the port
such that the restarted daemon binds to a new port. Also, consider opening the port range of 6800—7300 on each OSD node.
If you set separate public and cluster networks, you must add rules for both the public network and the cluster network, because clients will connect using the public
network and other Ceph OSD Daemons will connect using the cluster network.
Prerequisites
A running IBM Storage Ceph cluster.
Access to the Ceph software repository.
Root-level access to the Ceph OSD nodes.
Procedure
1. Add rules using the following example:
[root@mon ~]# sudo iptables -A INPUT -i IFACE
-m multiport -p tcp -s IP-ADDRESS/NETMASK --dports 6800:7300 -j ACCEPT
a. Replace IFACE with the public network interface (for example, eth0, eth1, and so on).
b. Replace IP-ADDRESS with the IP address of the public network and NETMASK with the netmask for the public network.
2. For the firewalld daemon, execute the following:
[root@mon ~] # firewall-cmd --zone=public --add-port=6800-7300/tcp
[root@mon ~] # firewall-cmd --zone=public --add-port=6800-7300/tcp --permanent
If you put the cluster network into another zone, open the ports within that zone as appropriate.
Ceph Monitor configuration
Use the default configuration values for the Ceph Monitor or customize them according to the intended workload.
Be sure to have a IBM Storage Ceph cluster installed before customizing values for Ceph Monitor.
For more information, see Ceph Monitor configuration options.
Ceph Monitor configuration
Cluster maps
Ceph Monitor quorum
Ceph Monitor consistency
Bootstrap the Ceph Monitor
Minimum configuration for a Ceph Monitor
The bare minimum monitor settings for a Ceph Monitor in the Ceph configuration file includes a host name for each monitor if it is not configured for DNS and the
monitor address. The Ceph Monitors run on port 6789 and 3300 by default.
Unique identifier for Ceph
Each IBM Storage Ceph cluster has a unique identifier (fsid).
Ceph Monitor data store
Ceph provides a default path where Ceph monitors store data.
198 IBM Storage Ceph
Storage capacity
Monitor your storage capacity to protect against data loss.
Ceph heartbeat
Ceph monitors know about the cluster by requiring reports from each OSD, and by receiving reports from OSDs about the status of their neighboring OSDs.
Ceph Monitor synchronization role
Time synchronization
Ceph Monitor configuration
All storage clusters have at least one monitor. A Ceph Monitor configuration usually remains fairly consistent, but you can add, remove or replace a Ceph Monitor in a
storage cluster.
Ceph monitors maintain a master copy of the cluster map. That means a Ceph client can determine the location of all Ceph monitors and Ceph OSDs just by connecting to
one Ceph monitor and retrieving a current cluster map.
Before Ceph clients can read from or write to Ceph OSDs, they must connect to a Ceph Monitor first. With a current copy of the cluster map and the CRUSH algorithm, a
Ceph client can compute the location for any object. The abilityto compute object locations allows a Ceph client to talk directly to Ceph OSDs, which is a very important
aspect of Ceph’s high scalability and performance.
The primary role of the Ceph Monitor is to maintain a master copy of the cluster map. Ceph Monitors also provide authentication and logging services. Ceph Monitors write
all changes in the monitor services to a single Paxos instance, and Paxos writes the changes to a key-value store for strong consistency. Ceph Monitors can query the most
recent version of the cluster map during synchronization operations. Ceph Monitors leverage the key-value store’s snapshots and iterators, using the rocksdb database,
to perform store-wide synchronization.
Figure 1. Paxos
Viewing the Ceph Monitor configuration database
Viewing the Ceph Monitor configuration database
You can view Ceph Monitor configuration in the configuration database.
Prerequisites
A running IBM Storage Ceph cluster.
Root-level access to a Ceph Monitor host.
Procedure
1. Log into the cephadm shell.
[root@host01 ~]# cephadm shell
2. Use the ceph config command to view the configuration database:
Example
[ceph: root@host01 /]# ceph config get mon
For more information about the options available for the ceph config command, use ceph config -h.
IBM Storage Ceph 199
Cluster maps
The cluster map is a composite of maps, including the monitor map, the OSD map, and the placement group map. The cluster map tracks a number of important events:
Which processes are in the IBM Storage Ceph cluster.
Which processes that are in the IBM Storage Ceph cluster are up and running or down.
Whether, the placement groups are active or inactive, and clean or in some other state.
Other details that reflect the current state of the cluster, such as:
the total amount of storage space or
the amount of storage used.
When there is a significant change in the state of the cluster, for example, a Ceph OSD goes down, a placement group falls into a degraded state, and so on. The cluster
map gets updated to reflect the current state of the cluster. Additionally, the Ceph monitor also maintains a history of the prior states of the cluster. The monitor map, OSD
map, and placement group map each maintain a history of their map versions. Each version is called an epoch.
When operating the IBM Storage Ceph cluster, keeping track of these states is an important part of the cluster administration.
Ceph Monitor quorum
A cluster will run sufficiently with a single monitor. However, a single monitor is a single-point-of-failure. To ensure high availability in a production Ceph storage cluster,
run Ceph with multiple monitors so that the failure of a single monitor will not cause a failure of the entire storage cluster.
When a Ceph storage cluster runs multiple Ceph Monitors for high availability, Ceph Monitors use the Paxos algorithm to establish consensus about the master cluster
map. A consensus requires a majority of monitors running to establish a quorum for consensus about the cluster map. For example, 1; 2 out of 3; 3 out of 5; 4 out of 6; and
so on.
IBM recommends running a production IBM Storage Ceph cluster with at least three Ceph Monitors to ensure high availability. When you run multiple monitors, you can
specify the initial monitors that must be members of the storage cluster to establish a quorum. This may reduce the time it takes for the storage cluster to come online.
[mon]
mon_initial_members = a,b,c
Note: Majority of the monitors in the storage cluster must be able to reach each other in to establish a quorum. You can decrease the initial number of monitors to
establish a quorum with the mon_initial_members option.
Ceph Monitor consistency
When you add monitor settings to the Ceph configuration file, you need to be aware of some of the architectural aspects of Ceph Monitors. Ceph imposes strict
consistency requirements for a Ceph Monitor when discovering another Ceph Monitor within the cluster. Whereas Ceph clients and other Ceph daemons use the Ceph
configuration file to discover monitors, monitors discover each other using the monitor map (monmap), not the Ceph configuration file.
A Ceph Monitor always refers to the local copy of the monitor map when discovering other Ceph Monitors in the IBM Storage Ceph cluster. Using the monitor map instead
of the Ceph configuration file avoids errors that could break the cluster. For example, typos in the Ceph configuration file when specifying a monitor address or port. Since
monitors use monitor maps for discovery and they share monitor maps with clients and other Ceph daemons, the monitor map provides monitors with a strict guarantee
that their consensus is valid.
Strict consistency when applying updates to the monitor maps
As with any other updates on the Ceph Monitor, changes to the monitor map always run through a distributed consensus algorithm called Paxos. The Ceph Monitors must
agree on each update to the monitor map, such as adding or removing a Ceph Monitor, to ensure that each monitor in the quorum has the same version of the monitor
map. Updates to the monitor map are incremental so that Ceph Monitors have the latest agreed-upon version and a set of previous versions.
Maintaining history
Maintaining a history enables a Ceph Monitor that has an older version of the monitor map to catch up with the current state of the IBM Storage Ceph cluster.
If Ceph Monitors discovered each other through the Ceph configuration file instead of through the monitor map, it would introduce additional risks because the Ceph
configuration files are not updated and distributed automatically. Ceph Monitors might inadvertently use an older Ceph configuration file, fail to recognize a Ceph Monitor,
fall out of a quorum, or develop a situation where Paxos is not able to determine the current state of the system accurately.
Bootstrap the Ceph Monitor
In most configuration and deployment cases, tools that deploy Ceph, such as cephadm, might help bootstrap the Ceph monitors by generating a monitor map for you.
A Ceph monitor requires a few explicit settings:
File System ID: The fsid is the unique identifier for your object store. Since you can run multiple storage clusters on the same hardware, you must specify the
unique ID of the object store when bootstrapping a monitor. Using deployment tools, such as cephadm, will generate a file system identifier, but you can also
specify the fsid manually.
200 IBM Storage Ceph
Monitor ID: A monitor ID is a unique ID assigned to each monitor within the cluster. By convention, the ID is set to the monitor’s hostname. This option can be set
using a deployment tool, using the ceph command, or in the Ceph configuration file. In the Ceph configuration file, sections are formed as follows:
Example
[mon.host1]
[mon.host2]
Keys: The monitor must have secret keys.
Reference
For more information about cephadm and the Ceph orchestrator, see Operations.
Minimum configuration for a Ceph Monitor
The bare minimum monitor settings for a Ceph Monitor in the Ceph configuration file includes a host name for each monitor if it is not configured for DNS and the monitor
address. The Ceph Monitors run on port 6789 and 3300 by default.
Important: Do not edit the Ceph configuration file.
Note: This minimum configuration for monitors assumes that a deployment tool generates the fsid and the mon. key for you.
You can use the following commands to set or read the storage cluster configuration options.
ceph config dump
Dumps the entire configuration database for the whole storage cluster.
ceph config generate-minimal-conf
Generates a minimal ceph.conf file.
ceph config get WHO
Dumps the configuration for a specific daemon or a client, as stored in the Ceph Monitor’s configuration database.
ceph config set WHO OPTION VALUE
Sets the configuration option in the Ceph Monitor’s configuration database.
ceph config show WHO
Shows the reported running configuration for a running daemon.
ceph config assimilate-conf -i INPUT_FILE -o OUTPUT_FILE
Ingests a configuration file from the input file and moves any valid options into the Ceph Monitors’ configuration database.
Here, the WHO parameter might be name of the section or a Ceph daemon, OPTION is a configuration file, and VALUE can be either true or false.
Important: When a Ceph daemon needs a config option prior to getting the option from the config store, you can set the configuration by running the following command:
ceph cephadm set-extra-ceph-conf
This command adds text to all the daemon’s ceph.conf files. It is a workaround and is NOT a recommended operation.
Unique identifier for Ceph
Each IBM Storage Ceph cluster has a unique identifier (fsid).
If specified, the unique identifier usually appears under the [global] section of the configuration file. Deployment tools usually generate the fsid and store it in the
monitor map, so the value may not appear in a configuration file. The fsid makes it possible to run daemons for multiple clusters on the same hardware.
Important: Do not set the fsid value if you use a deployment tool that does this automatically.
Ceph Monitor data store
Ceph provides a default path where Ceph monitors store data.
Important: For optimal performance in a production IBM Storage Ceph cluster, run Ceph monitors on separate drives from Ceph OSDs.
Note: A dedicated /var/lib/ceph partition should be used for the MON database with a size between 50 and 100 GB.
Ceph monitors call the fsync() function often, which can interfere with Ceph OSD workloads.
Ceph monitors store their data as key-value pairs. Using a data store prevents recovering Ceph monitors from running corrupted versions through Paxos, and it enables
multiple modification operations in one single atomic batch, among other advantages.
Important: Only change the default data location if necessary. If you modify the default location, make it uniform across Ceph monitors by setting it in the [mon] section
of the configuration file.
Storage capacity
Monitor your storage capacity to protect against data loss.
IBM Storage Ceph 201
When an IBM Storage Ceph cluster gets close to its maximum capacity (specified by the mon_osd_full_ratio parameter), Ceph prevents you from writing to or reading
from Ceph OSDs as a safety measure to prevent data loss. Therefore, letting a production IBM Storage Ceph cluster approach its full ratio is not a good practice, because it
sacrifices high availability. The default full ratio is .95, or 95% of capacity. This a very aggressive setting for a test cluster with a small number of OSDs.
Tip: When monitoring a cluster, be alert to warnings related to the nearfull ratio. This means that a failure of some OSDs could result in a temporary service disruption if
one or more OSDs fails. Consider adding more OSDs to increase storage capacity.
A common scenario for test clusters involves a system administrator removing a Ceph OSD from the IBM Storage Ceph cluster to watch the cluster re-balance. Then,
removing another Ceph OSD, and so on until the IBM Storage Ceph cluster eventually reaches the full ratio and locks up.
Important: IBM recommends a bit of capacity planning even with a test cluster. Planning enables you to gauge how much spare capacity you will need in to maintain high
availability.
Ideally, you want to plan for a series of Ceph OSD failures where the cluster can recover to an active + clean state without replacing those Ceph OSDs immediately.
You can run a cluster in an active + degraded state, but this is not ideal for normal operating conditions.
The following diagram depicts a simplistic IBM Storage Ceph cluster containing 33 Ceph Nodes with one Ceph OSD per host, each Ceph OSD Daemon reading from and
writing to a 3TB drive. So this exemplary IBM Storage Ceph cluster has a maximum actual capacity of 99TB. With a mon osd
full ratio of 0.95, if the IBM Storage Ceph cluster falls to 5 TB of remaining capacity, the cluster will not allow Ceph clients to read and write data. So, the IBM Storage
Ceph cluster's operating capacity is 95 TB, not 99 TB.
Figure 1. Storage capacity
It is normal in such a cluster for one or two OSDs to fail. A less frequent but reasonable scenario involves a rack’s router or power supply failing, which brings down
multiple OSDs simultaneously, for example, OSDs 7-12. In such a scenario, you should still strive for a cluster that can remain operational and achieve an active +
clean state, even if that means adding a few hosts with additional OSDs in short order. If your capacity utilization is too high, you might not lose data, but you could still
sacrifice data availability while resolving an outage within a failure domain if capacity utilization of the cluster exceeds the full ratio. For this reason, IBM recommends at
least some rough capacity planning.
Identify two numbers for your cluster:
the number of OSDs
the total capacity of the cluster
To determine the mean average capacity of an OSD within a cluster, divide the total capacity of the cluster by the number of OSDs in the cluster. Consider multiplying that
number by the number of OSDs you expect to fail simultaneously during normal operations (a relatively small number). Finally, multiply the capacity of the cluster by the
full ratio to arrive at a maximum operating capacity. Then, subtract the amount of data from the OSDs you expect to fail to arrive at a reasonable full ratio. Repeat the
foregoing process with a higher number of OSD failures (for example, a rack of OSDs) to arrive at a reasonable number for a near full ratio.
Ceph heartbeat
Ceph monitors know about the cluster by requiring reports from each OSD, and by receiving reports from OSDs about the status of their neighboring OSDs.
Ceph provides reasonable default settings for interaction between monitor and OSD, however, you can modify them as needed.
Ceph Monitor synchronization role
When you run a production cluster with multiple monitors which is recommended, each monitor checks to see if a neighboring monitor has a more recent version of the
cluster map. For example, a map in a neighboring monitor with one or more epoch numbers higher than the most current epoch in the map of the instant monitor.
202 IBM Storage Ceph
Periodically, one monitor in the cluster might fall behind the other monitors to the point where it must leave the quorum, synchronize to retrieve the most current
information about the cluster, and then rejoin the quorum.
Synchronization roles
For the purposes of synchronization, monitors can assume one of three roles:
Leader: The Leader is the first monitor to achieve the most recent Paxos version of the cluster map.
Provider: The Provider is a monitor that has the most recent version of the cluster map, but was not the first to achieve the most recent version.
Requester: The Requester is a monitor that has fallen behind the leader and must synchronize to retrieve the most recent information about the cluster before it
can rejoin the quorum.
These roles enable a leader to delegate synchronization duties to a provider, which prevents synchronization requests from overloading the leader and improving
performance. In the following diagram, the requester has learned that it has fallen behind the other monitors. The requester asks the leader to synchronize, and the leader
tells the requester to synchronize with a provider.
Figure 1. Monitor Synchronization
Monitor synchronization
Synchronization always occurs when a new monitor joins the cluster. During runtime operations, monitors can receive updates to the cluster map at different times. This
means the leader and provider roles may migrate from one monitor to another. If this happens while synchronizing, for example, a provider falls behind the leader, the
provider can terminate synchronization with a requester.
Once synchronization is complete, Ceph requires trimming across the cluster. Trimming requires that the placement groups are active + clean.
Time synchronization
Ceph daemons pass critical messages to each other, which must be processed before daemons reach a timeout threshold. If the clocks in Ceph monitors are not
synchronized, it can lead to a number of anomalies.
For example:
Daemons ignoring received messages such as outdated timestamps.
Timeouts triggered too soon or late when a message was not received in time.
Tip: Install NTP on the Ceph monitor hosts to ensure that the monitor cluster operates with synchronized clocks.
Clock drift may still be noticeable with NTP even though the discrepancy is not yet harmful. Ceph clock drift and clock skew warnings can get triggered even though NTP
maintains a reasonable level of synchronization. Increasing your clock drift may be tolerable under such circumstances. However, a number of factors such as workload,
network latency, configuring overrides to default timeouts, and other synchronization options can influence the level of acceptable clock drift without compromising Paxos
guarantees.
Ceph authentication configuration
Authenticating users and services is important to the security of the IBM Storage Ceph cluster. IBM Storage Ceph includes the Cephx protocol, as the default, for
cryptographic authentication, and the tools to manage authentication in the storage cluster.
Be sure to have a IBM Storage Ceph cluster installed before configuring user authentication.
IBM Storage Ceph 203
For more information, see Cephx configuration options.
Cephx authentication
The cephx protocol is enabled by default.
Enabling Cephx
Disabling Cephx
If your cluster environment is relatively safe, you can offset the computation expense of the running authentication, by disabling Cephx.
Cephx user keyrings
When you run Ceph with authentication enabled, the ceph administrative commands and Ceph clients require authentication keys to access the Ceph storage
cluster.
Cephx daemon keyrings
Administrative users or deployment tools might generate daemon keyrings in the same way as generating user keyrings. By default, Ceph stores daemons keyrings
inside their data directory. The default keyring locations, and the capabilities necessary for the daemon to function.
Cephx message signatures
Ceph provides fine-grained control so you can enable or disable signatures for service messages between the client and Ceph.
Cephx authentication
The cephx protocol is enabled by default.
Cryptographic authentication has some computational costs, though they are generally low. If the network environment connecting clients and hosts is considered safe
and you cannot afford authentication computational costs, you can disable it. When deploying a Ceph storage cluster, the deployment tool creates the client.admin
user and keyring.
Important: Use authentication. If authentication is disabled you are at risk of a man-in-the-middle attack altering client and server messages, which can lead to significant
security issues.
Enabling and disabling Cephx
Enabling Cephx requires that you have deployed keys for the Ceph Monitors and OSDs. When toggling Cephx authentication on or off, you do not have to repeat the
deployment procedures.
Enabling Cephx
When cephx is enabled, Ceph will look for the keyring in the default search path, which includes /etc/ceph/$cluster.$name.keyring. You can override this location
by adding a keyring option in the [global] section of the Ceph configuration file, but this is not recommended.
Execute the following procedures to enable cephx on a cluster with authentication disabled. If you or your deployment utility have already generated the keys, you may
skip the steps related to generating keys.
Prerequisites
A running IBM Storage Ceph cluster.
Root-level access to the Ceph Monitor node.
Procedure
1. Create a client.admin key, and save a copy of the key for your client host:
[root@mon ~]# ceph auth get-or-create client.admin mon 'allow *' osd 'allow *' -o /etc/ceph/ceph.client.admin.keyring
Warning: This will erase the contents of any existing /etc/ceph/client.admin.keyring file. Do not perform this step if a deployment tool has already done it for you.
2. Create a keyring for the monitor cluster and generate a monitor secret key:
[root@mon ~]# ceph-authtool --create-keyring /tmp/ceph.mon.keyring --gen-key -n mon. --cap mon 'allow *'
3. Copy the monitor keyring into a ceph.mon.keyring file in every monitor mon data directory. For example, to copy it to mon.a in cluster ceph, use the following:
[root@mon ~]# cp /tmp/ceph.mon.keyring /var/lib/ceph/mon/ceph-a/keyring
4. Generate a secret key for every OSD, where _ID_ is the OSD number:
ceph auth get-or-create osd.ID mon allow rwx osd allow * -o /var/lib/ceph/osd/ceph-ID/keyring
5. By default the cephx authentication protocol is enabled.
Note: If the cephx authentication protocol was disabled previously by setting the authentication options to none, then by removing the following lines under the
[global] section in the Ceph configuration file (/etc/ceph/ceph.conf) will reenable the cephx authentication protocol:
auth_cluster_required = none
auth_service_required = none
auth_client_required = none
6. Start or restart the Ceph storage cluster.
Important:
Enabling cephx requires downtime because the cluster needs to be completely restarted, or it needs to be shut down and then started while client I/O is disabled.
These flags need to be set before restarting or shutting down the storage cluster:
204 IBM Storage Ceph
[root@mon ~]# ceph osd set noout
[root@mon ~]# ceph osd set norecover
[root@mon ~]# ceph osd set norebalance
[root@mon ~]# ceph osd set nobackfill
[root@mon ~]# ceph osd set nodown
[root@mon ~]# ceph osd set pause
Once cephx is enabled and all PGs are active and clean, unset the flags:
[root@mon ~]# ceph osd unset noout
[root@mon ~]# ceph osd unset norecover
[root@mon ~]# ceph osd unset norebalance
[root@mon ~]# ceph osd unset nobackfill
[root@mon ~]# ceph osd unset nodown
[root@mon ~]# ceph osd unset pause
Disabling Cephx
If your cluster environment is relatively safe, you can offset the computation expense of the running authentication, by disabling Cephx.
The following procedure describes how to disable Cephx.
Important: IBM recommends enabling authentication.
However, it may be easier during setup or troubleshooting to temporarily disable authentication.
Prerequisites
A running IBM Storage Ceph cluster.
Root-level access to the Ceph Monitor node.
Procedure
1. Disable cephx authentication by setting the following options in the [global] section of the Ceph configuration file:
Example
auth_cluster_required = none
auth_service_required = none
auth_client_required = none
2. Start or restart the Ceph storage cluster.
Cephx user keyrings
When you run Ceph with authentication enabled, the ceph administrative commands and Ceph clients require authentication keys to access the Ceph storage cluster.
The most common way to provide these keys to the ceph administrative commands and clients is to include a Ceph keyring under the /etc/ceph/ directory. The file name
is usually ceph.client.admin.keyring or $cluster.client.admin.keyring. If you include the keyring under the /etc/ceph/ directory, you do not need to specify a keyring entry
in the Ceph configuration file.
Important: Copy the IBM Storage Ceph cluster keyring file to nodes where you will run administrative commands, because it contains the client.admin key.
To do so, run the following command:
# scp USER@HOSTNAME:/etc/ceph/ceph.client.admin.keyring /etc/ceph/ceph.client.admin.keyring
Replace USER with the user name used on the host with the client.admin key and HOSTNAME with the host name of that host.
Note: Ensure the ceph.keyring file has appropriate permissions set on the client machine.
You can specify the key itself in the Ceph configuration file using the key setting, which is not recommended, or a path to a key file using the keyfile setting.
Cephx daemon keyrings
Administrative users or deployment tools might generate daemon keyrings in the same way as generating user keyrings. By default, Ceph stores daemons keyrings inside
their data directory. The default keyring locations, and the capabilities necessary for the daemon to function.
Note: The monitor keyring contains a key but no capabilities, and is not part of the Ceph storage cluster auth database.
The daemon data directory locations default to directories of the form:
/var/lib/ceph/$type/CLUSTER-ID
For example:
/var/lib/ceph/osd/ceph-12
These locations can be overridden, but it is preferable to keep the default locations.
Cephx message signatures
Ceph provides fine-grained control so you can enable or disable signatures for service messages between the client and Ceph.
IBM Storage Ceph 205
Signatures for messages can be enabled or disabled between Ceph daemons.
Important: Ceph should authenticate all ongoing messages between the entities by using the session key set up for that initial authentication.
Note: Ceph kernel modules do not currently support signatures.
Pools, placement groups, and CRUSH configuration
As a storage administrator, you can choose to use the IBM Storage Ceph default options for pools, placement groups, and the CRUSH algorithm or customize them for the
intended workload.
Pools placement groups and CRUSH
When you create pools and set the number of placement groups for the pool, Ceph uses default values when you do not specifically override the defaults.
Pools placement groups and CRUSH
When you create pools and set the number of placement groups for the pool, Ceph uses default values when you do not specifically override the defaults.
Important: It is best to override some of the defaults. Specifically, set a pool’s replica size and override the default number of placement groups.
You can set these values when running pool commands.
By default, Ceph makes three replicas of objects. If you want to set four copies of an object as the default value, a primary copy and three replica copies, reset the default
values as shown in osd_pool_default_size. If you want to allow Ceph to write a lesser number of copies in a degraded state, set osd_pool_default_min_size to
a number less than the osd_pool_default_size value.
Example
[ceph: root@host01 /]# ceph config set global osd_pool_default_size 4 # Write an object four times.
[ceph: root@host01 /]# ceph config set global osd_pool_default_min_size 1 # Allow writing one copy in a degraded state.
Ensure you have a realistic number of placement groups. IBM recommends approximately 100 per OSD. Total number of OSDs multiplied by 100 divided by the number of
replicas, that is, osd_pool_default_size. For 10 OSDs and osd_pool_default_size = 4, we would recommend approximately (100*10)/4=250.
Example
[ceph: root@host01 /]# ceph config set global osd_pool_default_pg_num 250
[ceph: root@host01 /]# ceph config set global osd_pool_default_pgp_num 250
Ceph Object Storage Daemon (OSD) configuration
Configure the Ceph Object Storage Daemon (OSD) to be redundant and optimized based on the intended workload.
For more information, see Object Storage Daemon (OSD) configuration options.
Ceph OSD configuration
Scrubbing the OSD
In addition to making multiple copies of objects, Ceph ensures data integrity by scrubbing placement groups. Ceph scrubbing is analogous to the fsck command on
the object storage layer.
Backfilling an OSD
OSD recovery
Ceph OSD configuration
All Ceph clusters have a configuration, which defines:
Cluster identity
Authentication settings
Ceph daemon membership in the cluster
Network configuration
Host names and addresses
Paths to keyrings
Paths to OSD log files
Other runtime options
A deployment tool, such as cephadm, will typically create an initial Ceph configuration file for you. However, you can create one yourself if you prefer to bootstrap a cluster
without using a deployment tool.
For your convenience, each daemon has a series of default values. Many are set by the ceph/src/common/config_opts.h script. You can override these settings with
a Ceph configuration file or at runtime by using the monitor tell command or connecting directly to a daemon socket on a Ceph node.
Important: Changing the default paths makes it more difficult to troubleshoot Ceph later.
206 IBM Storage Ceph
Reference
For more information about cephadm and the Ceph orchestrator, see Operations.
Scrubbing the OSD
In addition to making multiple copies of objects, Ceph ensures data integrity by scrubbing placement groups. Ceph scrubbing is analogous to the fsck command on the
object storage layer.
For each placement group, Ceph generates a catalog of all objects and compares each primary object and its replicas to ensure that no objects are missing or mismatched.
Light scrubbing (daily) checks the object size and attributes. Deep scrubbing (weekly) reads the data and uses checksums to ensure data integrity.
Scrubbing is important for maintaining data integrity, but it can reduce performance. Adjust the following settings to increase or decrease scrubbing operations.
For more information, see Scrubbing options.
Backfilling an OSD
When you add Ceph OSDs to a cluster or remove them from the cluster, the CRUSH algorithm rebalances the cluster by moving placement groups to or from Ceph OSDs to
restore the balance. The process of migrating placement groups and the objects they contain can reduce the cluster operational performance considerably. To maintain
operational performance, Ceph performs this migration with the backfill process, which allows Ceph to set backfill operations to a lower priority than requests to read or
write data.
OSD recovery
When the cluster starts or when a Ceph OSD terminates unexpectedly and restarts, the OSD begins peering with other Ceph OSDs before a write operation can occur.
If a Ceph OSD crashes and comes back online, usually it will be out of sync with other Ceph OSDs containing more recent versions of objects in the placement groups.
When this happens, the Ceph OSD goes into recovery mode and seeks to get the latest copy of the data and bring its map back up to date. Depending upon how long the
Ceph OSD was down, the OSD’s objects and placement groups may be significantly out of date. Also, if a failure domain went down, for example, a rack, more than one
Ceph OSD might come back online at the same time. This can make the recovery process time consuming and resource intensive.
To maintain operational performance, Ceph performs recovery with limitations on the number of recovery requests, threads, and object chunk sizes which allows Ceph to
perform well in a degraded state.
Ceph Monitor and OSD interaction configuration
Be sure to properly configure the interactions between the Ceph Monitors and OSDs to ensure a stable working environment.
For more information, see Ceph Monitor and OSD configuration options.
Ceph Monitor and OSD interaction
OSD heartbeat
Reporting an OSD as down
Reporting a peering failure
OSD reporting status
Ceph Monitor and OSD interaction
After you have completed your initial Ceph configuration, you can deploy and run Ceph. When you execute a command such as ceph health or ceph -s, the Ceph
Monitor reports on the current state of the Ceph storage cluster. The Ceph Monitor knows about the Ceph storage cluster by requiring reports from each Ceph OSD
daemon, and by receiving reports from Ceph OSD daemons about the status of their neighboring Ceph OSD daemons. If the Ceph Monitor does not receive reports, or if it
receives reports of changes in the Ceph storage cluster, the Ceph Monitor updates the status of the Ceph cluster map.
Ceph provides reasonable default settings for Ceph Monitor and OSD interaction. However, you can override the defaults. The following sections describe how Ceph
Monitors and Ceph OSD daemons interact for the purposes of monitoring the Ceph storage cluster.
OSD heartbeat
Each Ceph OSD daemon checks the heartbeat of other Ceph OSD daemons every 6 seconds. To change the heartbeat interval, change the value at runtime:
Syntax
ceph config set osd osd_heartbeat_interval TIME_IN_SECONDS
Example
IBM Storage Ceph 207
[ceph: root@host01 /]# ceph config set osd osd_heartbeat_interval 60
If a neighboring Ceph OSD daemon does not send heartbeat packets within a 20 second grace period, the Ceph OSD daemon might consider the neighboring Ceph OSD
daemon down. It can report it back to a Ceph Monitor, which updates the Ceph cluster map. To change the grace period, set the value at runtime:
Syntax
ceph config set osd osd_heartbeat_grace TIME_IN_SECONDS
Example
[ceph: root@host01 /]# ceph config set osd osd_heartbeat_grace 30
Figure 1. Check heartbeats
Reporting an OSD as down
By default, two Ceph OSD Daemons from different hosts must report to the Ceph Monitors that another Ceph OSD Daemon is down before the Ceph Monitors
acknowledge that the reported Ceph OSD Daemon is down.
However, there is the chance that all the OSDs reporting the failure are in different hosts in a rack with a bad switch that causes connection problems between OSDs.
To avoid a "false alarm," Ceph considers the peers reporting the failure as a proxy for a "subcluster" that is similarly laggy. While this is not always the case, it may help
administrators localize the grace correction to a subset of the system that is performing poorly.
Ceph uses the mon_osd_reporter_subtree_level setting to group the peers into the "subcluster" by their common ancestor type in the CRUSH map.
By default, only two reports from a different subtree are required to report another Ceph OSD Daemon down. Administrators can change the number of reporters from
unique subtrees and the common ancestor type required to report a Ceph OSD Daemon down to a Ceph Monitor by setting the mon_osd_min_down_reporters and
mon_osd_reporter_subtree_level values at runtime:
Syntax
ceph config set mon mon_osd_min_down_reporters NUMBER
Example
[ceph: root@host01 /]# ceph config set mon mon_osd_min_down_reporters 4
Syntax
ceph config set mon mon_osd_reporter_subtree_level CRUSH_ITEM
Example
[ceph: root@host01 /]# ceph config set mon mon_osd_reporter_subtree_level host
[ceph: root@host01 /]# ceph config set mon mon_osd_reporter_subtree_level rack
[ceph: root@host01 /]# ceph config set mon mon_osd_reporter_subtree_level osd
Figure 1. Report down OSDs
208 IBM Storage Ceph
Reporting a peering failure
If a Ceph OSD daemon cannot peer with any of the Ceph OSD daemons defined in its Ceph configuration file or the cluster map, it pings a Ceph Monitor for the most recent
copy of the cluster map every 30 seconds. You can change the Ceph Monitor heartbeat interval by setting the value at runtime:
Syntax
ceph config set osd osd_mon_heartbeat_interval TIME_IN_SECONDS
Example
[ceph: root@host01 /]# ceph config set osd osd_mon_heartbeat_interval 60
Figure 1. Report peering failure
OSD reporting status
If a Ceph OSD Daemon does not report to a Ceph Monitor, the Ceph Monitor marks the Ceph OSD Daemon down after the mon_osd_report_timeout, which is 900
seconds, elapses. A Ceph OSD Daemon sends a report to a Ceph Monitor when a reportable event such as a failure, a change in placement group stats, a change in
up_thru or when it boots within 5 seconds.
You can change the Ceph OSD Daemon minimum report interval by setting the osd_mon_report_interval value at runtime:
Syntax
ceph config set osd osd_mon_report_interval TIME_IN_SECONDS
To get, set, and verify the config you can use the following example:
Example
[ceph: root@host01 /]# ceph config get osd osd_mon_report_interval
5
[ceph: root@host01 /]# ceph config set osd osd_mon_report_interval
IBM Storage Ceph 209
20
[ceph: root@host01 /]# ceph config dump | grep osd
global
advanced osd_pool_default_crush_rule
osd
basic
osd_memory_target
osd
advanced osd_mon_report_interval
-1
4294967296
20
Figure 1. Ceph configuration update
Debugging and logging configuration
Increase the amount of debugging and logging information in cephadm to help diagnose problems with IBM Storage Ceph.
Be sure to have a running IBM Storage Ceph cluster before beginning to debugging and logging information in cephadm.
For more information about troubleshooting cephadm, see Cephadm troubleshooting.
For more information about cephadm logging, see Cephadm operations.
General configuration options
Understand the general configuration options for Ceph.
Note: Typically, these configuration options are set automatically by deployment tools, such as cephadm.
fsid
Description: The file system ID. One per cluster.
Type: >UUID
Required: No.
Default: N/A. Usually generated by deployment tools.
admin_socket
Description: The socket for executing administrative commands on a daemon, irrespective of whether Ceph monitors have established a quorum.
Type: >String
Required: No
Default: /var/run/ceph/$cluster-$name.asok
pid_file
210 IBM Storage Ceph
Description: The file in which the monitor or OSD will write its PID. For instance, /var/run/$cluster/$type.$id.pid creates /var/run/ceph/mon.a.pid for the mon with
id a running in the ceph cluster. The pid file is removed when the daemon stops gracefully. If the process is not daemonized (meaning it runs with the -f or -d
option), the pid file is not created.
Type: >String
Required: No
Default: No
chdir
Description: The directory Ceph daemons change to once they are up and running. Default / directory recommended.
Type: >String
Required: No
Default: /
max_open_files
Description: If set, when the IBM Storage Ceph cluster starts, Ceph sets the max_open_fds at the OS level (that is, the max # of file descriptors). It helps prevent
Ceph OSDs from running out of file descriptors.
Type: >64-bit Integer
Required: No
Default: 0
fatal_signal_handlers
Description: If set, we will install signal handlers for SEGV, ABRT, BUS, ILL, FPE, XCPU, XFSZ, SYS signals to generate a useful log message.
Type: >Boolean
Default: true
Network configuration options
Understand the various network configuration options for Ceph.
Common options
Host options
TCP options
Bind options
Asynchronous messenger options
Common options
These are the common network configuration options for Ceph.
public_network
Description: The IP address and netmask of the public (front-side) network (for example, 192.168.0.0/24). Set in [global]. You can specify comma-delimited
subnets.
Type: <ip-address>/<netmask> [, <ip-address>/<netmask>]
Required: No
Default: N/A
public_addr
Description: The IP address for the public (front-side) network. Set for each daemon.
Type: IP Address
Required: No
Default: N/A
cluster_network
Description: The IP address and netmask of the cluster network (for example, 10.0.0.0/24). Set in [global]. You can specify comma-delimited subnets.
Type: <ip-address>/<netmask> [, <ip-address>/<netmask>]
Required: No
Default: NA
cluster_addr
Description: The IP address for the cluster network. Set for each daemon.
Type: Address
Required: No
IBM Storage Ceph 211
Default: NA
ms_type
Description: The messenger type for the network transport layer. IBM supports the simple and the async messenger type using posix semantics.
Type: String.
Required: No.
Default: async+posix
ms_public_type
Description: The messenger type for the network transport layer of the public network. It operates identically to ms_type, but is applicable only to the public or
front-side network. This setting enables Ceph to use a different messenger type for the public or front-side and cluster or back-side networks.
Type: String.
Required: No.
Default: None.
ms_cluster_type
Description: The messenger type for the network transport layer of the cluster network. It operates identically to ms_type, but is applicable only to the cluster or
back-side network. This setting enables Ceph to use a different messenger type for the public or front-side and cluster or back-side networks.
Type: String.
Required: No.
Default: None.
Host options
You must declare at least one Ceph Monitor in the Ceph configuration file, with a mon
addr setting under each declared monitor. Ceph expects a host setting under each declared monitor, metadata server and OSD in the Ceph configuration file.
Important: Do not use localhost. Use the short name of the node, not the fully-qualified domain name (FQDN). Do not specify any value for host when using a third
party deployment system that retrieves the node name for you.
mon_addr
Description: A list of <hostname>:<port> entries that clients can use to connect to a Ceph monitor. If not set, Ceph searches [mon.*] sections.
Type: String
Required: No
Default: NA
host
Description: The host name. Use this setting for specific daemon instances (for example, [osd.0]).
Type: String
Required: Yes, for daemon instances.
Default: localhost
TCP options
Ceph disables TCP buffering by default.
ms_tcp_nodelay
Description: Ceph enables ms_tcp_nodelay so that each request is sent immediately (no buffering). Disabling Nagle’s algorithm increases network traffic, which
can introduce congestion. If you experience large numbers of small packets, you may try disabling ms_tcp_nodelay, but be aware that disabling it will generally
increase latency.
Type: Boolean
Required: No
Default: true
ms_tcp_rcvbuf
Description: The size of the socket buffer on the receiving end of a network connection. Disabled by default.
Type: 32-bit Integer
Required: No
Default: 0
ms_tcp_read_timeout
Description: If a client or daemon makes a request to another Ceph daemon and does not drop an unused connection, the tcp read timeout defines the
connection as idle after the specified number of seconds.
Type: Unsigned 64-bit Integer
Required: No
212 IBM Storage Ceph
Default: 900 15 minutes.
Bind options
The bind options configure the default port ranges for the Ceph OSD daemons. The default range is 6800:7100. You can also enable Ceph daemons to bind to IPv6
addresses.
Important: Verify that the firewall configuration allows you to use the configured port range.
ms_bind_port_min
Description: The minimum port number to which an OSD daemon will bind.
Type: 32-bit Integer
Default: 6800
Required: No
ms_bind_port_max
Description: The maximum port number to which an OSD daemon will bind.
Type: 32-bit Integer
Default: 7300
Required: No.
ms_bind_ipv6
Description: Enables Ceph daemons to bind to IPv6 addresses.
Type: Boolean
Default: false
Required: No
Asynchronous messenger options
These Ceph messenger options configure the behavior of AsyncMessenger.
ms_async_transport_type
Description: Transport type used by the AsyncMessenger. IBM supports the posix setting, but does not support the dpdk or rdma settings at this time. POSIX
uses standard TCP/IP networking and is the default value. Other transport types are experimental and are NOT supported.
Type: String
Required: No
Default: posix
ms_async_op_threads
Description: Initial number of worker threads used by each AsyncMessenger instance. This configuration setting SHOULD equal the number of replicas or erasure
code chunks, but it may be set lower if the CPU core count is low or the number of OSDs on a single server is high.
Type: 64-bit Unsigned Integer
Required: No
Default: 3
ms_async_max_op_threads
Description: The maximum number of worker threads used by each AsyncMessenger instance. Set to lower values if the OSD host has limited CPU count, and
increase if Ceph is underutilizing CPUs are underutilized.
Type: 64-bit Unsigned Integer
Required: No
Default: 5
ms_async_set_affinity
Description:
Set to true to bind AsyncMessenger workers to particular CPU cores.
Type: Boolean
Required: No
Default: true
ms_async_affinity_cores
Description: When ms_async_set_affinity is true, this string specifies how AsyncMessenger workers are bound to CPU cores. For example, 0,2 will bind
workers #1 and #2 to CPU cores #0 and #2, respectively.
Note: When manually setting affinity, make sure to not assign workers to virtual CPUs created as an effect of hyper threading or similar technology, because they
are slower than physical CPU cores.
Type: String
Required: No
IBM Storage Ceph 213
Default: (empty)
ms_async_send_inline
Description: Send messages directly from the thread that generated them instead of queuing and sending from the AsyncMessenger thread. This option is known
to decrease performance on systems with a lot of CPU cores, so it’s disabled by default.
Type: Boolean
Required: No
Default: false
Ceph firewall ports
Understand the various firewall ports used in IBM Storage Ceph.
Table 1. Firewall ports
Port
6800-7300
Port type
TCP
Ceph OSDs.
Components
3300
TCP
Ceph clients and Ceph daemons connecting to the Ceph Monitor daemon. This port is preferred over port 6789.
6789
TCP
Ceph clients and Ceph daemons connecting to the Ceph Monitor daemon. This port is considered if port 3300 fails.
Ceph Monitor configuration options
Understand the various Ceph monitor configuration options that can be set up during deployment.
You can set these configuration options with the ceph config set mon CONFIGURATION_OPTION VALUE command.
mon_initial_members
Description: The IDs of initial monitors in a cluster during startup. If specified, Ceph requires an odd number of monitors to form an initial quorum (for example, 3).
Type: String
Default: None
mon_force_quorum_join
Description: Force monitor to join quorum even if it has been previously removed from the map
Type: Boolean
Default: False
mon_dns_srv_name
Description: The service name used for querying the DNS for the monitor hosts/addresses.
Type: String
Default: ceph-mon
fsid
Description: The cluster ID. One per cluster.
Type: UUID
Required: Yes.
Default: N/A. May be generated by a deployment tool if not specified.
mon_data
Description: The monitor’s data location.
Type: String
Default: /var/lib/ceph/mon/$cluster-$id
mon_data_size_warn
Description: Ceph issues a HEALTH_WARN status in the cluster log when the monitor’s data store reaches this threshold. The default value is 15GB.
Type: Integer
Default: 15*1024*1024*1024*
mon_data_avail_warn
Description: Ceph issues a HEALTH_WARN status in the cluster log when the available disk space of the monitor’s data store is lower than or equal to this
percentage.
Type: Integer
Default: 30
mon_data_avail_crit
Description: Ceph issues a HEALTH_ERR status in the cluster log when the available disk space of the monitor’s data store is lower or equal to this percentage.
214 IBM Storage Ceph
Type: Integer
Default: 5
mon_warn_on_cache_pools_without_hit_sets
Description: Ceph issues a HEALTH_WARN status in the cluster log if a cache pool does not have the hit_set_type parameter set.
Type: Boolean
Default: True
mon_warn_on_crush_straw_calc_version_zero
Description: Ceph issues a HEALTH_WARN status in the cluster log if the CRUSH’s straw_calc_version is zero.
Type: Boolean
Default: True
mon_warn_on_legacy_crush_tunables
Description: Ceph issues a HEALTH_WARN status in the cluster log if CRUSH tunables are too old (older than mon_min_crush_required_version).
Type: Boolean
Default: True
mon_crush_min_required_version
Description: This setting defines the minimum tunable profile version required by the cluster.
Type: String
Default: hammer
mon_warn_on_osd_down_out_interval_zero
Description: Ceph issues a HEALTH_WARN status in the cluster log if the mon_osd_down_out_interval setting is zero, because the Leader behaves in a similar
manner when the noout flag is set. Administrators find it easier to troubleshoot a cluster by setting the noout flag. Ceph issues the warning to ensure
administrators know that the setting is zero.
Type: Boolean
Default: True
mon_cache_target_full_warn_ratio
Description: Ceph issues a warning when between the ratio of cache_target_full and target_max_object.
Type: Float
Default: 0.66
mon_health_data_update_interval
Description: How often (in seconds) a monitor in the quorum shares its health status with its peers. A negative number disables health updates.
Type: Float
Default: 60
mon_health_to_clog
Description: This setting enables Ceph to send a health summary to the cluster log periodically.
Type: Boolean
Default: True
mon_health_detail_to_clog
Description: This setting enable Ceph to send a health details to the cluster log periodically.
Type: Boolean
Default: True
mon_op_complaint_time
Description: Number of seconds after which the Ceph Monitor operation is considered blocked after no updates.
Type: Integer
Default: 30
mon_health_to_clog_tick_interval
Description: How often (in seconds) the monitor sends a health summary to the cluster log. A non-positive number disables it. If the current health summary is
empty or identical to the last time, the monitor will not send the status to the cluster log.
Type: Integer
Default: 60.000000
mon_health_to_clog_interval
Description: How often (in seconds) the monitor sends a health summary to the cluster log. A non-positive number disables it. The monitor will always send the
summary to the cluster log.
Type: Integer
Default: 600
IBM Storage Ceph 215
mon_osd_full_ratio
Description: The percentage of disk space used before an OSD is considered full.
Type: Float:
Default: .95
mon_osd_nearfull_ratio
Description: The percentage of disk space used before an OSD is considered nearfull.
Type: Float
Default: .85
mon_sync_trim_timeout
Description:;
Type: Double
Default: 30.0
mon_sync_heartbeat_timeout
Description:;
Type: Double
Default: 30.0
mon_sync_heartbeat_interval
Description:;
Type: Double
Default: 5.0
mon_sync_backoff_timeout
Description:;
Type: Double
Default: 30.0
mon_sync_timeout
Description: The number of seconds the monitor will wait for the next update message from its sync provider before it gives up and bootstraps again.
Type: Double
Default: 60.000000
mon_sync_max_retries
Description:;
Type: Integer
Default: 5
mon_sync_max_payload_size
Description: The maximum size for a sync payload (in bytes).
Type: 32-bit Integer
Default: 1045676
paxos_max_join_drift
Description: The maximum Paxos iterations before we must first sync the monitor data stores. When a monitor finds that its peer is too far ahead of it, it will first
sync with data stores before moving on.
Type: Integer
Default: 10
paxos_stash_full_interval
Description: How often (in commits) to stash a full copy of the PaxosService state. Currently this setting only affects mds, mon, auth and mgr PaxosServices.
Type: Integer
Default: 25
paxos_propose_interval
Description: Gather updates for this time interval before proposing a map update.
Type: Double
Default: 1.0
paxos_min
Description: The minimum number of paxos states to keep around
Type: Integer
Default: 500
216 IBM Storage Ceph
paxos_min_wait
Description: The minimum amount of time to gather updates after a period of inactivity.
Type: Double
Default: 0.05
paxos_trim_min
Description: Number of extra proposals tolerated before trimming.
Type: Integer
Default: 250
paxos_trim_max
Description: The maximum number of extra proposals to trim at a time.
Type: Integer
Default: 500
paxos_service_trim_min
Description: The minimum amount of versions to trigger a trim (0 disables it).
Type: Integer
Default: 250
paxos_service_trim_max
Description: The maximum amount of versions to trim during a single proposal (0 disables it).
Type: Integer
Default: 500
mon_max_log_epochs
Description: The maximum amount of log epochs to trim during a single proposal.
Type: Integer
Default: 500
mon_max_pgmap_epochs
Description: The maximum amount of pgmap epochs to trim during a single proposal
Type: Integer
Default: 500
mon_mds_force_trim_to
Description: Force monitor to trim mdsmaps to this point (0 disables it. dangerous, use with care)
Type: Integer
Default: 0
mon_osd_force_trim_to
Description: Force monitor to trim osdmaps to this point, even if there is PGs not clean at the specified epoch (0 disables it. dangerous, use with care)
Type: Integer
Default: 0
mon_osd_cache_size
Description: The size of osdmaps cache, not to rely on underlying store’s cache.
Type: Integer
Default: 500
mon_election_timeout
Description: On election proposer, maximum waiting time for all ACKs in seconds.
Type: Float
Default: 5
mon_lease
Description: The length (in seconds) of the lease on the monitor’s versions.
Type: Float
Default: 5
mon_lease_renew_interval_factor
Description: mon lease * mon lease renew interval
factor will be the interval for the Leader to renew the other monitor’s leases. The factor should be less than 1.0.
Type: Float
Default: 0.6
IBM Storage Ceph 217
mon_lease_ack_timeout_factor
Description: The Leader will wait mon lease * mon
lease ack timeout factor for the Providers to acknowledge the lease extension.
Type: Float
Default: 2.0
mon_accept_timeout_factor
Description: The Leader will wait mon lease * mon
accept timeout factor for the Requesters to accept a Paxos update. It is also used during the Paxos recovery phase for similar purposes.
Type: Float
Default: 2.0
mon_min_osdmap_epochs
Description: Minimum number of OSD map epochs to keep at all times.
Type: 32-bit Integer
Default: 500
mon_max_pgmap_epochs
Description: Maximum number of PG map epochs the monitor should keep.
Type: 32-bit Integer
Default: 500
mon_max_log_epochs
Description: Maximum number of Log epochs the monitor should keep.
Type: 32-bit Integer
Default: 500
clock_offset
Description: How much to offset the system clock. See Clock.cc for details.
Type: Double
Default: 0
mon_tick_interval
Description: A monitor’s tick interval in seconds.
Type: 32-bit Integer
Default: 5
mon_clock_drift_allowed
Description: The clock drift in seconds allowed between monitors.
Type: Float
Default: .050
mon_clock_drift_warn_backoff
Description: Exponential backoff for clock drift warnings.
Type: Float
Default: 5
mon_timecheck_interval
Description: The time check interval (clock drift check) in seconds for the leader.
Type: Float
Default: 300.0
mon_timecheck_skew_interval
Description: The time check interval (clock drift check) in seconds when in the presence of a skew in seconds for the Leader.
Type: Float
Default: 30.0
mon_max_osd
Description: The maximum number of OSDs allowed in the cluster.
Type: 32-bit Integer
Default: 10000
mon_globalid_prealloc
Description: The number of global IDs to pre-allocate for clients and daemons in the cluster.
Type: 32-bit Integer
Default: 10000
218 IBM Storage Ceph
mon_sync_fs_threshold
Description: Synchronize with the filesystem when writing the specified number of objects. Set it to 0 to disable it.
Type: 32-bit Integer
Default: 5
mon_subscribe_interval
Description: The refresh interval, in seconds, for subscriptions. The subscription mechanism enables obtaining the cluster maps and log information.
Type: Double
Default: 86400.000000
mon_stat_smooth_intervals
Description: Ceph will smooth statistics over the last N PG maps.
Type: Integer
Default: 6
mon_probe_timeout
Description: Number of seconds the monitor will wait to find peers before bootstrapping.
Type: Double
Default: 2.0
mon_daemon_bytes
Description: The message memory cap for metadata server and OSD messages (in bytes).
Type: 64-bit Integer Unsigned
Default: 400ul << 20
mon_max_log_entries_per_event
Description: The maximum number of log entries per event.
Type: Integer
Default: 4096
mon_osd_prime_pg_temp
Description: Enables or disable priming the PGMap with the previous OSDs when an out OSD comes back into the cluster. With the true setting, the clients will
continue to use the previous OSDs until the newly in OSDs as that PG peered.
Type: Boolean
Default: true
mon_osd_prime_pg_temp_max_time
Description: How much time in seconds the monitor should spend trying to prime the PGMap when an out OSD comes back into the cluster.
Type: Float
Default: 0.5
mon_osd_prime_pg_temp_max_time_estimate
Description: Maximum estimate of time spent on each PG before we prime all PGs in parallel.
Type: Float
Default: 0.25
mon_osd_allow_primary_affinity
Description: Allow primary_affinity to be set in the osdmap.
Type: Boolean
Default: False
mon_osd_pool_ec_fast_read
Description: Whether turn on fast read on the pool or not. It will be used as the default setting of newly created erasure pools if fast_read is not specified at
create time.
Type: Boolean
Default: False
mon_mds_skip_sanity
Description: Skip safety assertions on FSMap, in case of bugs where we want to continue anyway. Monitor terminates if the FSMap sanity check fails, but we can
disable it by enabling this option.
Type: Boolean
Default: False
mon_max_mdsmap_epochs
Description: The maximum amount of mdsmap epochs to trim during a single proposal.
Type: Integer
IBM Storage Ceph 219
Default: 500
mon_config_key_max_entry_size
Description: The maximum size of config-key entry (in bytes).
Type: Integer
Default: 65536
mon_warn_pg_not_scrubbed_ratio
Description: How often, in seconds, the monitor scrub its store by comparing the stored checksums with the computed ones of all the stored keys.
Type: float
Default: 3600*24
mon_warn_pg_not_deep_scrubbed_ratio
Description: The percentage of the scrub max interval past the scrub max interval to warn.
Type: float
Default: 0.5
mon_scrub_interval
Description: The percentage of the deep scrub interval past the deep scrub interval to warn.
Type: Integer
Default: 0.75
mon_scrub_timeout
Description: The timeout to restart scrub of mon quorum participant does not respond for the latest chunk.
Type: Integer
Default: 5 min
mon_scrub_max_keys
Description: The maximum number of keys to scrub each time.
Type: Integer
Default: 100
mon_scrub_inject_crc_mismatch
Description:The probability of injecting CRC mismatches into Ceph Monitor scrub.
Type: integer
Default: 0.000000
mon_scrub_inject_missing_keys
Description: The probability of injecting missing keys into mon scrub.
Type: float
Default: 0
mon_compact_on_start
Description: Compact the database used as Ceph Monitor store on ceph-mon start. A manual compaction helps to shrink the monitor database and improve its
performance if the regular compaction fails to work.
Type: Boolean
Default: False
mon_compact_on_bootstrap
Description: Compact the database used as Ceph Monitor store on bootstrap. The monitor starts probing each other for creating a quorum after bootstrap. If it
times out before joining the quorum, it will start over and bootstrap itself again.
Type: Boolean
Default: False
mon_compact_on_trim
Description: Compact a certain prefix (including paxos) when we trim its old states.
Type: Boolean
Default: True
mon_cpu_threads
Description: Number of threads for performing CPU intensive work on monitor.
Type: Boolean
Default: True
mon_osd_mapping_pgs_per_chunk
Description: We calculate the mapping from the placement group to OSDs in chunks. This option specifies the number of placement groups per chunk.
Type: Integer
220 IBM Storage Ceph
Default: 4096
mon_osd_max_split_count
Description: Largest number of PGs per "involved" OSD to let split create. When we increase the pg_num of a pool, the placement groups will be split on all OSDs
serving that pool. We want to avoid extreme multipliers on PG splits.
Type: Integer
Default: 300
rados_mon_op_timeout
Description: Number of seconds to wait for a response from the monitor before returning an error from a rados operation. 0 means at limit, or no wait time.
Type: Double
Default: 0
Cephx configuration options
Understand the various Cephx configuration options that can be set up during deployment.
auth_cluster_required
Description: Valid settings are cephx or none.
Type: String
Required: No
Default: cephx.
auth_service_required
Description: Valid settings are cephx or none.
Type: String
Required: No
Default: cephx.
auth_client_required
Description: If enabled, the IBM Storage Ceph cluster daemons require Ceph clients to authenticate with the IBM Storage Ceph cluster in order to access Ceph
services. Valid settings are cephx or none.
Type: String
Required: No
Default: cephx.
keyring
Description: The path to the keyring file.
Type: String
Required: No
Default: /etc/ceph/$cluster.$name.keyring, /etc/ceph/$cluster.keyring, /etc/ceph/keyring, /etc/ceph/keyring.bin
keyfile
Description: The path to a key file (that is. a file containing only the key).
Type: String
Required: No
Default: None
key
Description: The key (that is, the text string of the key itself). Not recommended.
Type: String
Required: No
Default: None
ceph-mon
Location: $mon_data/keyring
Capabilities: mon 'allow *'
ceph-osd
Location: $osd_data/keyring
Capabilities: mon 'allow profile osd' osd 'allow *'
radosgw
Location: $rgw_data/keyring
IBM Storage Ceph 221
Capabilities: mon 'allow rwx' osd 'allow rwx'
cephx_require_signatures
Description: If set to true, Ceph requires signatures on all message traffic between the Ceph client and the IBM Storage Ceph cluster, and between daemons
comprising the IBM Storage Ceph cluster.
Type: Boolean
Required: No
Default: false
cephx_cluster_require_signatures
Description: If set to true, Ceph requires signatures on all message traffic between Ceph daemons comprising the IBM Storage Ceph cluster.
Type: Boolean
Required: No
Default: false
cephx_service_require_signatures
Description: If set to true, Ceph requires signatures on all message traffic between Ceph clients and the IBM Storage Ceph cluster.
Type: Boolean
Required: No
Default: false
cephx_sign_messages
Description: If the Ceph version supports message signing, Ceph will sign all messages so they cannot be spoofed.
Type: Boolean
Default: true
auth_service_ticket_ttl
Description: When the IBM Storage Ceph cluster sends a Ceph client a ticket for authentication, the cluster assigns the ticket a time to live.
Type: Double
Default: 60*60
Pools, placement groups, and CRUSH configuration options
Understand the various Ceph options that govern pools, placement groups, and the CRUSH algorithm.
mon_allow_pool_delete
Description: Allows a monitor to delete a pool. In RHCS 3 and later releases, the monitor cannot delete the pool by default as an added measure to protect data.
Type: Boolean
Default: false
mon_max_pool_pg_num
Description: The maximum number of placement groups per pool.
Type: Integer
Default: 65536
mon_pg_create_interval
Description: Number of seconds between PG creation in the same Ceph OSD Daemon.
Type: Float
Default: 30.0
mon_pg_stuck_threshold
Description: Number of seconds after which PGs can be considered as being stuck.
Type: 32-bit Integer
Default: 300
mon_pg_min_inactive
Description: Ceph issues a HEALTH_ERR status in the cluster log if the number of PGs that remain inactive longer than the mon_pg_stuck_threshold exceeds
this setting. The default setting is one PG. A non-positive number disables this setting.
Type: Integer
Default: 1
mon_pg_warn_min_per_osd
Description: Ceph issues a HEALTH_WARN status in the cluster log if the average number of PGs per OSD in the cluster is less than this setting. A non-positive
number disables this setting.
222 IBM Storage Ceph
Type: Integer
Default: 30
mon_pg_warn_max_per_osd
Description: Ceph issues a HEALTH_WARN status in the cluster log if the average number of PGs per OSD in the cluster is greater than this setting. A non-positive
number disables this setting.
Type: Integer
Default: 300
mon_pg_warn_min_objects
Description: Do not warn if the total number of objects in the cluster is below this number.
Type: Integer
Default: 1000
mon_pg_warn_min_pool_objects
Description: Do not warn on pools whose object number is below this number.
Type: Integer
Default: 1000
mon_pg_check_down_all_threshold
Description: The threshold of down OSDs by percentage after which Ceph checks all PGs to ensure they are not stuck or stale.
Type: Float
Default: 0.5
mon_pg_warn_max_object_skew
Description: Ceph issue a HEALTH_WARN status in the cluster log if the average number of objects in a pool is greater than mon pg warn max object
skew times the average number of objects for all pools. A non-positive number disables this setting.
Type: Float
Default: 10
mon_delta_reset_interval
Description: The number of seconds of inactivity before Ceph resets the PG delta to zero. Ceph keeps track of the delta of the used space for each pool to aid
administrators in evaluating the progress of recovery and performance.
Type: Integer
Default: 10
mon_osd_max_op_age
Description: The maximum age in seconds for an operation to complete before issuing a HEALTH_WARN status.
Type: Float
Default: 32.0
osd_pg_bits
Description: Placement group bits per Ceph OSD Daemon.
Type: 32-bit Integer
Default: 6
osd_pgp_bits
Description: The number of bits per Ceph OSD Daemon for Placement Groups for Placement purpose (PGPs).
Type: 32-bit Integer
Default: 6
osd_crush_chooseleaf_type
Description: The bucket type to use for chooseleaf in a CRUSH rule. Uses ordinal rank rather than name.
Type: 32-bit Integer
Default: 1. Typically a host containing one or more Ceph OSD Daemons.
osd_pool_default_crush_replicated_ruleset
Description: The default CRUSH ruleset to use when creating a replicated pool.
Type: 8-bit Integer
Default: 0
osd_pool_erasure_code_stripe_unit
Description: Sets the default size, in bytes, of a chunk of an object stripe for erasure coded pools. Every object of size S will be stored as N stripes, with each data
chunk receiving stripe unit bytes. Each stripe of N * stripe unit bytes will be encoded/decoded individually. This option can be overridden by the
stripe_unit setting in an erasure code profile.
Type: Unsigned 32-bit Integer
IBM Storage Ceph 223
Default: 4096
osd_pool_default_size
Description: Sets the number of replicas for objects in the pool. The default value is the same as ceph osd pool set {pool-name} size {size}.
Type: 32-bit Integer
Default: 3
osd_pool_default_min_size
Description: Sets the minimum number of written replicas for objects in the pool in order to acknowledge a write operation to the client. If the minimum is not met,
Ceph will not acknowledge the write to the client. This setting ensures a minimum number of replicas when operating in degraded mode.
Type: 32-bit Integer
Default: 0, which means no particular minimum. If 0, minimum is size - (size / 2).
osd_pool_default_pg_num
Description: The default number of placement groups for a pool. The default value is the same as pg_num with mkpool.
Type: 32-bit Integer
Default: 32
osd_pool_default_pgp_num
Description: The default number of placement groups for placement for a pool. The default value is the same as pgp_num with mkpool. PG and PGP should be
equal.
Type: 32-bit Integer
Default: 0
osd_pool_default_flags
Description: The default flags for new pools.
Type: 32-bit Integer
Default: 0
osd_max_pgls
Description: The maximum number of placement groups to list. A client requesting a large number can tie up the Ceph OSD Daemon.
Type: Unsigned 64-bit Integer
Default: 1024
Note: Default value is usually sufficient.
osd_min_pg_log_entries
Description: The minimum number of placement group logs to maintain when trimming log files.
Type: 32-bit Int Unsigned
Default: 250
osd_default_data_pool_replay_window
Description: The time, in seconds, for an OSD to wait for a client to replay a request.
Type: 32-bit Integer
Default: 45
Object Storage Daemon (OSD) configuration options
Understand the various Ceph Object Storage Daemon (OSD) configuration options that can be set during deployment.
You can set these configuration options with the ceph config set osd CONFIGURATION_OPTION VALUE command.
osd_uuid
Description: The universally unique identifier (UUID) for the Ceph OSD.
Type: UUID
Default: The UUID.
Note: The osd uuid applies to a single Ceph OSD. The fsid applies to the entire cluster.
osd_data
Description: The path to the OSD’s data. You must create the directory when deploying Ceph. Mount a drive for OSD data at this mount point.
Type: String
Default: /var/lib/ceph/osd/$cluster-$id
Important: IBM does not recommend changing the default.
osd_max_write_size
Description: The maximum size of a write in megabytes.
Type: 32-bit Integer
224 IBM Storage Ceph
Default 90
osd_client_message_size_cap
Description: The largest client data message allowed in memory.
Type: 64-bit Integer Unsigned
Default: 500MB. 500*1024L*1024L
osd_class_dir
Description: The class path for RADOS class plug-ins.
Type: String
Default $libdir/rados-classes
osd_max_scrubs
Description: The maximum number of simultaneous scrub operations for a Ceph OSD.
Type: 32-bit Int
Default 1
osd_scrub_thread_timeout
Description: The maximum time in seconds before timing out a scrub thread.
Type: 32-bit Integer
Default 60
osd_scrub_finalize_thread_timeout
Description: The maximum time in seconds before timing out a scrub finalize thread.
Type: 32-bit Integer
Default 60*10
osd_scrub_begin_hour
Description: This restricts scrubbing to this hour of the day or later. Use osd_scrub_begin_hour = 0 and osd_scrub_end_hour = 0 to allow scrubbing the
entire day. Along with osd_scrub_end_hour, they define a time window, in which the scrubs can happen. But a scrub is performed no matter whether the time
window allows or not, as long as the placement group’s scrub interval exceeds osd_scrub_max_interval.
Type: Integer
Default: 0
Allowed range: [0, 23]
osd_scrub_end_hour
Description: This restricts scrubbing to the hour earlier than this. Use osd_scrub_begin_hour = 0 and osd_scrub_end_hour = 0 to allow scrubbing for the
entire day. Along with osd_scrub_begin_hour, they define a time window, in which the scrubs can happen. But a scrub is performed no matter whether the time
window allows or not, as long as the placement group's scrub interval exceeds osd_scrub_max_interval.
Type:Integer
Default: 0
Allowed range: [0, 23]
osd_scrub_load_threshold
Description: The maximum load. Ceph will not scrub when the system load (as defined by the getloadavg() function) is higher than this number. Default is 0.5.
Type: Float
Default 0.5
osd_scrub_min_interval
Description: The minimum interval in seconds for scrubbing the Ceph OSD when the IBM Storage Ceph cluster load is low.
Type: Float
Default: Once per day. 60*60*24
osd_scrub_max_interval
Description: The maximum interval in seconds for scrubbing the Ceph OSD irrespective of cluster load.
Type: Float
Default: Once per week. 7*60*60*24
osd_scrub_interval_randomize_ratio
Description: Takes the ratio and randomizes the scheduled scrub between osd scrub min interval and osd scrub max interval.
Type: Float
Default: 0.5.
mon_warn_not_scrubbed
Description: Number of seconds after osd_scrub_interval to warn about any PGs that were not scrubbed.
Type: Integer
IBM Storage Ceph 225
Default: 0 (no warning).
osd_scrub_chunk_min
Description: The object store is partitioned into chunks which end on hash boundaries. For chunky scrubs, Ceph scrubs objects one chunk at a time with writes
blocked for that chunk. The osd scrub chunk min setting represents the minimum number of chunks to scrub.
Type: 32-bit Integer
Default 5
osd_scrub_chunk_max
Description: The maximum number of chunks to scrub.
Type: 32-bit Integer
Default 25
osd_scrub_sleep
Description: The time to sleep between deep scrub operations.
Type: Float
Default: 0 (or off).
osd_scrub_during_recovery
Description: Allows scrubbing during recovery.
Type: Boolean
Default false
osd_scrub_invalid_stats
Description: Forces extra scrub to fix stats marked as invalid.
Type: Boolean
Default true
osd_scrub_priority
Description: Controls queue priority of scrub operations versus client I/O.
Type: Unsigned 32-bit Integer
Default 5
osd_requested_scrub_priority
Description: The priority set for user requested scrub on the work queue. If this value were to be smaller than osd_client_op_priority, it can be boosted to
the value of osd_client_op_priority when scrub is blocking client operations.
Type: Unsigned 32-bit Integer
Default 120
osd_scrub_cost
Description: Cost of scrub operations in megabytes for queue scheduling purposes.
Type: Unsigned 32-bit Integer
Default 52428800
osd_deep_scrub_interval
Description: The interval for deep scrubbing, that is fully reading all data. The osd scrub load threshold parameter does not affect this setting.
Type: Float
Default: Once per week. 60*60*24*7
osd_deep_scrub_stride
Description: Read size when doing a deep scrub.
Type: 32-bit Integer
Default: 512 KB. 524288
mon_warn_not_deep_scrubbed
Description: Number of seconds after osd_deep_scrub_interval to warn about any PGs that were not scrubbed.
Type: Integer
Default: 0 (no warning)
osd_deep_scrub_randomize_ratio
Description: The rate at which scrubs will randomly become deep scrubs (even before osd_deep_scrub_interval has passed).
Type: Float
Default: 0.15 or 15%
osd_deep_scrub_update_digest_min_age
Description: How many seconds old objects must be before scrub updates the whole-object digest.
Type: Integer
226 IBM Storage Ceph
Default: 7200 (120 hours)
osd_deep_scrub_large_omap_object_key_threshold
Description: Warning when you encounter an object with more OMAP keys than this.
Type: Integer
Default 200000
osd_deep_scrub_large_omap_object_value_sum_threshold
Description: Warning when you encounter an object with more OMAP key bytes than this.
Type:Integer
Default 1 G
osd_delete_sleep
Description: Time in seconds to sleep before the next removal transaction. This throttles the placement group deletion process.
Type: Float
Default 0.0
osd_delete_sleep_hdd
Description: Time in seconds to sleep before the next removal transaction for HDDs.
Type: Float
Default 5.0
osd_delete_sleep_ssd
Description: Time in seconds to sleep before the next removal transaction for SSDs.
Type: Float
Default 1.0
osd_delete_sleep_hybrid
Description: Time in seconds to sleep before the next removal transaction when Ceph OSD data is on HDD and OSD journal or WAL and DB is on SSD.
Type: Float
Default 1.0
osd_op_num_shards
Description: The number of shards for client operations.
Type: 32-bit Integer
Default 0
osd_op_num_threads_per_shard
Description: The number of threads per shard for client operations.
Type: 32-bit Integer
Default 0
osd_op_num_shards_hdd
Description: The number of shards for HDD operations.
Type: 32-bit Integer
Default 5
osd_op_num_threads_per_shard_hdd
Description: The number of threads per shard for HDD operations.
Type: 32-bit Integer
Default 1
osd_op_num_shards_ssd
Description: The number of shards for SSD operations.
Type: 32-bit Integer
Default 8
osd_op_num_threads_per_shard_ssd
Description: The number of threads per shard for SSD operations.
Type: 32-bit Integer
Default 2
osd_client_op_priority
Description: The priority set for client operations. It is relative to osd recovery op priority.
Type: 32-bit Integer
Default 63
IBM Storage Ceph 227
Valid Range: 1-63
osd_recovery_op_priority
Description: The priority set for recovery operations. It is relative to osd client op priority.
Type: 32-bit Integer
Default 3
Valid Range: 1-63
osd_op_thread_timeout
Description: The Ceph OSD operation thread timeout in seconds.
Type: 32-bit Integer
Default 15
osd_op_complaint_time
Description: An operation becomes complaint worthy after the specified number of seconds have elapsed.
Type: Float
Default 30
osd_disk_threads
Description: The number of disk threads, which are used to perform background disk intensive OSD operations such as scrubbing and snap trimming.
Type: 32-bit Integer
Default 1
osd_op_history_size
Description: The maximum number of completed operations to track.
Type: 32-bit Unsigned Integer
Default 20
osd_op_history_duration
Description: The oldest completed operation to track.
Type: 32-bit Unsigned Integer
Default 600
osd_op_log_threshold
Description: How many operations logs to display at once.
Type: 32-bit Integer
Default 5
osd_op_timeout
Description: The time in seconds after which running OSD operations time out.
Type: Integer
Default 0
Important: Do not set the osd op timeout option unless your clients can handle the consequences. For example, setting this parameter on clients running in
virtual machines can lead to data corruption because the virtual machines interpret this timeout as a hardware failure.
osd_max_backfills
Description: The maximum number of backfill operations allowed to or from a single OSD.
Type: 64-bit Unsigned Integer
Default 1
osd_backfill_scan_min
Description: The minimum number of objects per backfill scan.
Type: 32-bit Integer
Default 64
osd_backfill_scan_max
Description: The maximum number of objects per backfill scan.
Type: 32-bit Integer
Default 512
osd_backfill_full_ratio
Description: Refuse to accept backfill requests when the Ceph OSD’s full ratio is above this value.
Type: Float
Default 0.85
osd_backfill_retry_interval
Description: The number of seconds to wait before retrying backfill requests.
228 IBM Storage Ceph
Type: Double
Default 30.000000
osd_map_dedup
Description: Enable removing duplicates in the OSD map.
Type: Boolean
Default true
osd_map_cache_size
Description: The size of the OSD map cache in megabytes.
Type: 32-bit Integer
Default 50
osd_map_cache_bl_size
Description: The size of the in-memory OSD map cache in OSD daemons.
Type: 32-bit Integer
Default 50
osd_map_cache_bl_inc_size
Description: The size of the in-memory OSD map cache incrementals in OSD daemons.
Type: 32-bit Integer
Default 100
osd_map_message_max
Description: The maximum map entries allowed per MOSDMap message.
Type: 32-bit Integer
Default 40
osd_snap_trim_thread_timeout
Description: The maximum time in seconds before timing out a snap trim thread.
Type: 32-bit Integer
Default 60*60*1
osd_pg_max_concurrent_snap_trims
Description: The max number of parallel snap trims/PG. This controls how many objects per PG to trim at once.
Type: 32-bit Integer
Default 2
osd_snap_trim_sleep
Description: Insert a sleep between every trim operation a PG issues.
Type: 32-bit Integer
Default 0
osd_max_trimming_pgs
Description: The max number of trimming PGs
Type: 32-bit Integer
Default 2
osd_backlog_thread_timeout
Description: The maximum time in seconds before timing out a backlog thread.
Type: 32-bit Integer
Default 60*60*1
osd_default_notify_timeout
Description: The OSD default notification timeout (in seconds).
Type: 32-bit Integer Unsigned
Default 30
osd_check_for_log_corruption
Description: Check log files for corruption. Can be computationally expensive.
Type: Boolean
Default false
osd_remove_thread_timeout
Description: The maximum time in seconds before timing out a remove OSD thread.
Type: 32-bit Integer
IBM Storage Ceph 229
Default 60*60
osd_command_thread_timeout
Description: The maximum time in seconds before timing out a command thread.
Type: 32-bit Integer
Default 10*60
osd_command_max_records
Description: Limits the number of lost objects to return.
Type: 32-bit Integer
Default 256
osd_auto_upgrade_tmap
Description: Uses tmap for omap on old objects.
Type: Boolean
Default true
osd_tmapput_sets_users_tmap
Description: Uses tmap for debugging only.
Type: Boolean
Default false
osd_preserve_trimmed_log
Description: Preserves trimmed log files, but uses more disk space.
Type: Boolean
Default false
osd_recovery_delay_start
Description: After peering completes, Ceph delays for the specified number of seconds before starting to recover objects.
Type: Float
Default 0
osd_recovery_max_active
Description: The number of active recovery requests per OSD at one time. More requests will accelerate recovery, but the requests place an increased load on the
cluster.
Type: 32-bit Integer
Default 0
osd_recovery_max_chunk
Description: The maximum size of a recovered chunk of data to push.
Type: 64-bit Integer Unsigned
Default 8388608
osd_recovery_threads
Description: The number of threads for recovering data.
Type: 32-bit Integer
Default 1
osd_recovery_thread_timeout
Description: The maximum time in seconds before timing out a recovery thread.
Type: 32-bit Integer
Default 30
osd_recover_clone_overlap
Description: Preserves clone overlap during recovery. Should always be set to true.
Type: Boolean
Default true
rados_osd_op_timeout
Description: Number of seconds that RADOS waits for a response from the OSD before returning an error from a RADOS operation. A value of 0 means no limit.
Type: Double
Default: 0
Ceph Monitor and OSD configuration options
230 IBM Storage Ceph
Understand the various Ceph Monitor and OSD configuration options.
When modifying heartbeat settings, include them in the [global] section of the Ceph configuration file.
mon_osd_min_up_ratio
Description: The minimum ratio of up Ceph OSD Daemons before Ceph will mark Ceph OSD Daemons down.
Type: Double
Default: .3
mon_osd_min_in_ratio
Description: The minimum ratio of in Ceph OSD Daemons before Ceph will mark Ceph OSD Daemons out.
Type Double
Default: 0.750000
mon_osd_laggy_halflife
Description: The number of seconds laggy estimates will decay.
Type: Integer
Default: 60*60
mon_osd_laggy_weight
Description: The weight for new samples in laggy estimation decay.
Type: Double
Default: 0.3
mon_osd_laggy_max_interval
Description: Maximum value of laggy_interval in laggy estimations (in seconds). The monitor uses an adaptive approach to evaluate the laggy_interval of a
certain OSD. This value will be used to calculate the grace time for that OSD.
Type: Integer
Default: 300
mon_osd_adjust_heartbeat_grace
Description: If set to true, Ceph will scale based on laggy estimations.
Type: Boolean
Default: true
mon_osd_adjust_down_out_interval
Description: If set to true, Ceph will scaled based on laggy estimations.
Type: Boolean
Default: true
mon_osd_auto_mark_in
Description: Ceph will mark any booting Ceph OSD Daemons as in the Ceph Storage Cluster.
Type: Boolean
Default: false
mon_osd_auto_mark_auto_out_in
Description: Ceph will mark booting Ceph OSD Daemons auto marked out of the Ceph Storage Cluster as in the cluster.
Type: Boolean
Default: true
mon_osd_auto_mark_new_in
Description: Ceph will mark booting new Ceph OSD Daemons as in the Ceph Storage Cluster.
Type: Boolean
Default: true
mon_osd_down_out_interval
Description: The number of seconds Ceph waits before marking a Ceph OSD Daemon down and out if it does not respond.
Type: 32-bit Integer
Default: 600
mon_osd_downout_subtree_limit
Description: The largest CRUSH unit type that Ceph will automatically mark out.
Type: String
Default: rack
mon_osd_reporter_subtree_level
Description: This setting defines the parent CRUSH unit type for the reporting OSDs. The OSDs send failure reports to the monitor if they find an unresponsive peer.
The monitor may mark the reported OSD down and then out after a grace period.
IBM Storage Ceph 231
Type: String
Default: host
mon_osd_report_timeout
Description: The grace period in seconds before declaring unresponsive Ceph OSD Daemons down.
Type: 32-bit Integer
Default: 900
mon_osd_min_down_reporters
Description: The minimum number of Ceph OSD Daemons required to report a down Ceph OSD Daemon.
Type: 32-bit Integer
Default: 2
osd_heartbeat_address
Description: A Ceph OSD Daemon’s network address for heartbeats.
Type: Address
Default: The host address.
osd_heartbeat_interval
Description: How often a Ceph OSD Daemon pings its peers (in seconds).
Type: 32-bit Integer
Default: 6
osd_heartbeat_grace
Description: The elapsed time when a Ceph OSD Daemon has not shown a heartbeat that the Ceph Storage Cluster considers it down.
Type: 32-bit Integer
Default: 20
osd_mon_heartbeat_interval
Description: Frequency of Ceph OSD Daemon pinging a Ceph Monitor if it has no Ceph OSD Daemon peers.
Type: 32-bit Integer
Default: 30
osd_mon_report_interval_max
Description The maximum time in seconds that a Ceph OSD Daemon can wait before it must report to a Ceph Monitor.
Type: 32-bit Integer
Default: 120
osd_mon_report_interval_min
Description: The minimum number of seconds a Ceph OSD Daemon may wait from startup or another reportable event before reporting to a Ceph Monitor.
Type: 32-bit Integer
Default: 5
Valid Range: Should be less than osd mon report interval
max
osd_mon_ack_timeout
Description: The number of seconds to wait for a Ceph Monitor to acknowledge a request for statistics.
Type: 32-bit Integer
Default: 30
Debugging and logging configuration options
Logging and debugging settings are not required in a Ceph configuration file, but you can override default settings as needed.
The options take a single item that is assumed to be the default for all daemons regardless of channel. For example, specifying "info" is interpreted as "default=info".
However, options can also take key/value pairs. For example, default=daemon audit=local0 is interpreted as "default all to daemon, override audit with local0."
log_file
Description: The location of the logging file for the cluster.
Type: String
Required: No
Default: /var/log/ceph/$cluster-$name.log
mon_cluster_log_file
Description The location of the monitor cluster’s log file.
232 IBM Storage Ceph
Type String
Required No
Default: /var/log/ceph/$cluster.log
log_max_new
Description: The maximum number of new log files.
Type: Integer
Required: No
Default: 1000
log_max_recent
Description: The maximum number of recent events to include in a log file.
Type: Integer
Required: No
Default: 10000
log_flush_on_exit
Description: Determines if Ceph flushes the log files after exit.
Type: Boolean
Required: No
Default: true
mon_cluster_log_file_level
Description: The level of file logging for the monitor cluster. Valid settings include "debug", "info", "sec", "warn", and "error".
Type: String
Default: "info"
log_to_stderr
Description: Determines if logging messages appear in stderr.
Type: Boolean
Required: No
Default: true
err_to_stderr
Description: Determines if error messages appear in stderr.
Type: Boolean
Required: No
Default: true
log_to_syslog
Description: Determines if logging messages appear in syslog.
Type: Boolean
Required: No
Default: false
err_to_syslog
Description: Determines if error messages appear in syslog.
Type: Boolean
Required: No
Default: false
clog_to_syslog
Description: Determines if clog messages will be sent to syslog.
Type: Boolean
Required: No
Default: false
mon_cluster_log_to_syslog
Description: Determines if the cluster log will be output to syslog.
Type: Boolean
Required: No
IBM Storage Ceph 233
Default: false
mon_cluster_log_to_syslog_level
Description: The level of syslog logging for the monitor cluster. Valid settings include "debug", "info", "sec", "warn", and "error".
Type: String
Default: "info"
mon_cluster_log_to_syslog_facility
Description: The facility generating the syslog output. This is usually set to "daemon" for the Ceph daemons.
Type: String
Default: "daemon"
clog_to_monitors
Description: Determines if clog messages will be sent to monitors.
Type: Boolean
Required: No
Default: true
mon_cluster_log_to_graylog
Description: Determines if the cluster will output log messages to graylog.
Type: String
Default: "false"
mon_cluster_log_to_graylog_host
Description: The IP address of the graylog host. If the graylog host is different from the monitor host, override this setting with the appropriate IP address.
Type: String
Default: "127.0.0.1"
mon_cluster_log_to_graylog_port
Description: Graylog logs will be sent to this port. Ensure the port is open for receiving data.
Type: String
Default: "12201"
osd_preserve_trimmed_log
Description: Preserves trimmed logs after trimming.
Type: Boolean
Required: No
Default: false
osd_tmapput_sets_uses_tmap
Description: Uses tmap. For debug only.
Type: Boolean
Required: No
Default: false
osd_min_pg_log_entries
Description: The minimum number of log entries for placement groups.
Type: 32-bit Unsigned Integer
Required: No
Default: 1000
osd_op_log_threshold
Description: Number of op log messages to show up in one pass.
Type: Integer
Required: No
Default: 5
Scrubbing options
Ceph ensures data integrity by scrubbing placement groups.
The following are the Ceph scrubbing options that you can adjust to increase or decrease scrubbing operations.
234 IBM Storage Ceph
You can set these configuration options with the ceph config set global CONFIGURATION_OPTION VALUE command.
mds_max_scrub_ops_in_progress
Description: The maximum number of scrub operations performed in parallel. You can set this value with ceph config set mds_max_scrub_ops_in_progress
VALUE command.
Type: integer
Default: 5
osd_max_scrubs
Description: The maximum number of simultaneous scrub operations for a Ceph OSD Daemon.
Type: integer
Default: 1
osd_scrub_begin_hour
Description: The specific hour at which the scrubbing begins. Along with osd_scrub_end_hour, you can define a time window in which the scrubs can happen.
Use osd_scrub_begin_hour = 0 and osd_scrub_end_hour = 0 to allow scrubbing the entire day.
Type: integer
Default: 0
Allowed range: [0, 23]
osd_scrub_end_hour
Description: The specific hour at which the scrubbing ends. Along with osd_scrub_begin_hour, you can define a time window, in which the scrubs can happen.
Use osd_scrub_begin_hour = 0 and osd_scrub_end_hour = 0 to allow scrubbing for the entire day.
Type: integer
Default: 0
Allowed range: [0, 23]
osd_scrub_begin_week_day
Description: The specific day on which the scrubbing begins. 0 = Sunday, 1 = Monday, etc. Along with osd_scrub_end_week_day, you can define a time window
in which scrubs can happen. Use osd_scrub_begin_week_day = 0 and osd_scrub_end_week_day = 0 to allow scrubbing for the entire week.
Type: integer
Default: 0
Allowed range: [0, 6]
osd_scrub_end_week_day
Description: This defines the day on which the scrubbing ends. 0 = Sunday, 1 = Monday, etc. Along with osd_scrub_begin_week_day, they define a time
window, in which the scrubs can happen. Use osd_scrub_begin_week_day = 0 and osd_scrub_end_week_day = 0 to allow scrubbing for the entire week.
Type: integer
Default: 0
Allowed range: [0, 6]
osd_scrub_during_recovery
Description: Allow scrub during recovery. Setting this to false disables scheduling new scrub, and deep-scrub, while there is an active recovery. The already
running scrubs continue which is useful to reduce load on busy storage clusters.
Type: boolean
Default: false
osd_scrub_extended_sleep
Description: Duration to inject a delay during scrubbing out of scrubbing hours or seconds.
Type: float
Default: 0.0
osd_scrub_extended_sleep
Description: Backoff ratio for scheduling scrubs. This is the percentage of ticks that do NOT schedule scrubs, 66% means that 1 out of 3 ticks schedules scrubs.
Type: float
Default: 0.66
osd_scrub_load_threshold
Description: The normalized maximum load. Scrubbing does not happen when the system load, as defined by getloadavg()/number of online CPUs, is higher
than this defined number.
Type: float
Default: 0.5
osd_scrub_min_interval
Description: The minimal interval in seconds for scrubbing the Ceph OSD daemon when the Ceph storage Cluster load is low.
Type: float
IBM Storage Ceph 235
Default: 1 day
osd_scrub_max_interval
Description: The maximum interval in seconds for scrubbing the Ceph OSD daemon irrespective of cluster load.
Type: float
Default: 7 days
osd_scrub_chunk_min
Description: The minimal number of object store chunks to scrub during a single operation. Ceph blocks writes to a single chunk during scrub.
Type: integer
Default: 5
osd_scrub_chunk_max
Description: The maximum number of object store chunks to scrub during a single operation.
Type: integer
Default: 25
osd_scrub_sleep
Description: Time to sleep before scrubbing the next group of chunks. Increasing this value slows down the overall rate of scrubbing, so that client operations are
less impacted.
Type: float
Default: 0.0
osd_deep_scrub_interval
Description: The interval for deep scrubbing, fully reading all data. The osd_scrub_load_threshold does not affect this setting.
Type: float
Default: 7 days
osd_debug_deep_scrub_sleep
Description: Inject an expensive sleep during deep scrub IO to make it easier to induce preemption.
Type: float
Default: 0
osd_scrub_interval_randomize_ratio
Description: Add a random delay to osd_scrub_min_interval when scheduling the next scrub job for a placement group. The delay is a random value less than
osd_scrub_min_interval * osd_scrub_interval_randomized_ratio. The default setting spreads scrubs throughout the allowed time window of [1,
1.5] * osd_scrub_min_interval.
Type: float
Default: 0.5
osd_deep_scrub_stride
Description: Read size when doing a deep scrub.
Type: size
Default: 512 KB
osd_scrub_auto_repair_num_errors
Description: Auto repair does not occur if more than this many errors are found.
Type: integer
Default: 5
osd_scrub_auto_repair
Description: Setting this to true enables automatic Placement Group (PG) repair when errors are found by scrubs or deep-scrubs. However, if more than
osd_scrub_auto_repair_num_errors errors are found, a repair is NOT performed.
Type: boolean
Default: false
osd_scrub_max_preemptions
Description: Set the maximum number of times you need to preempt a deep scrub due to a client operation before blocking client IO to complete the scrub.
Type: integer
Default: 5
osd_scrub_auto_repair
Description: Number of keys to read from an object at a time during deep scrub.
Type: integer
Default: 1024
236 IBM Storage Ceph
BlueStore configuration options
Ceph BlueStore configuration options can be configured during deployment.
Note: This list is not complete.
rocksdb_cache_size
Description: The size of the RocksDB cache in MB.
Type: 32-bit Integer
Default: 512
Administering
Learn how to properly administer and operate IBM Storage Ceph.
Administration
Operations
Administration
Learn how to manage processes, monitor cluster states, manage users, and add and remove daemons for IBM Storage Ceph.
Understanding process management
Monitoring a Ceph cluster
Stretch clusters for Ceph storage
Override Ceph behavior
Ceph user management
Using ceph-volume utility
Use the ceph-volume utility to prepare, list, create, activate, deactivate, batch, trigger, zap, and migrate Ceph OSDs.
Ceph performance benchmark
Ceph performance counters
mClock OSD scheduler
BlueStore
BlueStore is the back-end object store for the OSD daemons and puts objects directly on the block device.
Cephadm troubleshooting
Cephadm operations
Using cephadm-ansible modules
Use cephadm-ansible modules in Ansible playbooks to administer your IBM Storage Ceph cluster.
Ceph administration
An IBM Storage Ceph cluster is the foundation for all Ceph deployments. After deploying the cluster, there are administrative operations for keeping the cluster healthy
and perform optimally.
This section helps storage administrators to perform such tasks as:
How do I check the health of my cluster?
How do I start and stop the services?
How do I add or remove an OSD from a running cluster?
How do I manage user authentication and access controls to the objects stored in the cluster?
I want to understand how to use overrides with the cluster.
I want to monitor the performance of the cluster.
A basic Ceph storage cluster consist of two types of daemons:
A Ceph Object Storage Device (OSD) stores data as objects within placement groups assigned to the OSD
A Ceph Monitor maintains a master copy of the cluster map
A production system will have three or more Ceph Monitors for high availability and typically a minimum of 50 OSDs for acceptable load balancing, data re-balancing and
data recovery.
Reference
For more information, see Initial installation.
Understanding process management
IBM Storage Ceph 237
As a storage administrator, you can manipulate the various Ceph daemons by type or instance in an IBM Storage Ceph cluster. Manipulating these daemons allows you to
start, stop, and restart all of the Ceph services as needed.
Process management
Starting, stopping, and restarting all Ceph daemons using the systemctl command
Starting, stopping, and restarting all Ceph services
Ceph services are logical groups of Ceph daemons of the same type, configured to run in the same IBM Storage Ceph cluster. The orchestration layer in Ceph allows
the user to manage these services in a centralized way, making it easy to execute operations that affect all the Ceph daemons that belong to the same logical
service. The Ceph daemons running in each host are managed through the Systemd service. You can start, stop, and restart all Ceph services from the host where
you want to manage the Ceph services.
Viewing log files of Ceph daemons
Powering down and rebooting the cluster
Process management
In IBM Storage Ceph, all process management is done through the systemd service. Each time you want to start, restart, and stop the Ceph daemons, you must
specify the daemon type or the daemon instance.
For more information on using systemd, see the Introduction to systemd and Managing system services with systemctl chapters within the Configuring basic system
settings guide within the Product Documentation for Red Hat Enterprise Linux for your OS version, on the Managing system services with systemctl.
Starting, stopping, and restarting all Ceph daemons using the systemctl command
You can start, stop, and restart all Ceph daemons as the root user from the host where you want to stop the Ceph daemons.
Prerequisites
A running IBM Storage Ceph cluster.
Having root access to the node.
Procedure
1. On the host where you want to start, stop, and restart the daemons, run the systemctl service to get the SERVICE_ID of the service.
Example
[root@host01 ~]# systemctl --type=service
ceph-499829b4-832f-11eb-8d6d-001a4a000635@mon.host01.service
2. Starting all Ceph daemons:
Syntax
systemctl start SERVICE_ID
Example
[root@host01 ~]# systemctl start ceph-499829b4-832f-11eb-8d6d-001a4a000635@mon.host01.service
3. Stopping all Ceph daemons:
Syntax
systemctl stop SERVICE_ID
Example
[root@host01 ~]# systemctl stop ceph-499829b4-832f-11eb-8d6d-001a4a000635@mon.host01.service
4. Restarting all Ceph daemons:
Syntax
systemctl restart SERVICE_ID
Example
[root@host01 ~]# systemctl restart ceph-499829b4-832f-11eb-8d6d-001a4a000635@mon.host01.service
Starting, stopping, and restarting all Ceph services
Ceph services are logical groups of Ceph daemons of the same type, configured to run in the same IBM Storage Ceph cluster. The orchestration layer in Ceph allows the
user to manage these services in a centralized way, making it easy to execute operations that affect all the Ceph daemons that belong to the same logical service. The
Ceph daemons running in each host are managed through the Systemd service. You can start, stop, and restart all Ceph services from the host where you want to manage
the Ceph services.
238 IBM Storage Ceph
Important: To start,stop, or restart a specific Ceph daemon in a specific host, you need to use the SystemD service. To obtain a list of the SystemD services running in a
specific host, connect to the host, and run the following command:
[root@host01 ~]# systemctl list-units “ceph*”
The output provides a list of the service names that you can use to manage each Ceph daemon.
Prerequisites
A running IBM Storage Ceph cluster.
Having root access to the node.
Procedure
1. Log into the Cephadm shell:
Example
[root@host01 ~]# cephadm shell
2. Run the ceph orch ls command to get a list of Ceph services configured in the IBM Storage Ceph cluster and to get the specific service ID.
Example
[ceph: root@host01 /]# ceph orch ls
NAME
IMAGE ID
alertmanager
b7bae610cd46
crash
c88a5d60f510
grafana
bd3d7748747b
mgr
c88a5d60f510
mon
c88a5d60f510
node-exporter
osd.all-available-devices
c88a5d60f510
prometheus
bebb0ddef7f0
rgw.test_realm.test_zone
c88a5d60f510
RUNNING
REFRESHED
AGE
PLACEMENT
IMAGE NAME
1/1
4m ago
4M
count:1
cp.icr.io/cp/ibm-ceph/prometheus-alertmanager:v4.10
3/3
4m ago
4M
*
cp.icr.io/cp/ibm-ceph/ceph-6-rhel9:latest
1/1
4m ago
4M
count:1
cp.icr.io/cp/ibm-ceph/ceph-6-dashboard-rhe98:latest
2/2
4m ago
4M
count:2
cp.icr.io/cp/ibm-ceph/ceph-6-rhel9:latest
2/2
4m ago
10w
count:2
cp.icr.io/cp/ibm-ceph/ceph-6-rhel9:latest
1/3
5/5
4m ago
4m ago
4M
3M
*
*
cp.icr.io/cp/ibm-ceph/prometheus-node-exporter:v4.10
cp.icr.io/cp/ibm-ceph/ceph-6-rhel9:latest
1/1
4m ago
4M
count:1
cp.icr.io/cp/ibm-ceph/prometheus:v4.10
2/2
4m ago
3M
count:2
cp.icr.io/cp/ibm-ceph/ceph-6-rhel9:latest
mix
3. To start a specific service, run the following command:
Syntax
ceph orch start SERVICE_ID
Example
[ceph: root@host01 /]# ceph orch start node-exporter
4. To stop a specific service, run the following command:
Important: The ceph orch stop SERVICE_ID command results in the IBM Storage Ceph cluster being inaccessible, only for the MON and MGR service. It is
recommended to use the systemctl stop SERVICE_ID command to stop a specific daemon in the host.
Syntax
ceph orch stop SERVICE_ID
Example
[ceph: root@host01 /]# ceph orch stop node-exporter
In the example the ceph orch stop node-exporter command removes all the daemons of the node exporter service.
5. To restart a specific service, run the following command:
Syntax
ceph orch restart SERVICE_ID
Example
[ceph: root@host01 /]# ceph orch restart node-exporter
Viewing log files of Ceph daemons
Use the journald daemon from the container host to view a log file of a Ceph daemon from a container.
Prerequisites
IBM Storage Ceph 239
A running IBM Storage Ceph cluster.
Root-level access to the node.
Procedure
1. To view the entire Ceph log file, run a journalctl command as root composed in the following format:
Syntax
journalctl -u ceph SERVICE_ID
Example
[root@host01 ~]# journalctl -u ceph-499829b4-832f-11eb-8d6d-001a4a000635@osd.8.service
In the above example, you can view the entire log for the OSD with ID osd.8.
2. To show only the recent journal entries, use the -f option.
Syntax
journalctl -fu SERVICE_ID
Example
[root@host01 ~]# journalctl -fu ceph-499829b4-832f-11eb-8d6d-001a4a000635@osd.8.service
Note: The sosreport utility can be used to view the journald logs. For more details about SOS reports, see the What is an sosreport and how to create one in Red Hat
Enterprise Linux? solution on the Red Hat Customer Portal.
Reference
The journalctl manual page.
Powering down and rebooting the cluster
You can power down and restart the IBM Storage Ceph cluster that uses two different approaches: systemctl commands and the Ceph Orchestrator. You can choose
either of the approaches to power down and restart the cluster.
Note: When powering down or rebooting a IBM Storage Ceph cluster with the Ceph Object gateway multi-site, please ensure no IOs are in progress. Also, power off/on the
sites one at a time.
Powering down and rebooting that uses Ceph Orchestrator
Powering down and rebooting that uses systemctl commands
Prerequisites
A running IBM Storage Ceph cluster.
Root-level access.
Powering down and rebooting that uses Ceph Orchestrator
You can also use the capabilities of the Ceph Orchestrator to power down and restart the IBM Storage Ceph cluster. Usually, it is a single system login that can help in
powering off the cluster.
The Ceph Orchestrator supports several operations, such as start, stop, and restart. You can use these commands with systemctl, for some cases, in powering
down or rebooting the cluster.
Prerequisites
A running IBM Storage Ceph cluster.
Root-level access to the node.
Procedure
Powering down the IBM Storage Ceph cluster
1. Stop the clients from using the user Block Device Image and Ceph Object Gateway on this cluster and any other clients.
2. Log in to the Cephadm shell.
Example
[root@host01 ~]# cephadm shell
240 IBM Storage Ceph
3. The cluster must be in healthy state (Health_OK and all PGs active+clean) before proceeding. Run ceph status on the host with the client keyrings, for
example, the Ceph Monitor or OpenStack controller nodes, to ensure that cluster is healthy.
Example
[ceph: root@host01 /]# ceph -s
4. If you use the Ceph File System (CephFS), bring down the CephFS cluster. Syntax
ceph fs set FS_NAME max_mds 1
ceph fs fail FS_NAME
ceph status
ceph fs set FS_NAME joinable false
ceph mds fail FS_NAME:N
Example
[ceph: root@host01 /]# ceph fs set cephfs max_mds 1
[ceph: root@host01 /]# ceph fs fail cephfs
[ceph: root@host01 /]# ceph status
[ceph: root@host01 /]# ceph fs set cephfs joinable false
[ceph: root@host01 /]# ceph mds fail cephfs:1
5. Set the noout, norecover, norebalance, nobackfill, nodown, and pause flags. Run the following commands on a node with the client keyrings, for example,
the Ceph Monitor or OpenStack controller node:
Example
[ceph: root@host01 /]# ceph osd set noout
[ceph: root@host01 /]# ceph osd set norecover
[ceph: root@host01 /]# ceph osd set norebalance
[ceph: root@host01 /]# ceph osd set nobackfill
[ceph: root@host01 /]# ceph osd set nodown
[ceph: root@host01 /]# ceph osd set pause
6. Stop the MDS service.
a. Fetch the MDS service name.
Example
[ceph: root@host01 /]# ceph orch ls --service-type mds
b. Stop the MDS service that uses the fetched name in the previous step.
Syntax
ceph orch stop SERVICE_NAME
7. Stop the Ceph Object Gateway services. Repeat for each deployed service.
a. Fetch the Ceph Object Gateway service name.
Example
[ceph: root@host01 /]# ceph orch ls --service-type rgw
b. Stop the Ceph Object Gateway service that uses the fetched name.
Syntax
ceph orch stop SERVICE_NAME
8. Stop the Alertmanager service.
Example
[ceph: root@host01 /]# ceph orch stop alertmanager
9. Stop the node-exporter service, which is a part of the monitoring stack:
Example
[ceph: root@host01 /]# ceph orch stop node-exporter
10. Stop the Prometheus service.
Example
[ceph: root@host01 /]# ceph orch stop prometheus
11. Stop the Grafana dashboard service.
Example
[ceph: root@host01 /]# ceph orch stop grafana
12. Stop the crash service.
Example
[ceph: root@host01 /]# ceph orch stop crash
13. Shut down the OSD nodes from the cephadm node, one by one. Repeat this step for all the OSDs in the cluster.
IBM Storage Ceph 241
a. Fetch the OSD ID.
Example
[ceph: root@host01 /]# ceph orch ps --daemon-type=osd
b. Shut down the OSD node that uses the OSD ID you fetched.
Example
[ceph: root@host01 /]# ceph orch daemon stop osd.1
Scheduled to stop osd.1 on host 'host02'
14. Stop the monitors one by one.
a. Identify the hosts with the monitors.
Example
[ceph: root@host01 /]# ceph orch ps --daemon-type mon
b. On each host, stop the monitor.
i. Identify the systemctl unit name.
Example
[ceph: root@host01 /]# systemctl list-units ceph-* | grep mon
ii. Stop the service:
Syntax
systemct stop SERVICE-NAME
15. Shut down all the hosts.
Rebooting the IBM Storage Ceph cluster
1. If network equipment was involved, ensure that it is powered ON and stable before powering ON any Ceph hosts or nodes.
2. Power ON all the Ceph hosts.
3. Log in to the administration node from the Cephadm shell
Example
[root@host01 ~]# cephadm shell
4. Verify all the services are in running state:
Example
[ceph: root@host01 /]# ceph orch ls
5. Ensure that the cluster health is Health_OK status.
Example
[ceph: root@host01 /]# ceph -s
6. Unset the noout, norecover, norebalance, nobackfill, nodown, and pause flags. Run the following commands on a node with the client keyrings, for
example, the Ceph Monitor or OpenStack controller node.
Example
[ceph: root@host01 /]# ceph osd unset noout
[ceph: root@host01 /]# ceph osd unset norecover
[ceph: root@host01 /]# ceph osd unset norebalance
[ceph: root@host01 /]# ceph osd unset nobackfill
[ceph: root@host01 /]# ceph osd unset nodown
[ceph: root@host01 /]# ceph osd unset pause
7. If you use the Ceph File System (CephFS), bring the CephFS cluster back up by setting the joinable flag to true:
Syntax
ceph fs set FS_NAME joinable true
Example
[ceph: root@host01 /]# ceph fs set cephfs joinable true
Verification
Verify that the cluster is in healthy state (Health_OK and all PGs active+clean). Run ceph status on a node with the client keyrings, for example, the Ceph
Monitor or OpenStack controller nodes, to ensure that the cluster is healthy.
Example
[ceph: root@host01 /]# ceph -s
242 IBM Storage Ceph
Reference
For more information, see Initial installation.
Powering down and rebooting that uses systemctl commands
You can use the systemctl commands approach to power down and restart the IBM Storage Ceph cluster. This approach follows the Linux way of stopping the services.
Prerequisites
A running IBM Storage Ceph cluster.
Root-level access.
Procedure
Powering down the IBM Storage Ceph cluster
1. Stop the clients from using the Block Device images RADOS Gateway - Ceph Object Gateway on this cluster and any other clients.
2. Log in to the Cephadm shell.
Example
[root@host01 ~]# cephadm shell
3. The cluster must be healthy (Health_OK and all PGs active+clean) before proceeding. Run ceph status on the host with the client keyrings, for example, the
Ceph Monitor or OpenStack controller nodes, to ensure that the cluster is healthy.
Example
[ceph: root@host01 /]# ceph -s
4. If you use the Ceph File System (CephFS), bring down the CephFS cluster.
Syntax
ceph fs set FS_NAME max_mds 1
ceph fs fail FS_NAME
ceph status
ceph fs set FS_NAME joinable false
Example
[ceph: root@host01 /]# ceph fs set cephfs max_mds 1
[ceph: root@host01 /]# ceph fs fail cephfs
[ceph: root@host01 /]# ceph status
[ceph: root@host01 /]# ceph fs set cephfs joinable false
5. Set the noout, norecover, norebalance, nobackfill, nodown, and pause flags. Run the following commands on a node with the client keyrings, for example,
the Ceph Monitor or OpenStack controller node.
Example
[ceph: root@host01 /]# ceph osd set noout
[ceph: root@host01 /]# ceph osd set norecover
[ceph: root@host01 /]# ceph osd set norebalance
[ceph: root@host01 /]# ceph osd set nobackfill
[ceph: root@host01 /]# ceph osd set nodown
[ceph: root@host01 /]# ceph osd set pause
6. If the MDS and Ceph Object Gateway nodes are on their own dedicated nodes, power them off.
7. Get the systemd target of the daemons.
systemctl list-units --type target | grep ceph
For example,
[root@host01 ~]# systemctl list-units --type target | grep ceph
ceph-0b007564-ec48-11ee-b736-525400fd02f8.target loaded active active Ceph cluster 0b007564-ec48-11ee-b736-525400fd02f8
ceph.target
loaded active active All Ceph clusters and services
8. Disable the target that includes the cluster FSID.
systemctl disable CLUSTER_FSID.target
For example,
[root@host01 ~]# systemctl disable ceph-0b007564-ec48-11ee-b736-525400fd02f8.target
Removed "/etc/systemd/system/multi-user.target.wants/ceph-0b007564-ec48-11ee-b736-525400fd02f8.target".
Removed "/etc/systemd/system/ceph.target.wants/ceph-0b007564-ec48-11ee-b736-525400fd02f8.target".
9. Stop the target.
systemctl stop CLUSTER_FSID.target
IBM Storage Ceph 243
For example,
[root@host01 ~]# systemctl stop ceph-0b007564-ec48-11ee-b736-525400fd02f8.target
This stops all the daemons on the host that needs to be stopped.
10. Shutdown the node.
shutdown
For example,
[root@host01 ~]# shutdown
Shutdown scheduled for Wed 2024-03-27 11:47:19 EDT, use 'shutdown -c' to cancel.
11. Repeat the above steps for all the nodes of the cluster.
Rebooting the IBM Storage Ceph cluster
1. If network equipment was involved, ensure that it is powered ON and stable before powering ON any Ceph hosts or nodes.
2. Enable the systemd target to get all the daemons running.
systemctl enable CLUSTER_FSID.target
For example,
[root@host01 ~]# systemctl enable ceph-0b007564-ec48-11ee-b736-525400fd02f8.target
Created symlink /etc/systemd/system/multi-user.target.wants/ceph-0b007564-ec48-11ee-b736-525400fd02f8.target →
/etc/systemd/system/ceph-0b007564-ec48-11ee-b736-525400fd02f8.target.
Created symlink /etc/systemd/system/ceph.target.wants/ceph-0b007564-ec48-11ee-b736-525400fd02f8.target →
/etc/systemd/system/ceph-0b007564-ec48-11ee-b736-525400fd02f8.target.
3. Start the systemd target.
systemctl start CLUSTER_FSID.target
For example,
[root@host01 ~]# systemctl start ceph-0b007564-ec48-11ee-b736-525400fd02f8.target
4. Wait for all the nodes to come up. Verify all the services are up and no connectivity issues are there between the nodes.
5. Unset the noout, norecover, norebalance, nobackfill, nodown, and pause flags. Run the following commands on a node with the client keyrings, for
example, the Ceph Monitor or OpenStack controller node:
Example
[ceph: root@host01 /]# ceph osd unset noout
[ceph: root@host01 /]# ceph osd unset norecover
[ceph: root@host01 /]# ceph osd unset norebalance
[ceph: root@host01 /]# ceph osd unset nobackfill
[ceph: root@host01 /]# ceph osd unset nodown
[ceph: root@host01 /]# ceph osd unset pause
6. If you use the Ceph File System (CephFS), bring the CephFS cluster back up by setting the joinable flag to true.
Syntax
ceph fs set FS_NAME joinable true
Example
[ceph: root@host01 /]# ceph fs set cephfs joinable true
Verification
Verify that the cluster is healthy (Health_OK and all PGs active+clean). Run ceph status on a node with the client keyrings, for example, the Ceph Monitor or
OpenStack controller nodes, to ensure that the cluster is healthy.
Example
[ceph: root@host01 /]# ceph -s
Reference
For more information about installing Ceph, see Initial installation.
Monitoring a Ceph cluster
As a storage administrator, you can monitor the overall health of the IBM Storage Cephcluster, along with monitoring the health of the individual components of Ceph.
Once you have a running cluster, you might begin monitoring the storage cluster to ensure that the Ceph Monitor and Ceph OSD daemons are running, at a high-level. Ceph
storage cluster clients connect to a Ceph Monitor and receive the latest version of the storage cluster map before they can read and write data to the Ceph pools within the
storage cluster. So the monitor cluster must have agreement on the state of the cluster before Ceph clients can read and write data.
244 IBM Storage Ceph
Ceph OSDs must peer the placement groups on the primary OSD with the copies of the placement groups on secondary OSDs. If faults arise, peering will reflect something
other than the active + clean state.
High-level monitoring
Low-level monitoring
High-level monitoring
As a storage administrator, you can monitor the health of the Ceph daemons to ensure that they are up and running. High level monitoring also involves checking the
storage cluster capacity to ensure that the storage cluster does not exceed its full ratio. The IBM Storage Ceph Dashboard is the most common way to conduct highlevel monitoring. However, you can also use the command-line interface, the Ceph admin socket or the Ceph API to monitor the storage cluster.
Checking storage cluster health
Watching storage cluster events
How Ceph calculates data usage
Understanding storage clusters usage stats
Checking storage cluster status
Understanding OSD usage stats
Understanding Ceph OSD status
Checking Ceph Monitor status
Using Ceph administration socket
Checking storage cluster health
After you start the Ceph storage cluster, and before you start reading or writing data, check the storage cluster’s health first.
Prerequisites
A running IBM Storage Ceph cluster.
Root-level access to the node.
Procedure
1. Log into the Cephadm shell:
Example
[root@host01 ~]# cephadm shell
2. You can check on the health of the Ceph storage cluster with the following command:
Example
[ceph: root@host01 /]# ceph health
HEALTH_OK
3. You can check the status of the Ceph storage cluster by running ceph status command:
Example
[ceph: root@host01 /]# ceph status
The output provides the following information:
Cluster ID
Cluster health status
The monitor map epoch and the status of the monitor quorum.
The OSD map epoch and the status of OSDs.
The status of Ceph Managers.
The status of Object Gateways.
The placement group map version.
The number of placement groups and pools.
The notional amount of data stored and the number of objects stored.
The total amount of data stored.
The IO client operations.
An update on the upgrade process if the cluster is upgrading.
IBM Storage Ceph 245
Upon starting the Ceph cluster, you will likely encounter a health warning such as HEALTH_WARN XXX num placement groups stale. Wait a few
moments and check it again. When the storage cluster is ready, ceph health should return a message such as HEALTH_OK. At that point, it is okay to begin
using the cluster.
Watching storage cluster events
You can watch events that are happening with the Ceph storage cluster using the command-line interface.
Prerequisites
A running IBM Storage Ceph cluster.
Root-level access to the node.
Procedure
1. Log into the Cephadm shell:
Example
root@host01 ~]# cephadm shell
2. To watch the cluster’s ongoing events, run the following command:
Example
[ceph: root@host01 /]# ceph -w
cluster:
id:
8c9b0072-67ca-11eb-af06-001a4a0002a0
health: HEALTH_OK
services:
mon: 2 daemons, quorum Ceph5-2,Ceph5-adm (age 3d)
mgr: Ceph5-1.nqikfh(active, since 3w), standbys: Ceph5-adm.meckej
osd: 5 osds: 5 up (since 2d), 5 in (since 8w)
rgw: 2 daemons active (test_realm.test_zone.Ceph5-2.bfdwcn, test_realm.test_zone.Ceph5-adm.acndrh)
data:
pools:
11 pools, 273 pgs
objects: 459 objects, 32 KiB
usage:
2.6 GiB used, 72 GiB / 75 GiB avail
pgs:
273 active+clean
io:
client:
170 B/s rd, 730 KiB/s wr, 0 op/s rd, 729 op/s wr
2022-11-02 15:45:21.655871 osd.0 [INF] 17.71 deep-scrub ok
2022-11-02 15:45:47.880608 osd.1 [INF] 1.0 scrub ok
2022-11-02 15:45:48.865375 osd.1 [INF] 1.3 scrub ok
2022-11-02 15:45:50.866479 osd.1 [INF] 1.4 scrub ok
2022-11-02 15:45:01.345821 mon.0 [INF] pgmap v41339: 952 pgs: 952 active+clean; 17130 MB data, 115 GB used, 167 GB / 297
GB avail
2022-11-02 15:45:05.718640 mon.0 [INF] pgmap v41340: 952 pgs: 1 active+clean+scrubbing+deep, 951 active+clean; 17130 MB
data, 115 GB used, 167 GB / 297 GB avail
2022-11-02 15:45:53.997726 osd.1 [INF] 1.5 scrub ok
2022-11-02 15:45:06.734270 mon.0 [INF] pgmap v41341: 952 pgs: 1 active+clean+scrubbing+deep, 951 active+clean; 17130 MB
data, 115 GB used, 167 GB / 297 GB avail
2022-11-02 15:45:15.722456 mon.0 [INF] pgmap v41342: 952 pgs: 952 active+clean; 17130 MB data, 115 GB used, 167 GB / 297
GB avail
2022-11-02 15:46:06.836430 osd.0 [INF] 17.75 deep-scrub ok
2022-11-02 15:45:55.720929 mon.0 [INF] pgmap v41343: 952 pgs: 1 active+clean+scrubbing+deep, 951 active+clean; 17130 MB
data, 115 GB used, 167 GB / 297 GB avail
How Ceph calculates data usage
The used value reflects the actual amount of raw storage used. The xxx GB / xxx GB value means the amount available, the lesser of the two numbers, of the overall
storage capacity of the cluster. The notional number reflects the size of the stored data before it is replicated, cloned or snapshotted. Therefore, the amount of data
actually stored typically exceeds the notional amount stored, because Ceph creates replicas of the data and may also use storage capacity for cloning and snapshotting.
Understanding storage clusters usage stats
To check a cluster’s data usage and data distribution among pools, use the df option. It is similar to the Linux df command. You can run either the ceph
df command or ceph df detail command.
The SIZE/AVAIL/RAW USED in the ceph
df and ceph status command output are different if some OSDs are marked OUT of the cluster compared to when all OSDs are IN. The SIZE/AVAIL/RAW USED is
calculated from sum of SIZE (osd disk size), RAW USE (total used space on disk), and AVAIL of all OSDs which are in IN state. You can see the total of SIZE/AVAIL/RAW
USED for all OSDs in ceph osd df tree command output.
246 IBM Storage Ceph
Example
[ceph: root@host01 /]#ceph df
--- RAW STORAGE --CLASS
SIZE
AVAIL
USED
hdd
5 TiB 2.9 TiB 2.1 TiB
TOTAL 5 TiB 2.9 TiB 2.1 TiB
RAW USED
2.1 TiB
2.1 TiB
--- POOLS --POOL
.mgr
.rgw.root
default.rgw.log
default.rgw.control
default.rgw.meta
default.rgw.buckets.index
default.rgw.buckets.data
default.rgw.buckets.non-ec
source-ecpool-86
PGS
1
32
32
32
32
32
32
32
32
ID
1
2
3
4
5
7
8
9
11
%RAW USED
42.98
42.98
STORED
5.3 MiB
1.3 KiB
3.6 KiB
0 B
1.7 KiB
5.5 MiB
807 KiB
1.0 MiB
1.2 TiB
OBJECTS
3
4
209
8
10
22
3
1
391.13k
USED
16 MiB
48 KiB
408 KiB
0 B
96 KiB
17 MiB
2.4 MiB
3.1 MiB
2.1 TiB
%USED
0
0
0
0
0
0
0
0
53.49
MAX AVAIL
629 GiB
629 GiB
629 GiB
629 GiB
629 GiB
629 GiB
629 GiB
629 GiB
1.1 TiB
The ceph df detail command gives more details about other pool statistics such as quota objects, quota bytes, used compression, and under compression.
The RAW STORAGE section of the output provides an overview of the amount of storage the storage cluster manages for data.
CLASS: The class of OSD device.
SIZE: The amount of storage capacity managed by the storage cluster.
In the above example, if the SIZE is 90 GiB, it is the total size without the replication factor, which is three by default. The total available capacity with the
replication factor is 90 GiB/3 = 30 GiB. Based on the full ratio, which is 0.85% by default, the maximum available space is 30 GiB * 0.85 = 25.5 GiB
AVAIL: The amount of free space available in the storage cluster.
In the above example, if the SIZE is 90 GiB and the USED space is 6 GiB, then the AVAIL space is 84 GiB. The total available space with the replication factor, which
is three by default, is 84 GiB/3 = 28 GiB
USED: The amount of raw storage consumed by user data.
In the above example, 100 MiB is the total space available after considering the replication factor. The actual available size is 33 MiB.
RAW USED: The amount of raw storage consumed by user data, internal overhead, or reserved capacity.
% RAW USED: The percentage of RAW USED. Use this number in conjunction with the full ratio and near full ratio to ensure that you are not reaching
the storage cluster’s capacity.
The POOLS section of the output provides a list of pools and the notional usage of each pool. The output from this section DOES NOT reflect replicas, clones or snapshots.
For example, if you store an object with 1 MB of data, the notional usage will be 1 MB, but the actual usage may be 3 MB or more depending on the number of replicas for
example, size =
3, clones and snapshots.
POOL: The name of the pool.
ID: The pool ID.
STORED: The actual amount of data stored by the user in the pool. This value changes based on the raw usage data based on (k+M)/K values, number of object
copies, and the number of objects degraded at the time of pool stats calculation.
OBJECTS: The notional number of objects stored per pool. It is STORED size * replication factor.
USED: The notional amount of data stored in kilobytes, unless the number appends M for megabytes or G for gigabytes.
%USED: The notional percentage of storage used per pool.
MAX AVAIL: An estimate of the notional amount of data that can be written to this pool. It is the amount of data that can be used before the first OSD becomes full.
It considers the projected distribution of data across disks from the CRUSH map and uses the first OSD to fill up as the target.
In the above example, MAX AVAIL is 153.85 MB without considering the replication factor, which is three by default.
See the Knowledgebase article titled ceph df MAX AVAIL is incorrect for simple replicated pool to calculate the value of MAX AVAIL.
QUOTA OBJECTS: The number of quota objects.
QUOTA BYTES: The number of bytes in the quota objects.
USED COMPR: The amount of space allocated for compressed data including his includes compressed data, allocation, replication and erasure coding overhead.
UNDER COMPR: The amount of data passed through compression and beneficial enough to be stored in a compressed form.
Note:
The numbers in the POOLS section are notional. They are not inclusive of the number of replicas, snapshots or clones. As a result, the sum of the USED and %USED
amounts will not add up to the RAW USED and %RAW USED amounts in the GLOBAL section of the output.
The MAX AVAIL value is a complicated function of the replication or erasure code used, the CRUSH rule that maps storage to devices, the utilization of those
devices, and the configured mon_osd_full_ratio.
For more information, see How Ceph calculates data usage and Understanding OSD usage stats.
Checking storage cluster status
IBM Storage Ceph 247
You can check the status of the IBM Storage Ceph cluster from the command-line interface. The status sub command or the -s argument will display the current status
of the storage cluster.
Prerequisites
A running IBM Storage Ceph cluster.
Root-level access to the node.
Procedure
1. Log into the Cephadm shell:
Example
[root@host01 ~]# cephadm shell
2. To check a storage cluster’s status, execute the following:
Example
[ceph: root@host01 /]# ceph status
Or
Example
[ceph: root@host01 /]# ceph -s
3. In interactive mode, type ceph and press Enter:
Example
[ceph: root@host01 /]# ceph
ceph> status
cluster:
id:
499829b4-832f-11eb-8d6d-001a4a000635
health: HEALTH_WARN
1 stray daemon(s) not managed by cephadm
1/3 mons down, quorum host03,host02
too many PGs per OSD (261 > max 250)
services:
mon:
mgr:
osd:
rgw:
3 daemons, quorum host03,host02 (age 3d), out of quorum: host01
host01.hdhzwn(active, since 9d), standbys: host05.eobuuv, host06.wquwpj
12 osds: 11 up (since 2w), 11 in (since 5w)
2 daemons active (test_realm.test_zone.host04.hgbvnq, test_realm.test_zone.host05.yqqilm)
data:
pools:
8 pools, 960 pgs
objects: 414 objects, 1.0 MiB
usage:
5.7 GiB used, 214 GiB / 220 GiB avail
pgs:
960 active+clean
io:
client:
41 KiB/s rd, 0 B/s wr, 41 op/s rd, 27 op/s wr
ceph> health
HEALTH_WARN 1 stray daemon(s) not managed by cephadm; 1/3 mons down, quorum host03,host02; too many PGs per OSD (261 > max
250)
ceph> mon stat
e3: 3 mons at {host01=[v2:10.74.255.0:3300/0,v1:10.74.255.0:6789/0],host02=
[v2:10.74.249.253:3300/0,v1:10.74.249.253:6789/0],host03=[v2:10.74.251.164:3300/0,v1:10.74.251.164:6789/0]}, election
epoch 6688, leader 1 host03, quorum 1,2 host03,host02
Understanding OSD usage stats
Use the ceph osd df command to view OSD utilization stats.
Example
[ceph: root@host01 /]# ceph osd df
ID CLASS WEIGHT REWEIGHT SIZE
USE
DATA
OMAP
META
AVAIL
%USE VAR PGS
3
hdd 0.90959 1.00000 931GiB 70.1GiB 69.1GiB
0B
1GiB 861GiB 7.53 2.93 66
4
hdd 0.90959 1.00000 931GiB 1.30GiB 308MiB
0B
1GiB 930GiB 0.14 0.05 59
0
hdd 0.90959 1.00000 931GiB 18.1GiB 17.1GiB
0B
1GiB 913GiB 1.94 0.76 57
MIN/MAX VAR: 0.02/2.98 STDDEV: 2.91
ID: The name of the OSD.
CLASS: The type of devices the OSD uses.
WEIGHT: The weight of the OSD in the CRUSH map.
REWEIGHT: The default reweight value.
248 IBM Storage Ceph
SIZE: The overall storage capacity of the OSD.
USE: The OSD capacity.
DATA: The amount of OSD capacity that is used by user data.
OMAP: An estimate value of the bluefs storage that is being used to store object map (omap) data (key value pairs stored in rocksdb).
META: The bluefs space allocated, or the value set in the bluestore_bluefs_min parameter, whichever is larger, for internal metadata which is calculated as
the total space allocated in bluefs minus the estimated omap data size.
AVAIL: The amount of free space available on the OSD.
%USE: The notional percentage of storage used by the OSD
VAR: The variation above or below average utilization.
PGS: The number of placement groups in the OSD.
MIN/MAX VAR: The minimum and maximum variation across all OSDs.
Reference
For more information, see:
How Ceph calculates data usage
CRUSH weights
Understanding Ceph OSD status
A Ceph OSD’s status is either in the storage cluster, or out of the storage cluster. It is either up and running, or it is down and not running. If a Ceph OSD is up, it can be
either in the storage cluster, where data can be read and written, or it is out of the storage cluster. If it was in the storage cluster and recently moved out of the storage
cluster, Ceph starts migrating placement groups to other Ceph OSDs. If a Ceph OSD is out of the storage cluster, CRUSH will not assign placement groups to the Ceph
OSD. If a Ceph OSD is down, it should also be out.
Note: If a Ceph OSD is down and in, there is a problem, and the storage cluster will not be in a healthy state.
Figure 1. OSD States
If you execute a command such as ceph health, ceph -s or ceph -w, you might notice that the storage cluster does not always echo back HEALTH OK. Do not panic.
With respect to Ceph OSDs, you can expect that the storage cluster will NOT echo HEALTH OK in a few expected circumstances:
You have not started the storage cluster yet, and it is not responding.
You have just started or restarted the storage cluster, and it is not ready yet, because the placement groups are getting created and the Ceph OSDs are in the
process of peering.
You just added or removed a Ceph OSD.
You just modified the storage cluster map.
An important aspect of monitoring Ceph OSDs is to ensure that when the storage cluster is up and running that all Ceph OSDs that are in the storage cluster are up and
running, too.
To see if all OSDs are running, execute:
Example
[ceph: root@host01 /]# ceph osd stat
or
Example
[ceph: root@host01 /]# ceph osd dump
The result should tell you the map epoch, eNNNN, the total number of OSDs, x, how many, y, are up, and how many, z, are in:
IBM Storage Ceph 249
eNNNN: x osds: y up, z in
If the number of Ceph OSDs that are in the storage cluster are more than the number of Ceph OSDs that are up. Execute the following command to identify the ceph-osd
daemons that are not running:
Example
[ceph: root@host01 /]# ceph osd tree
# id
-1 3
-3 3
-2 3
0
1
1
1
2
1
weight type name
up/down reweight
pool default
rack mainrack
host osd-host
osd.0
up 1
osd.1
up 1
osd.2
up 1
Tip: The ability to search through a well-designed CRUSH hierarchy can help you troubleshoot the storage cluster by identifying the physical locations faster.
If a Ceph OSD is down, connect to the node and start it. You can use IBM Storage Ceph Console to restart the Ceph OSD daemon, or you can use the command line.
Syntax
systemctl start CEPH_OSD_SERVICE_ID
Example
[root@host01 ~]# systemctl start ceph-499829b4-832f-11eb-8d6d-001a4a000635@osd.6.service
Reference
For more information, see Dashboard.
Checking Ceph Monitor status
If the storage cluster has multiple Ceph Monitors, which is a requirement for a production IBM Storage Ceph cluster, then you can check the Ceph Monitor quorum status
after starting the storage cluster, and before doing any reading or writing of data.
A quorum must be present when multiple Ceph Monitors are running.
Check the Ceph Monitor status periodically to ensure that they are running. If there is a problem with the Ceph Monitor, that prevents an agreement on the state of the
storage cluster, the fault can prevent Ceph clients from reading and writing data.
Prerequisites
A running IBM Storage Ceph cluster.
Root-level access to the node.
Procedure
1. Log into the Cephadm shell:
Example
[root@host01 ~]# cephadm shell
2. To display the Ceph Monitor map, execute the following:
Example
[ceph: root@host01 /]# ceph mon stat
or
Example
[ceph: root@host01 /]# ceph mon dump
3. To check the quorum status for the storage cluster, execute the following:
[ceph: root@host01 /]# ceph quorum_status -f json-pretty
Ceph returns the quorum status.
Example
{
"election_epoch": 6686,
"quorum": [
0,
1,
2
],
"quorum_names": [
"host01",
250 IBM Storage Ceph
"host03",
"host02"
],
"quorum_leader_name": "host01",
"quorum_age": 424884,
"features": {
"quorum_con": "4540138297136906239",
"quorum_mon": [
"kraken",
"luminous",
"mimic",
"osdmap-prune",
"nautilus",
"octopus",
"pacific",
"elector-pinging"
]
},
"monmap": {
"epoch": 3,
"fsid": "499829b4-832f-11eb-8d6d-001a4a000635",
"modified": "2021-03-15T04:51:38.621737Z",
"created": "2021-03-12T12:35:16.911339Z",
"min_mon_release": 16,
"min_mon_release_name": "pacific",
"election_strategy": 1,
"disallowed_leaders: ": "",
"stretch_mode": false,
"features": {
"persistent": [
"kraken",
"luminous",
"mimic",
"osdmap-prune",
"nautilus",
"octopus",
"pacific",
"elector-pinging"
],
"optional": []
},
"mons": [
{
"rank": 0,
"name": "host01",
"public_addrs": {
"addrvec": [
{
"type": "v2",
"addr": "10.74.255.0:3300",
"nonce": 0
},
{
"type": "v1",
"addr": "10.74.255.0:6789",
"nonce": 0
}
]
},
"addr": "10.74.255.0:6789/0",
"public_addr": "10.74.255.0:6789/0",
"priority": 0,
"weight": 0,
"crush_location": "{}"
},
{
"rank": 1,
"name": "host03",
"public_addrs": {
"addrvec": [
{
"type": "v2",
"addr": "10.74.251.164:3300",
"nonce": 0
},
{
"type": "v1",
"addr": "10.74.251.164:6789",
"nonce": 0
}
]
},
"addr": "10.74.251.164:6789/0",
"public_addr": "10.74.251.164:6789/0",
"priority": 0,
"weight": 0,
"crush_location": "{}"
},
{
"rank": 2,
"name": "host02",
"public_addrs": {
"addrvec": [
{
"type": "v2",
"addr": "10.74.249.253:3300",
IBM Storage Ceph 251
},
{
]
}
]
}
}
}
"nonce": 0
"type": "v1",
"addr": "10.74.249.253:6789",
"nonce": 0
},
"addr": "10.74.249.253:6789/0",
"public_addr": "10.74.249.253:6789/0",
"priority": 0,
"weight": 0,
"crush_location": "{}"
Using Ceph administration socket
Use the administration socket to interact with a given daemon directly by using a UNIX socket file. For example, the socket enables you to:
List the Ceph configuration at runtime
Set configuration values at runtime directly without relying on Monitors. This is useful when Monitors are down.
Dump historic operations
Dump the operation priority queue state
Dump operations without rebooting
Dump performance counters
In addition, using the socket is helpful when troubleshooting problems related to Ceph Monitors or OSDs.
Regardless, if the daemon is not running, a following error is returned when attempting to use the administration socket:
Error 111: Connection Refused
Important: The administration socket is only available while a daemon is running. When you shut down the daemon properly, the administration socket is removed.
However, if the daemon terminates unexpectedly, the administration socket might persist.
Prerequisites
A running IBM Storage Ceph cluster.
Root-level access to the node.
Procedure
1. Log into the Cephadm shell:
Example
[root@host01 ~]# cephadm shell
2. To use the socket:
Syntax
ceph daemon MONITOR_ID COMMAND
Replace:
MONITOR_ID of the daemon
COMMAND with the command to run. Use help to list the available commands for a given daemon.
To view the status of a Ceph Monitor:
Example
[ceph: root@host01 /]# ceph daemon mon.host01 help
{
"add_bootstrap_peer_hint": "add peer address as potential bootstrap peer for cluster bringup",
"add_bootstrap_peer_hintv": "add peer address vector as potential bootstrap peer for cluster bringup",
"compact": "cause compaction of monitor's leveldb/rocksdb storage",
"config diff": "dump diff of current config and default config",
"config diff get": "dump diff get <field>: dump diff of current and default config setting <field>",
"config get": "config get <field>: get the config value",
"config help": "get config setting schema and descriptions",
"config set": "config set <field> <val> [<val> ...]: set a config variable",
"config show": "dump current config settings",
"config unset": "config unset <field>: unset a config variable",
"connection scores dump": "show the scores used in connectivity-based elections",
252 IBM Storage Ceph
}
"connection scores reset": "reset the scores used in connectivity-based elections",
"counter dump": "dump all labeled and non-labeled counters and their values",
"counter schema": "dump all labeled and non-labeled counters schemas",
"dump_historic_ops": "show recent ops",
"dump_historic_slow_ops": "show recent slow ops",
"dump_mempools": "get mempool stats",
"get_command_descriptions": "list available commands",
"git_version": "get git sha1",
"heap": "show heap usage info (available only if compiled with tcmalloc)",
"help": "list available commands",
"injectargs": "inject configuration arguments into running daemon",
"log dump": "dump recent log entries to log file",
"log flush": "flush log entries to log file",
"log reopen": "reopen log file",
"mon_status": "report status of monitors",
"ops": "show the ops currently in flight",
"perf dump": "dump non-labeled counters and their values",
"perf histogram dump": "dump perf histogram values",
"perf histogram schema": "dump perf histogram schema",
"perf reset": "perf reset <name>: perf reset all or one perfcounter name",
"perf schema": "dump non-labeled counters schemas",
"quorum enter": "force monitor back into quorum",
"quorum exit": "force monitor out of the quorum",
"sessions": "list existing sessions",
"smart": "Query health metrics for underlying device",
"sync_force": "force sync of and clear monitor store",
"version": "get ceph version"
Example
[ceph: root@host01 /]# ceph daemon mon.host01 mon_status
{
"name": "host01",
"rank": 0,
"state": "leader",
"election_epoch": 120,
"quorum": [
0,
1,
2
],
"quorum_age": 206358,
"features": {
"required_con": "2449958747317026820",
"required_mon": [
"kraken",
"luminous",
"mimic",
"osdmap-prune",
"nautilus",
"octopus",
"pacific",
"elector-pinging"
],
"quorum_con": "4540138297136906239",
"quorum_mon": [
"kraken",
"luminous",
"mimic",
"osdmap-prune",
"nautilus",
"octopus",
"pacific",
"elector-pinging"
]
},
"outside_quorum": [],
"extra_probe_peers": [],
"sync_provider": [],
"monmap": {
"epoch": 3,
"fsid": "81a4597a-b711-11eb-8cb8-001a4a000740",
"modified": "2021-05-18T05:50:17.782128Z",
"created": "2021-05-17T13:13:13.383313Z",
"min_mon_release": 16,
"min_mon_release_name": "pacific",
"election_strategy": 1,
"disallowed_leaders: ": "",
"stretch_mode": false,
"features": {
"persistent": [
"kraken",
"luminous",
"mimic",
"osdmap-prune",
"nautilus",
"octopus",
"pacific",
"elector-pinging"
],
"optional": []
},
"mons": [
{
IBM Storage Ceph 253
},
{
},
{
}
"rank": 0,
"name": "host01",
"public_addrs": {
"addrvec": [
{
"type": "v2",
"addr": "10.74.249.41:3300",
"nonce": 0
},
{
"type": "v1",
"addr": "10.74.249.41:6789",
"nonce": 0
}
]
},
"addr": "10.74.249.41:6789/0",
"public_addr": "10.74.249.41:6789/0",
"priority": 0,
"weight": 0,
"crush_location": "{}"
"rank": 1,
"name": "host02",
"public_addrs": {
"addrvec": [
{
"type": "v2",
"addr": "10.74.249.55:3300",
"nonce": 0
},
{
"type": "v1",
"addr": "10.74.249.55:6789",
"nonce": 0
}
]
},
"addr": "10.74.249.55:6789/0",
"public_addr": "10.74.249.55:6789/0",
"priority": 0,
"weight": 0,
"crush_location": "{}"
"rank": 2,
"name": "host03",
"public_addrs": {
"addrvec": [
{
"type": "v2",
"addr": "10.74.249.49:3300",
"nonce": 0
},
{
"type": "v1",
"addr": "10.74.249.49:6789",
"nonce": 0
}
]
},
"addr": "10.74.249.49:6789/0",
"public_addr": "10.74.249.49:6789/0",
"priority": 0,
"weight": 0,
"crush_location": "{}"
}
]
},
"feature_map": {
"mon": [
{
"features": "0x3f01cfb9fffdffff",
"release": "luminous",
"num": 1
}
],
"osd": [
{
"features": "0x3f01cfb9fffdffff",
"release": "luminous",
"num": 3
}
]
},
"stretch_mode": false
3. Alternatively, specify the Ceph daemon by using its socket file:
Syntax
ceph daemon /var/run/ceph/SOCKET_FILE COMMAND
254 IBM Storage Ceph
4. To view the status of a Ceph OSD named osd.0 on the specific host:
Example
[ceph: root@host01 /]# ceph daemon /var/run/ceph/ceph-osd.0.asok status
{
"cluster_fsid": "9029b252-1668-11ee-9399-001a4a000429",
"osd_fsid": "1de9b064-b7a5-4c54-9395-02ccda637d21",
"whoami": 0,
"state": "active",
"oldest_map": 1,
"newest_map": 58,
"num_pgs": 33
}
Note: You can use help instead of status for the various options that are available for the specific daemon.
5. List all socket files for the Ceph processes:
Example
[ceph: root@host01 /]# ls /var/run/ceph
Reference
For more information, see Identifying problems.
Low-level monitoring
As a storage administrator, you can monitor the health of an IBM Storage Ceph cluster from a low-level perspective. Low-level monitoring typically involves ensuring that
Ceph OSDs are peering properly. When peering faults occur, placement groups operate in a degraded state. This degraded state can be the result of many different things,
such as hardware failure, a hung or crashed Ceph daemon, network latency, or a complete site outage.
Monitoring placement group sets
Ceph OSD peering
Placement Group States
Placement Group creating state
Placement group peering state
Placement group active state
Placement Group clean state
Placement Group degraded state
Placement Group recovering state
Back fill state
Placement Group remapped state
Placement Group stale state
Placement Group misplaced state
Placement Group incomplete state
Identifying stuck Placement Groups
Finding object’s location
Monitoring placement group sets
When CRUSH assigns placement groups to Ceph OSDs, it looks at the number of replicas for the pool and assigns the placement group to Ceph OSDs such that each
replica of the placement group gets assigned to a different Ceph OSD. For example, if the pool requires three replicas of a placement group, CRUSH may assign them to
osd.1, osd.2 and osd.3 respectively. CRUSH actually seeks a pseudo-random placement that will take into account failure domains you set in the CRUSH map, so you
will rarely see placement groups assigned to nearest neighbor Ceph OSDs in a large cluster. We refer to the set of Ceph OSDs that should contain the replicas of a
particular placement group as the Acting Set. In some cases, an OSD in the Acting Set is down or otherwise not able to service requests for objects in the placement
group. When these situations arise, do not panic. Common examples include:
You added or removed an OSD. Then, CRUSH reassigned the placement group to other Ceph OSDs, thereby changing the composition of the acting set and
spawning the migration of data with a "backfill" process.
A Ceph OSD was down, was restarted and is now recovering.
A Ceph OSD in the acting set is down or unable to service requests, and another Ceph OSD has temporarily assumed its duties.
Ceph processes a client request using the Up Set, which is the set of Ceph OSDs that actually handle the requests. In most cases, the up set and the Acting Set are
virtually identical. When they are not, it can indicate that Ceph is migrating data, a Ceph OSD is recovering, or that there is a problem, that is, Ceph usually echoes a
HEALTH WARN state with a "stuck stale" message in such scenarios.
Prerequisites
A running IBM Storage Ceph cluster.
Root-level access to the node.
Procedure
IBM Storage Ceph 255
1. Log into the Cephadm shell:
Example
[root@host01 ~]# cephadm shell
2. To retrieve a list of placement groups:
Example
[ceph: root@host01 /]# ceph pg dump
3. View which Ceph OSDs are in the Acting Set or in the Up Set for a given placement group:
Syntax
ceph pg map PG_NUM
Example
[ceph: root@host01 /]# ceph pg map 128
Note: If the Up Set and Acting Set do not match, this may be an indicator that the storage cluster rebalancing itself or of a potential problem with the storage
cluster.
Ceph OSD peering
Before you can write data to a placement group, it must be in an active state, and it should be in a clean state. For Ceph to determine the current state of a placement
group, the primary OSD of the placement group that is, the first OSD in the acting set, peers with the secondary and tertiary OSDs to establish agreement on the current
state of the placement group. Assuming a pool with three replicas of the PG.
Peering
Figure 1. Peering
Placement Group States
If you execute a command such as ceph health, ceph -s or ceph -w, you may notice that the cluster does not always echo back HEALTH
OK. After you check to see if the OSDs are running, you should also check placement group states. You should expect that the cluster does NOT echo HEALTH OK in a
number of placement group peering-related circumstances:
You have just created a pool and placement groups have not peered yet.
The placement groups are recovering.
You have just added an OSD to or removed an OSD from the cluster.
You have just modified the CRUSH map and the placement groups are migrating.
There is inconsistent data in different replicas of a placement group.
Ceph is scrubbing a placement group’s replicas.
Ceph does not have enough storage capacity to complete backfilling operations.
If one of the foregoing circumstances causes Ceph to echo HEALTH WARN, do not panic. In many cases, the cluster will recover on its own. In some cases, you may need
to take action. An important aspect of monitoring placement groups is to ensure that when the cluster is up and running that all placement groups are active, and
preferably in the clean state.
To see the status of all placement groups, execute:
Example
[ceph: root@host01 /]# ceph pg stat
256 IBM Storage Ceph
The result should tell you the placement group map version, vNNNNNN, the total number of placement groups, x, and how many placement groups, y, are in a particular
state such as active+clean:
vNNNNNN: x pgs: y active+clean; z bytes data, aa MB used, bb GB / cc GB avail
Note: It is common for Ceph to report multiple states for placement groups.
Snapshot Trimming PG States
When snapshots exist, two additional PG states will be reported.
snaptrim : The PGs are currently being trimmed
snaptrim_wait : The PGs are waiting to be trimmed
Example Output:
244 active+clean+snaptrim_wait
32 active+clean+snaptrim
In addition to the placement group states, Ceph will also echo back the amount of data used, aa, the amount of storage capacity remaining, bb, and the total storage
capacity cc for the placement group. These numbers can be important in a few cases:
You are reaching the near full ratio or full ratio.
Your data isn’t getting distributed across the cluster due to an error in the CRUSH configuration.
Placement Group IDs
Placement group IDs consist of the pool number, and not the pool name, followed by a period (.) and the placement group ID - a hexadecimal number. You can view pool
numbers and their names from the output of ceph osd lspools. The default pool names data, metadata and rbd correspond to pool numbers 0, 1 and 2 respectively.
A fully qualified placement group ID has the following form:
Syntax
POOL_NUM.PG_ID
Example output:
0.1f
To retrieve a list of placement groups:
Example
[ceph: root@host01 /]# ceph pg dump
To format the output in JSON format and save it to a file:
Syntax
ceph pg dump -o FILE_NAME --format=json
Example
[ceph: root@host01 /]# ceph pg dump -o test --format=json
Query a particular placement group:
Syntax
ceph pg POOL_NUM.PG_ID query
Example
[ceph: root@host01 /]# ceph pg 5.fe query
{
"snap_trimq": "[]",
"snap_trimq_len": 0,
"state": "active+clean",
"epoch": 2449,
"up": [
3,
8,
10
],
"acting": [
3,
8,
10
],
"acting_recovery_backfill": [
"3",
"8",
"10"
],
"info": {
"pgid": "5.ff",
"last_update": "0'0",
"last_complete": "0'0",
"log_tail": "0'0",
"last_user_version": 0,
"last_backfill": "MAX",
"purged_snaps": [],
"history": {
IBM Storage Ceph 257
"epoch_created": 114,
"epoch_pool_created": 82,
"last_epoch_started": 2402,
"last_interval_started": 2401,
"last_epoch_clean": 2402,
"last_interval_clean": 2401,
"last_epoch_split": 114,
"last_epoch_marked_full": 0,
"same_up_since": 2401,
"same_interval_since": 2401,
"same_primary_since": 2086,
"last_scrub": "0'0",
"last_scrub_stamp": "2022-10-17T01:32:03.763988+0000",
"last_deep_scrub": "0'0",
"last_deep_scrub_stamp": "2022-10-17T01:32:03.763988+0000",
"last_clean_scrub_stamp": "2022-10-17T01:32:03.763988+0000",
"prior_readable_until_ub": 0
....
},
"stats": {
"version": "0'0",
"reported_seq": "2989",
"reported_epoch": "2449",
"state": "active+clean",
"last_fresh": "2022-10-18T05:16:59.401080+0000",
"last_change": "2022-10-17T01:32:03.764162+0000",
"last_active": "2022-10-18T05:16:59.401080+0000",
Reference
For more details about the snapshot trimming settings, see Object Storage Daemon (OSD) configuration options.
Placement Group creating state
When you create a pool, it will create the number of placement groups you specified. Ceph will echo creating when it is creating one or more placement groups. Once
they are created, the OSDs that are part of a placement group’s Acting Set will peer. Once peering is complete, the placement group status should be active+clean,
which means a Ceph client can begin writing to the placement group.
Figure 1. Creating PGs
Placement group peering state
When Ceph is Peering a placement group, Ceph is bringing the OSDs that store the replicas of the placement group into agreement about the state of the objects and
metadata in the placement group. When Ceph completes peering, this means that the OSDs that store the placement group agree about the current state of the placement
group. However, completion of the peering process does NOT mean that each replica has the latest contents.
Authoritative History
Ceph will NOT acknowledge a write operation to a client, until all OSDs of the acting set persist the write operation. This practice ensures that at least one member of the
acting set will have a record of every acknowledged write operation since the last successful peering operation.
With an accurate record of each acknowledged write operation, Ceph can construct and disseminate a new authoritative history of the placement group. A complete, and
fully ordered set of operations that, if performed, would bring an OSD’s copy of a placement group up to date.
Placement group active state
Once Ceph completes the peering process, a placement group may become active. The active state means that the data in the placement group is generally available
in the primary placement group and the replicas for read and write operations.
Placement Group clean state
258 IBM Storage Ceph
When a placement group is in the clean state, the primary OSD and the replica OSDs have successfully peered and there are no stray replicas for the placement group.
Ceph replicated all objects in the placement group the correct number of times.
Placement Group degraded state
When a client writes an object to the primary OSD, the primary OSD is responsible for writing the replicas to the replica OSDs. After the primary OSD writes the object to
storage, the placement group will remain in a degraded state until the primary OSD has received an acknowledgement from the replica OSDs that Ceph created the
replica objects successfully.
The reason a placement group can be active+degraded is that an OSD may be active even though it doesn’t hold all of the objects yet. If an OSD goes down, Ceph
marks each placement group assigned to the OSD as degraded. The Ceph OSDs must peer again when the Ceph OSD comes back online. However, a client can still write
a new object to a degraded placement group if it is active.
If an OSD is down and the degraded condition persists, Ceph may mark the down OSD as out of the cluster and remap the data from the down OSD to another OSD. The
time between being marked down and being marked out is controlled by mon_osd_down_out_interval, which is set to 600 seconds by default.
A placement group can also be degraded, because Ceph cannot find one or more objects that Ceph thinks should be in the placement group. While you cannot read or
write to unfound objects, you can still access all of the other objects in the degraded placement group.
For example, if there are nine OSDs in a three way replica pool. If OSD number 9 goes down, the PGs assigned to OSD 9 goes into a degraded state. If OSD 9 does not
recover, it goes out of the storage cluster and the storage cluster rebalances. In that scenario, the PGs are degraded and then recover to an active state.
Placement Group recovering state
Ceph was designed for fault-tolerance at a scale where hardware and software problems are ongoing. When an OSD goes down, its contents may fall behind the current
state of other replicas in the placement groups. When the OSD is back up, the contents of the placement groups must be updated to reflect the current state. During that
time period, the OSD may reflect a recovering state.
Recovery is not always trivial, because a hardware failure might cause a cascading failure of multiple Ceph OSDs. For example, a network switch for a rack or cabinet may
fail, which can cause the OSDs of a number of host machines to fall behind the current state of the storage cluster. Each one of the OSDs must recover once the fault is
resolved.
Ceph provides a number of settings to balance the resource contention between new service requests and the need to recover data objects and restore the placement
groups to the current state. The osd recovery delay start setting allows an OSD to restart, re-peer and even process some replay requests before starting the
recovery process. The osd recovery
threads setting limits the number of threads for the recovery process, by default one thread. The osd recovery thread timeout sets a thread timeout, because
multiple Ceph OSDs can fail, restart and re-peer at staggered rates. The osd recovery max
active setting limits the number of recovery requests a Ceph OSD works on simultaneously to prevent the Ceph OSD from failing to serve. The osd recovery max
chunk setting limits the size of the recovered data chunks to prevent network congestion.
Back fill state
When a new Ceph OSD joins the storage cluster, CRUSH will reassign placement groups from OSDs in the cluster to the newly added Ceph OSD. Forcing the new OSD to
accept the reassigned placement groups immediately can put excessive load on the new Ceph OSD. Backfilling the OSD with the placement groups allows this process to
begin in the background. Once backfilling is complete, the new OSD will begin serving requests when it is ready.
During the backfill operations, you might see one of several states:
backfill_wait indicates that a backfill operation is pending, but isn’t underway yet
backfill indicates that a backfill operation is underway
backfill_too_full indicates that a backfill operation was requested, but couldn’t be completed due to insufficient storage capacity.
When a placement group cannot be backfilled, it can be considered incomplete.
Ceph provides a number of settings to manage the load spike associated with reassigning placement groups to a Ceph OSD, especially a new Ceph OSD. By default,
osd_max_backfills sets the maximum number of concurrent backfills to or from a Ceph OSD to 10. The osd backfill
full ratio enables a Ceph OSD to refuse a backfill request if the OSD is approaching its full ratio, by default 85%. If an OSD refuses a backfill request, the osd
backfill retry
interval enables an OSD to retry the request, by default after 10 seconds. OSDs can also set osd backfill scan min and osd backfill scan max to manage
scan intervals, by default 64 and 512.
For some workloads, it is beneficial to avoid regular recovery entirely and use backfill instead. Since backfilling occurs in the background, this allows I/O to proceed on the
objects in the OSD. You can force a backfill rather than a recovery by setting the osd_min_pg_log_entries option to 1, and setting the osd_max_pg_log_entries
option to 2. Contact your IBM Support account team for details on when this situation is appropriate for your workload.
Placement Group remapped state
When the Acting Set that services a placement group changes, the data migrates from the old acting set to the new acting set. It may take some time for a new primary
OSD to service requests. So it may ask the old primary to continue to service requests until the placement group migration is complete. Once data migration completes,
the mapping uses the primary OSD of the new acting set.
IBM Storage Ceph 259
Placement Group stale state
While Ceph uses heartbeats to ensure that hosts and daemons are running, the ceph-osd daemons may also get into a stuck state where they are not reporting
statistics in a timely manner. For example, a temporary network fault. By default, OSD daemons report their placement group, up through boot and failure statistics every
half a second, that is, 0.5, which is more frequent than the heartbeat thresholds. If the Primary OSD of a placement group’s acting set fails to report to the monitor or if
other OSDs have reported the primary OSD down, the monitors will mark the placement group stale.
When you start the storage cluster, it is common to see the stale state until the peering process completes. After the storage cluster has been running for a while, seeing
placement groups in the stale state indicates that the primary OSD for those placement groups is down or not reporting placement group statistics to the monitor.
Placement Group misplaced state
There are some temporary backfilling scenarios where a PG gets mapped temporarily to an OSD. When that temporary situation should no longer be the case, the PGs
might still reside in the temporary location and not in the proper location. In which case, they are said to be misplaced. That’s because the correct number of extra
copies actually exist, but one or more copies is in the wrong place.
For example, there are 3 OSDs: 0,1,2 and all PGs map to some permutation of those three. If you add another OSD (OSD 3), some PGs will now map to OSD 3 instead of
one of the others. However, until OSD 3 is backfilled, the PG will have a temporary mapping allowing it to continue to serve I/O from the old mapping. During that time, the
PG is misplaced, because it has a temporary mapping, but not degraded, since there are three copies.
Syntax
pg 1.5: up=acting: [0,1,2]
ADD_OSD_3
pg 1.5: up: [0,3,1] acting: [0,1,2]
[0,1,2] is a temporary mapping, so the up set is not equal to the acting set and the PG is misplaced but not degraded since [0,1,2] is still three copies.
Example
pg 1.5: up=acting: [0,3,1]
OSD 3 is now backfilled and the temporary mapping is removed, not degraded and not misplaced.
Placement Group incomplete state
A PG goes into a incomplete state when there is incomplete content and peering fails, that is, when there are no complete OSDs which are current enough to perform
recovery.
Lets say OSD 1, 2, and 3 are the acting OSD set and it switches to OSD 1, 4, and 3, then osd.1 will request a temporary acting set of OSD 1, 2, and 3 while backfilling 4.
During this time, if OSD 1, 2, and 3 all go down, osd.4 will be the only one left which might not have fully backfilled all the data. At this time, the PG will go incomplete
indicating that there are no complete OSDs which are current enough to perform recovery.
Alternately, if osd.4 is not involved and the acting set is simply OSD 1, 2, and 3 when OSD 1, 2, and 3 go down, the PG would likely go stale indicating that the mons
have not heard anything on that PG since the acting set changed. The reason being there are no OSDs left to notify the new OSDs.
Identifying stuck Placement Groups
A placement group is not necessarily problematic just because it is not in a active+clean state. Generally, Ceph’s ability to self repair might not be working when
placement groups get stuck. The stuck states include:
Unclean: Placement groups contain objects that are not replicated the desired number of times. They should be recovering.
Inactive: Placement groups cannot process reads or writes because they are waiting for an OSD with the most up-to-date data to come back up.
Stale: Placement groups are in an unknown state, because the OSDs that host them have not reported to the monitor cluster in a while, and can be configured with
the mon osd report
timeout setting.
Prerequisites
A running IBM Storage Ceph cluster.
Root-level access to the node.
Procedure
To identify stuck placement groups, execute the following:
Syntax
ceph pg dump_stuck {inactive|unclean|stale|undersized|degraded [inactive|unclean|stale|undersized|degraded...]} {INT}
260 IBM Storage Ceph
Example
[ceph: root@host01 /]# ceph pg dump_stuck stale
OK
Finding object’s location
The Ceph client retrieves the latest cluster map and the CRUSH algorithm calculates how to map the object to a placement group, and then calculates how to assign the
placement group to an OSD dynamically.
Prerequisites
A running IBM Storage Ceph cluster.
Root-level access to the node.
Procedure
To find the object location, all you need is the object name and the pool name:
Syntax
ceph osd map POOL_NAME OBJECT_NAME
Example
[ceph: root@host01 /]# ceph osd map mypool myobject
Stretch clusters for Ceph storage
As a storage administrator, you can configure stretch clusters by entering stretch mode with 2-site clusters.
IBM Storage Ceph is capable of withstanding the loss of Ceph OSDs because of its network and cluster, which are equally reliable with failures randomly distributed across
the CRUSH map. If a number of OSDs is shut down, the remaining OSDs and monitors still manage to operate.
However, this might not be the best solution for some stretched cluster configurations where a significant part of the Ceph cluster can use only a single network
component. The example is a single cluster located in multiple data centers, for which the user wants to sustain a loss of a full data center.
The standard configuration is with two data centers. Other configurations are in clouds or availability zones. Each site holds two copies of the data, therefore, the
replication size is four. The third site should have a tiebreaker monitor, this can be a virtual machine or high-latency compared to the main sites. This monitor chooses one
of the sites to restore data if the network connection fails and both data centers remain active.
Note: The standard Ceph configuration survives many failures of the network or data centers and it never compromises data consistency. If you restore enough Ceph
servers following a failure, it recovers. Ceph maintains availability if you lose a data center, but can still form a quorum of monitors and have all the data available with
enough copies to satisfy pools’ min_size, or CRUSH rules that replicate again to meet the size.
Note: There are no additional steps to power down a stretch cluster. See Powering down and rebooting the cluster for more information.
Stretch cluster failures
IBM Storage Ceph never compromises on data integrity and consistency. If there is a network failure or a loss of nodes and the services can still be restored, Ceph returns
to normal functionality on its own.
However, there are situations where you lose data availability even if you have enough servers available to meet Ceph’s consistency and sizing constraints, or where you
unexpectedly do not meet the constraints.
First important type of failure is caused by inconsistent networks. If there is a network split, Ceph might be unable to mark OSD as down to remove it from the acting
placement group (PG) sets despite the primary OSD being unable to replicate data. When this happens, the I/O is not permitted because Ceph cannot meet its durability
guarantees.
The second important category of failures is when it appears that you have data replicated across data enters, but the constraints are not sufficient to guarantee this. For
example, you might have data centers A and B, and the CRUSH rule targets three copies and places a copy in each data center with a min_size of 2. The PG might go
active with two copies in site A and no copies in site B, which means that if you lose site A, you lose the data and Ceph cannot operate on it. This situation is difficult to
avoid with standard CRUSH rules.
Stretch mode
Setting crush location for daemons
Entering stretch mode
Adding OSD hosts in stretch mode
Stretch mode
To configure stretch clusters, you must enter the stretch mode. When stretch mode is enabled, the Ceph OSDs only take PGs as active when they peer across data centers,
or whichever other CRUSH bucket type you specified, assuming both are active. Pools increase in size from the default three to four, with two copies on each site.
In stretch mode, Ceph OSDs are only allowed to connect to monitors within the same data center. New monitors are not allowed to join the cluster without specified
location.
IBM Storage Ceph 261
If all the OSDs and monitors from a data center become inaccessible at once, the surviving data center will enter a degraded stretch mode. This issues a warning,
reduces the min_size to 1, and allows the cluster to reach an active state with the data from the remaining site.
Note: The degraded state also triggers warnings that the pools are too small, because the pool size does not get changed. However, a special stretch mode flag prevents
the OSDs from creating extra copies in the remaining data center, therefore it still keeps 2 copies.
When the missing data center becomes accesible again, the cluster enters recovery stretch mode. This changes the warning and allows peering, but still requires only
the OSDs from the data center, which was up the whole time.
When all PGs are in a known state and are not degraded or incomplete, the cluster goes back to the regular stretch mode, ends the warning, and restores min_size to its
starting value 2. The cluster again requires both sites to peer, not only the site that stayed up the whole time, therefore you can fail over to the other site, if necessary.
Stretch mode limitations
It is not possible to exit from stretch mode once it is entered.
You cannot use erasure-coded pools with clusters in stretch mode. You can neither enter the stretch mode with erasure-coded pools, nor create an erasure-coded
pool when the stretch mode is active.
Stretch mode with no more than two sites is supported.
The weights of the two sites should be the same. If they are not, you receive the following error:
Example
[ceph: root@host01 /]# ceph mon enable_stretch_mode host05 stretch_rule datacenter
Error EINVAL: the 2 datacenter instances in the cluster have differing weights 25947 and 15728 but stretch mode
currently requires they be the same!
To achieve same weights on both sites, the Ceph OSDs deployed in the two sites should be of equal size, that is, storage capacity in the first site is equivalent to storage
capacity in the second site.
While it is not enforced, you should run two Ceph monitors on each site and a tiebreaker, for a total of five. This is because OSDs can only connect to monitors in
their own site when in stretch mode.
You have to create your own CRUSH rule, which provides two copies on each site, which totals to four on both sites.
You cannot enable stretch mode if you have existing pools with non-default size or min_size.
Because the cluster runs with min_size 1 when degraded, you should only use stretch mode with all-flash OSDs. This minimizes the time needed to recover once
connectivity is restored, and minimizes the potential for data loss.
Reference
Troubleshooting clusters in stretch mode
Setting crush location for daemons
Before you enter the stretch mode, you need to prepare the cluster by setting the crush location to the daemons in the cluster. There are two ways to do this:
Bootstrap the cluster through a service configuration file, where the locations are added to the hosts as part of deployment.
Set the locations manually through ceph osd crush add-bucket and ceph
osd crush move commands after the cluster is deployed.
Setting the crush location during bootstrap
Bootstrap the cluster through a service configuration file, where the locations are added to the hosts as part of deployment.
Setting the crush location for daemons manually
Set the locations manually through ceph osd crush add-bucket and ceph osd crush move commands after the cluster is deployed.
Setting the crush location during bootstrap
Bootstrap the cluster through a service configuration file, where the locations are added to the hosts as part of deployment.
Prerequisites
Root-level access to the nodes.
Procedure
1. If you are bootstrapping your new storage cluster, you can create the service configuration .yaml file that adds the nodes to the IBM Storage Ceph cluster and also
sets specific labels for where the services should run.
Example
service_type: host
addr: host01
hostname: host01
262 IBM Storage Ceph
location:
root: default
datacenter: DC1
labels:
- osd
- mon
- mgr
--service_type: host
addr: host02
hostname: host02
location:
datacenter: DC1
labels:
- osd
- mon
--service_type: host
addr: host03
hostname: host03
location:
datacenter: DC1
labels:
- osd
- mds
- rgw
--service_type: host
addr: host04
hostname: host04
location:
root: default
datacenter: DC2
labels:
- osd
- mon
- mgr
--service_type: host
addr: host05
hostname: host05
location:
datacenter: DC2
labels:
- osd
- mon
--service_type: host
addr: host06
hostname: host06
location:
datacenter: DC2
labels:
- osd
- mds
- rgw
--service_type: host
addr: host07
hostname: host07
labels:
- mon
--service_type: mon
placement:
label: "mon"
--service_id: cephfs
placement:
label: "mds"
--service_type: mgr
service_name: mgr
placement:
label: "mgr"
--service_type: osd
service_id: all-available-devices
service_name: osd.all-available-devices
placement:
label: "osd"
spec:
data_devices:
all: true
--service_type: rgw
service_id: objectgw
service_name: rgw.objectgw
placement:
count: 2
label: "rgw"
spec:
rgw_frontend_port: 8080
2. Bootstrap the storage cluster with the --apply-spec option.
IBM Storage Ceph 263
Syntax
cephadm bootstrap --apply-spec CONFIGURATION_FILE_NAME --mon-ip MONITOR_IP_ADDRESS --ssh-private-key PRIVATE_KEY --sshpublic-key PUBLIC_KEY --registry-url REGISTRY_URL --registry-username USER_NAME --registry-password PASSWORD
Example
[root@host01 ~]# cephadm bootstrap --apply-spec initial-config.yaml --mon-ip 10.10.128.68 --ssh-private-key
/home/ceph/.ssh/id_rsa --ssh-public-key /home/ceph/.ssh/id_rsa.pub --registry-url registry.redhat.io --registry-username
myuser1 --registry-password mypassword1
Important: You can use different command options with the cephadm
bootstrap command. However, always include the --apply-spec option to use the service configuration file and configure the host locations.
Reference
For more information about Ceph bootstrapping and different cephadm bootstrap command options, see Bootstrapping a new storage cluster
Setting the crush location for daemons manually
Set the locations manually through ceph osd crush add-bucket and ceph osd crush move commands after the cluster is deployed.
Prerequisites
Root-level access to the nodes.
Procedure
1. Add two buckets to which you plan to set the location of your nontiebreaker monitors to the CRUSH map, specifying the bucket type as datacenter.
Syntax
ceph osd crush add-bucket BUCKET_NAME BUCKET_TYPE
Example
[ceph: root@host01 /]# ceph osd crush add-bucket DC1 datacenter
[ceph: root@host01 /]# ceph osd crush add-bucket DC2 datacenter
2. Move the buckets under root=default.
Syntax
ceph osd crush move BUCKET_NAME root=default
Example
[ceph: root@host01 /]# ceph osd crush move DC1 root=default
[ceph: root@host01 /]# ceph osd crush move DC2 root=default
3. Move the OSD hosts according to the required CRUSH placement.
Syntax
ceph osd crush move HOST datacenter=DATACENTER
Example
[ceph: root@host01 /]# ceph osd crush move host01 datacenter=DC1
Entering stretch mode
The new stretch mode is designed to handle two sites. There is a lower risk of component availability outages with 2-site clusters.
Prerequisites
Root-level access to the nodes.
The crush location is set to the hosts.
Procedure
1. Set the location of each monitor, matching your CRUSH map:
Syntax
ceph mon set_location HOST datacenter=DATACENTER
Example
264 IBM Storage Ceph
[ceph: root@host01 /]# ceph mon set_location host01 datacenter=DC1
[ceph: root@host01 /]# ceph mon set_location host02 datacenter=DC1
[ceph: root@host01 /]# ceph mon set_location host04 datacenter=DC2
[ceph: root@host01 /]# ceph mon set_location host05 datacenter=DC2
[ceph: root@host01 /]# ceph mon set_location host07 datacenter=DC3
2. Generate a CRUSH rule which places two copies on each data center:
Syntax
ceph osd getcrushmap > COMPILED_CRUSHMAP_FILENAME crushtool -d COMPILED_CRUSHMAP_FILENAME -o DECOMPILED_CRUSHMAP_FILENAME
Example
[ceph: root@host01 /]# ceph osd getcrushmap > crush.map.bin
[ceph: root@host01 /]# crushtool -d crush.map.bin -o crush.map.txt
3. Edit the decompiled CRUSH map file to add a new rule:
Example
rule stretch_rule {
id 1
type replicated
min_size 1
max_size 10
step take DC1 <2>
step chooseleaf firstn 2 type host
step emit
step take DC2
step chooseleaf firstn 2 type host
step emit
}
The rule `id` has to be unique. In this example, there is only one more rule with `id 0`, thereby the `id 1` is used,
however you might need to use a different rule ID depending on the number of existing rules. In this example, there are
two data center buckets named `DC1` and `DC2`.
Note: This rule makes the cluster have read-affinity towards data center DC1. Therefore, all the reads or writes happen through Ceph OSDs placed in DC1. If this is
not desirable, and reads or writes are to be distributed evenly across the zones, the crush rule is the following:
Example
rule stretch_rule {
id 1
type replicated
min_size 1
max_size 10
step take default
step choose firstn 0 type datacenter
step chooseleaf firstn 2 type host
step emit
}
In this rule, the data center is selected randomly and automatically. See CRUSH rules for more information on firstn and indep options.
4. Inject the CRUSH map to make the rule available to the cluster:
Syntax
crushtool -c DECOMPILED_CRUSHMAP_FILENAME -o COMPILED_CRUSHMAP_FILENAME
ceph osd setcrushmap -i COMPILED_CRUSHMAP_FILENAME
Example
[ceph: root@host01 /]# crushtool -c crush.map.txt -o crush2.map.bin
[ceph: root@host01 /]# ceph osd setcrushmap -i crush2.map.bin
5. If you do not run the monitors in connectivity mode, set the election strategy to connectivity:
Example
[ceph: root@host01 /]# ceph mon set election_strategy connectivity
6. Enter stretch mode by setting the location of the tiebreaker monitor to split across the data centers:
Syntax
ceph mon set_location HOST datacenter=DATACENTER
ceph mon enable_stretch_mode HOST stretch_rule datacenter
Example
[ceph: root@host01 /]# ceph mon set_location host07 datacenter=DC3
[ceph: root@host01 /]# ceph mon enable_stretch_mode host07 stretch_rule datacenter
In this example the monitor mon.host07 is the tiebreaker.
Important: The location of the tiebreaker monitor should differ from the data centers to which you previously set the non-tiebreaker monitors. In the example
above, it is data center DC3.
Important: Do not add this data center to the CRUSH map as it results in the following error when you try to enter stretch mode:
Error EINVAL: there are 3 datacenters in the cluster but stretch mode currently only works with 2!
IBM Storage Ceph 265
Note: If you are writing your own tooling for deploying Ceph, you can use a new --set-crush-location option when booting monitors, instead of running the
ceph mon set_location command. This option accepts only a single bucket=location pair, for example ceph-mon --set-crush-location
'datacenter=DC1', which must match the bucket type you specified when running the enable_stretch_mode command.
7. Verify that the stretch mode is enabled successfully:
Example
[ceph: root@host01 /]# ceph osd dump
epoch 361
fsid 1234ab78-1234-11ed-b1b1-de456ef0a89d
created 2023-01-16T05:47:28.4827170000
modified 2023-01-17T17:36:50.0661830000
flags sortbitwise,recovery_deletes,purged_snapdirs,pglog_hardlimit
crush_version 31
full_ratio 0.95
backfillfull_ratio 0.92
nearfull_ratio 0.85
require_min_compat_client luminous
min_compat_client luminous
require_osd_release quincy
stretch_mode_enabled true
stretch_bucket_count 2
degraded_stretch_mode 0
recovering_stretch_mode 0
stretch_mode_bucket 8
The stretch_mode_enabled should be set to true. You can also see the number of stretch buckets, stretch mode buckets, and if the stretch mode is degraded
or recovering.
8. Verify that the monitors are in an appropriate locations:
Example
[ceph: root@host01 /]# ceph mon dump
epoch 19
fsid 1234ab78-1234-11ed-b1b1-de456ef0a89d
last_changed 2023-01-17T04:12:05.7094750000
created 2023-01-16T05:47:25.6316840000
min_mon_release 16 (pacific)
election_strategy: 3
stretch_mode_enabled 1
tiebreaker_mon host07
disallowed_leaders host07
0: [v2:132.224.169.63:3300/0,v1:132.224.169.63:6789/0] mon.host07; crush_location {datacenter=DC3}
1: [v2:220.141.179.34:3300/0,v1:220.141.179.34:6789/0] mon.host04; crush_location {datacenter=DC2}
2: [v2:40.90.220.224:3300/0,v1:40.90.220.224:6789/0] mon.host01; crush_location {datacenter=DC1}
3: [v2:60.140.141.144:3300/0,v1:60.140.141.144:6789/0] mon.host02; crush_location {datacenter=DC1}
4: [v2:186.184.61.92:3300/0,v1:186.184.61.92:6789/0] mon.host05; crush_location {datacenter=DC2}
dumped monmap epoch 19
You can also see which monitor is the tiebreaker, and the monitor election strategy.
Reference
Configuring monitor election strategy
Adding OSD hosts in stretch mode
You can add Ceph OSDs in the stretch mode. The procedure is similar to the addition of the OSD hosts on a cluster where stretch mode is not enabled.
Prerequisites
A running IBM Storage Ceph cluster.
Stretch mode in enabled on a cluster.
Root-level access to the nodes.
Procedure
1. List the available devices to deploy OSDs:
Syntax
ceph orch device ls [--hostname=HOST_1 HOST_2] [--wide] [--refresh]
Example
[ceph: root@host01 /]# ceph orch device ls
2. Deploy the OSDs on specific hosts or on all the available devices:
Create an OSD from a specific device on a specific host:
266 IBM Storage Ceph
Syntax
ceph orch daemon add osd HOST:DEVICE_PATH
Example
[ceph: root@host01 /]# ceph orch daemon add osd host03:/dev/sdb
Deploy OSDs on any available and unused devices:
Important: This command creates collocated WAL and DB devices. If you want to create non-collocated devices, do not use this command.
Example
[ceph: root@host01 /]# ceph orch apply osd --all-available-devices
3. Move the OSD hosts under the CRUSH bucket:
Syntax
ceph osd crush move HOST datacenter=DATACENTER
Example
[ceph: root@host01 /]# ceph osd crush move host03 datacenter=DC1
[ceph: root@host01 /]# ceph osd crush move host06 datacenter=DC2
Note: Ensure you add the same topology nodes on both sites. Issues might arise if hosts are added only on one site.
Reference
Adding OSDs
Override Ceph behavior
As a storage administrator, you need to understand how to use overrides for the IBM Storage Ceph cluster to change Ceph options during runtime.
Setting and unsetting override options
Override use cases
Setting and unsetting override options
You can set and unset Ceph options to override Ceph’s default behavior.
Prerequisites
A running IBM Storage Ceph cluster.
Root-level access to the node.
Procedure
1. To override Ceph’s default behavior, use the ceph osd set command and the behavior you wish to override:
Syntax
ceph osd set FLAG
Once you set the behavior, ceph health will reflect the override(s) that you have set for the cluster.
Example
[ceph: root@host01 /]# ceph osd set noout
2. To cease overriding Ceph’s default behavior, use the ceph osd unset command and the override you wish to cease.
Syntax
ceph osd unset FLAG
Example
[ceph: root@host01 /]# ceph osd unset noout
Flag
Description
noin
Prevents OSDs from being treated as in the cluster.
noout
Prevents OSDs from being treated as out of the cluster.
noup
Prevents OSDs from being treated as up and running.
nodown
Prevents OSDs from being treated as down.
full
Makes a cluster appear to have reached its full_ratio, and thereby prevents write operations.
IBM Storage Ceph 267
Flag
pause
Ceph stops processing read and write operations, but will not affect OSD in, out, up or down statuses.
Description
nobackfill
Ceph prevents new backfill operations.
norebalance
Ceph prevents new rebalancing operations.
norecover
Ceph prevents new recovery operations.
noscrub
Ceph prevents new scrubbing operations.
nodeep-scrub Ceph prevents new deep scrubbing operations.
notieragent
Ceph disables the process that is looking for cold/dirty objects to flush and evict.
Override use cases
noin: Commonly used with noout to address flapping OSDs.
noout: If the mon osd report timeout is exceeded and an OSD has not reported to the monitor, the OSD will get marked out. If this happens erroneously, you
can set noout to prevent the OSD(s) from getting marked out while you troubleshoot the issue.
noup: Commonly used with nodown to address flapping OSDs.
nodown: Networking issues may interrupt Ceph heartbeat processes, and an OSD may be up but still get marked down. You can set nodown to prevent OSDs from
getting marked down while troubleshooting the issue.
full: If a cluster is reaching its full_ratio, you can pre-emptively set the cluster to full and expand capacity.
Note: Setting the cluster to full will prevent write operations.
pause: If you need to troubleshoot a running Ceph cluster without clients reading and writing data, you can set the cluster to pause to prevent client operations.
nobackfill: If you need to take an OSD or node down temporarily, for example, upgrading daemons, you can set nobackfill so that Ceph will not backfill while
the OSDs is down.
norecover: If you need to replace an OSD disk and don’t want the PGs to recover to another OSD while you are hotswapping disks, you can set norecover to
prevent the other OSDs from copying a new set of PGs to other OSDs.
noscrub and nodeep-scrubb: If you want to prevent scrubbing for example, to reduce overhead during high loads, recovery, backfilling, and rebalancing you can
set noscrub and/or nodeep-scrub to prevent the cluster from scrubbing OSDs.
notieragent: If you want to stop the tier agent process from finding cold objects to flush to the backing storage tier, you may set notieragent.
Ceph user management
As a storage administrator, you can manage the Ceph user base by providing authentication, and access control to objects in the IBM Storage Ceph cluster.
Figure 1. OSD States
Note: Cephadm manages the client keyrings for the IBM Storage Ceph cluster as long as the clients are within the scope of Cephadm. Do not modify the keyrings that are
managed by the Cephadm, unless otherwise specified. If troubleshooting is needed, see Cephadm troubleshooting.
Ceph user management background
When Ceph runs with authentication and authorization that is enabled, you must specify a username. If a username is not specified, Ceph uses the client.admin
administrative user as the default username.
Managing Ceph users
Ceph user management background
When Ceph runs with authentication and authorization that is enabled, you must specify a username. If a username is not specified, Ceph uses the client.admin
administrative user as the default username.
The CEPH_ARGS environment variable may also be used to avoid re-entry of the username and secret.
Irrespective of the type of Ceph client, for example, block device, object store, file system, native API, or the Ceph command line, Ceph stores all data as objects within
pools. Ceph users must have access to pools in order to read and write data. Also, administrative Ceph users must have permissions to run Ceph’s administrative
commands.
For more information on configuring the use of authentication, see Configuring.
268 IBM Storage Ceph
The following concepts can help you understand Ceph user management:
Storage cluster users
Authorization capabilities
Pool
Namespace
Storage cluster users
A user of the IBM Storage Ceph cluster is either an individual or as an application. Creating users allows you to control who can access the storage cluster, its pools, and
the data within those pools.
Ceph has the notion of a type of user. For the purposes of user management, the type is always client. Ceph identifies users in period (.) delimited form consisting of
the user type and the user ID. For example, TYPE.ID, client.admin, or client.user1. The reason for user typing is that Ceph Monitors, and OSDs also use the Cephx
protocol, but they are not clients. Distinguishing the user type helps to distinguish between client users and other users, streamlining access control, user monitoring, and
traceability.
Sometimes Ceph’s user type may seem confusing because the Ceph command line allows you to specify a user with or without the type, depending upon the commandline usage. If you specify --user or --id, you can omit the type. So client.user1 can be entered, instead, as user1. If you specify --name or -n, you must specify the
type and name, such as client.user1. Use the type and name, wherever possible.
Note: An IBM Storage Ceph cluster user is not the same as a Ceph Object Gateway user. The object gateway uses an IBM Storage Ceph cluster user to communicate
between the gateway daemon and the storage cluster. The gateway also has its own user management functionality for its end users.
Authorization capabilities
Ceph uses the term "capabilities" (caps) to describe authorizing an authenticated user to exercise the functionality of the Ceph Monitors and OSDs. Capabilities can also
restrict access to data within a pool or a namespace within a pool. A Ceph administrative user sets a user’s capabilities when creating or updating a user. Capability syntax
follows the form:
DAEMON_TYPE 'allow CAPABILITY' [DAEMON_TYPE 'allow CAPABILITY']
Monitor Caps
Monitor capabilities include r, w, x, allow profile CAP, and profile rbd.
For example:
mon 'allow rwx'
mon 'allow profile osd'
OSD Caps
OSD capabilities include r, w, x, class-read, class-write, profile osd, profile rbd, and profile rbd-read-only. OSD capabilities also allow
for pool and namespace settings.
For example:
osd 'allow CAPABILITY' [pool=POOL_NAME] [namespace=NAMESPACE_NAME]
Note: The Ceph Object Gateway daemon (radosgw) is a client of the Ceph storage cluster, so it is not represented as a Ceph storage cluster daemon type.
Table 1 describes each capability.
Table 1. OSD Caps capabilities
Capability
allow
Precedes access settings for a daemon.
Description
r
Gives the user read access. Required with monitors to retrieve the CRUSH map.
w
Gives the user write access to objects.
x
Gives the user the capability to call class methods (that is, both read and write) and to conduct auth operations on monitors.
class-read
Gives the user the capability to call class read methods. Subset of x.
class-write
Gives the user the capability to call class write methods. Subset of x.
*
Gives the user read, write and execute permissions for a particular daemon or pool, and the ability to execute admin commands.
profile osd
Gives a user permissions to connect as an OSD to other OSDs or monitors. Conferred on OSDs to enable OSDs to handle replication
heartbeat traffic and status reporting.
profile
bootstrap-osd
profile rbd
Gives a user permissions to bootstrap an OSD so that they have permissions to add keys when bootstrapping an OSD.
profile rbdread-only
Gives a user read-only access to the Ceph Block Devices.
Gives a user read/write access to the Ceph Block Devices.
Pool
A pool defines a storage strategy for Ceph clients, and acts as a logical partition for that strategy.
In Ceph deployments, it is common to create a pool to support different types of use cases. For example, cloud volumes or images, object storage, hot storage, cold
storage, and so on. When deploying Ceph as a back end for OpenStack, a typical deployment would have pools for volumes, images, backups and virtual machines, and
users such as client.glance, client.cinder, and so on.
Namespace
Objects within a pool can be associated to a namespace, a logical group of objects within the pool. A user’s access to a pool can be associated with a namespace such that
reads and writes by the user take place only within the namespace. Objects written to a namespace within the pool can only be accessed by users who have access to the
namespace.
Note: Currently, namespaces are only useful for applications that are written on top of librados. Ceph clients such as block device and object storage do not currently
support this feature.
IBM Storage Ceph 269
The rationale for namespaces is that pools can be a computationally expensive method of segregating data by use case because each pool creates a set of placement
groups that get mapped to OSDs. If multiple pools use the same CRUSH hierarchy and ruleset, OSD performance can degrade as load increases.
For example, a pool should have approximately 100 placement groups per OSD. So an exemplary cluster with 1000 OSDs would have 100,000 placement groups for one
pool. Each pool mapped to the same CRUSH hierarchy and ruleset would create another 100,000 placement groups in the exemplary cluster. By contrast, writing an
object to a namespace simply associates the namespace to the object name without the computational overhead of a separate pool. Rather than creating a separate pool
for a user or set of users, you may use a namespace.
Note: Only available using librados at this time.
Managing Ceph users
As a storage administrator, you can manage Ceph users by creating, modifying, deleting, and importing users.
A Ceph client user can be either individuals or applications, which use Ceph clients to interact with the IBM Storage Ceph cluster daemons.
Listing Ceph users
Displaying Ceph user information
Adding Ceph user
Modifying Ceph user
Deleting Ceph user
Printing Ceph user key
Listing Ceph users
You can list the users in the storage cluster using the command-line interface.
Prerequisites
A running IBM Storage Ceph cluster.
Root-level access to the node.
Procedure
To list the users in the storage cluster, execute the following:
Example
[ceph: root@host01 /]# ceph auth list
installed auth entries:
osd.10
osd.11
osd.9
key: AQBW7U5gqOsEExAAg/CxSwZ/gSh8iOsDV3iQOA==
caps: [mgr] allow profile osd
caps: [mon] allow profile osd
caps: [osd] allow *
key: AQBX7U5gtj/JIhAAPsLBNG+SfC2eMVEFkl3vfA==
caps: [mgr] allow profile osd
caps: [mon] allow profile osd
caps: [osd] allow *
key: AQBV7U5g1XDULhAAKo2tw6ZhH1jki5aVui2v7g==
caps: [mgr] allow profile osd
caps: [mon] allow profile osd
caps: [osd] allow *
client.admin
key: AQADYEtgFfD3ExAAwH+C1qO7MSLE4TWRfD2g6g==
caps: [mds] allow *
caps: [mgr] allow *
caps: [mon] allow *
caps: [osd] allow *
client.bootstrap-mds
key: AQAHYEtgpbkANBAANqoFlvzEXFwD8oB0w3TF4Q==
caps: [mon] allow profile bootstrap-mds
client.bootstrap-mgr
key: AQAHYEtg3dcANBAAVQf6brq3sxTSrCrPe0pKVQ==
caps: [mon] allow profile bootstrap-mgr
client.bootstrap-osd
key: AQAHYEtgD/QANBAATS9DuP3DbxEl86MTyKEmdw==
caps: [mon] allow profile bootstrap-osd
client.bootstrap-rbd
key: AQAHYEtgjxEBNBAANho25V9tWNNvIKnHknW59A==
caps: [mon] allow profile bootstrap-rbd
client.bootstrap-rbd-mirror
key: AQAHYEtgdE8BNBAAr6rLYxZci0b2hoIgH9GXYw==
caps: [mon] allow profile bootstrap-rbd-mirror
client.bootstrap-rgw
key: AQAHYEtgwGkBNBAAuRzI4WSrnowBhZxr2XtTFg==
caps: [mon] allow profile bootstrap-rgw
client.crash.host04
key: AQCQYEtgz8lGGhAAy5bJS8VH9fMdxuAZ3CqX5Q==
270 IBM Storage Ceph
caps: [mgr] profile crash
caps: [mon] profile crash
client.crash.host02
key: AQDuYUtgqgfdOhAAsyX+Mo35M+HFpURGad7nJA==
caps: [mgr] profile crash
caps: [mon] profile crash
client.crash.host03
key: AQB98E5g5jHZAxAAklWSvmDsh2JaL5G7FvMrrA==
caps: [mgr] profile crash
caps: [mon] profile crash
client.rgw.test_realm.test_zone.host01.hgbvnq
key: AQD5RE9gAQKdCRAAJzxDwD/dJObbInp9J95sXw==
caps: [mgr] allow rw
caps: [mon] allow *
caps: [osd] allow rwx tag rgw *=*
client.rgw.test_realm.test_zone.host02.yqqilm
key: AQD0RE9gkxA4ExAAFXp3pLJWdIhsyTe2ZR6Ilw==
caps: [mgr] allow rw
caps: [mon] allow *
caps: [osd] allow rwx tag rgw *=*
mgr.host01.hdhzwn
key: AQAEYEtg3lhIBxAAmHodoIpdvnxK0llWF80ltQ==
caps: [mds] allow *
caps: [mon] profile mgr
caps: [osd] allow *
mgr.host02.eobuuv
key: AQAn6U5gzUuiABAA2Fed+jPM1xwb4XDYtrQxaQ==
caps: [mds] allow *
caps: [mon] profile mgr
caps: [osd] allow *
mgr.host03.wquwpj
key: AQAd6U5gIzWsLBAAbOKUKZlUcAVe9kBLfajMKw==
caps: [mds] allow *
caps: [mon] profile mgr
caps: [osd] allow *
Note: The TYPE.ID notation for users applies such that osd.0 is a user of type osd and its ID is 0. client.admin is a user of type client and its ID is admin, that is,
the default client.admin user. Note also that each entry has a key: VALUE entry, and one or more caps: entries.
You may use the -o FILE_NAME option with ceph auth list to save the output to a file.
Displaying Ceph user information
You can display a Ceph’s user information using the command-line interface.
Prerequisites
A running IBM Storage Ceph cluster.
Root-level access to the node.
Procedure
1. To retrieve a specific user, key and capabilities, execute the following:
Syntax
ceph auth export TYPE.ID
Example
[ceph: root@host01 /]# ceph auth export mgr.host02.eobuuv
2. You can also use the -o FILE_NAME option.
Syntax
ceph auth export TYPE.ID -o FILE_NAME
Example
[ceph: root@host01 /]# ceph auth export osd.9 -o filename
export auth(key=AQBV7U5g1XDULhAAKo2tw6ZhH1jki5aVui2v7g==)
The auth export command is identical to auth get, but also prints out the internal auid, which isn’t relevant to end users.
Adding Ceph user
Adding a user creates a username, that is, TYPE.ID, a secret key and any capabilities included in the command you use to create the user.
A user’s key enables the user to authenticate with the Ceph storage cluster. The user’s capabilities authorize the user to read, write, or execute on Ceph monitors (mon),
Ceph OSDs (osd) or Ceph Metadata Servers (mds).
There are a few ways to add a user:
IBM Storage Ceph 271
ceph auth add: This command is the canonical way to add a user. It will create the user, generate a key and add any specified capabilities.
ceph auth get-or-create: This command is often the most convenient way to create a user, because it returns a keyfile format with the user name (in
brackets) and the key. If the user already exists, this command simply returns the user name and key in the keyfile format. You may use the -o FILE_NAME option
to save the output to a file.
ceph auth get-or-create-key: This command is a convenient way to create a user and return the user’s key only. This is useful for clients that need the key
only, for example, libvirt. If the user already exists, this command simply returns the key. You may use the -o FILE_NAME option to save the output to a file.
When creating client users, you may create a user with no capabilities. A user with no capabilities is useless beyond mere authentication, because the client cannot
retrieve the cluster map from the monitor. However, you can create a user with no capabilities if you wish to defer adding capabilities later using the ceph auth caps
command.
A typical user has at least read capabilities on the Ceph monitor and read and write capability on Ceph OSDs. Additionally, a user’s OSD permissions are often restricted to
accessing a particular pool. :
[ceph: root@host01 /]# ceph auth add client.john mon 'allow r' osd 'allow rw pool=mypool'
[ceph: root@host01 /]# ceph auth get-or-create client.paul mon 'allow r' osd 'allow rw pool=mypool'
[ceph: root@host01 /]# ceph auth get-or-create client.george mon 'allow r' osd 'allow rw pool=mypool' -o george.keyring
[ceph: root@host01 /]# ceph auth get-or-create-key client.ringo mon 'allow r' osd 'allow rw pool=mypool' -o ringo.key
Important: If you provide a user with capabilities to OSDs, but you DO NOT restrict access to particular pools, the user will have access to ALL pools in the cluster.
Modifying Ceph user
The ceph auth caps command allows you to specify a user and change the user’s capabilities.
Prerequisites
A running IBM Storage Ceph cluster.
Root-level access to the node.
Procedure
1. To add capabilities, use the form:
Syntax
ceph auth caps USERTYPE.USERID DAEMON 'allow [r|w|x|*|...] [pool=POOL_NAME] [namespace=NAMESPACE_NAME]'
Example
[ceph: root@host01 /]# ceph auth caps client.john mon 'allow r' osd 'allow rw pool=mypool'
[ceph: root@host01 /]# ceph auth caps client.paul mon 'allow rw' osd 'allow rwx pool=mypool'
[ceph: root@host01 /]# ceph auth caps client.brian-manager mon 'allow *' osd 'allow *'
2. To remove a capability, you may reset the capability. If you want the user to have no access to a particular daemon that was previously set, specify an empty string:
Example
[ceph: root@host01 /]# ceph auth caps client.ringo mon ' ' osd ' '
Reference
For more information about capabilities, see Ceph user management background.
Deleting Ceph user
You can delete a user from the Ceph storage cluster using the command-line interface.
Prerequisites
A running IBM Storage Ceph cluster.
Root-level access to the node.
Procedure
1. To delete a user, use ceph auth del:
Syntax
ceph auth del TYPE.ID
Example
[ceph: root@host01 /]# ceph auth del osd.6
272 IBM Storage Ceph
Printing Ceph user key
You can display a Ceph user’s key information using the command-line interface.
Prerequisites
A running IBM Storage Ceph cluster.
Root-level access to the node.
Procedure
Print a user’s authentication key to standard output:
Syntax
ceph auth print-key TYPE.ID
Example
[ceph: root@host01 /]# ceph auth print-key osd.6
AQBQ7U5gAry3JRAA3NoPrqBBThpFMcRL6Sr+5w==[ceph: root@host01 /]#
Using ceph-volume utility
Use the ceph-volume utility to prepare, list, create, activate, deactivate, batch, trigger, zap, and migrate Ceph OSDs.
The ceph-volume utility is a single-purpose command-line tool to deploy logical volumes as OSDs. It uses a plugin-type framework to deploy OSDs with different device
technologies. The ceph-volume utility follows a similar workflow of the ceph-disk utility for deploying OSDs, with a predictable, and robust way of preparing, activating,
and starting OSDs. Currently, the ceph-volume utility only supports the lvm plug-in, with the plan to support others technologies in the future.
Important: The ceph-disk command is deprecated.
Ceph volume lvm plugin
Why does ceph-volume replace ceph-disk?
Previous versions of Ceph used the ceph-disk utility to prepare, activate, and create OSDs. Starting with IBM Storage Ceph 5, ceph-disk is replaced by the
ceph-volume utility that aims to be a single purpose command-line tool to deploy logical volumes as OSDs, while maintaining a similar API to ceph-disk when
preparing, activating, and creating OSDs.
Preparing OSDs
Listing devices
Activating OSDs
Deactivating OSDs
Creating Ceph OSDs
Migrating BlueFS data
Expanding BlueFS DB device
Expand the storage of BlueStore File System.
Using batch mode
Zapping data
Ceph volume lvm plugin
By making use of LVM tags, the lvm sub-command is able to store and re-discover by querying devices associated with OSDs so they can be activated. This includes
support for lvm-based technologies like dm-cache as well.
When using ceph-volume, the use of dm-cache is transparent, and treats dm-cache like a logical volume. The performance gains and losses when using dm-cache will
depend on the specific workload. Generally, random and sequential reads will see an increase in performance at smaller block sizes. While random and sequential writes
will see a decrease in performance at larger block sizes.
To use the LVM plugin, add lvm as a subcommand to the ceph-volume command within the cephadm shell:
[ceph: root@host01 /]# ceph-volume lvm
Following are the lvm subcommands:
prepare - Format an LVM device and associate it with an OSD.
activate - Discover and mount the LVM device associated with an OSD ID and start the Ceph OSD.
list - List logical volumes and devices associated with Ceph.
batch - Automatically size devices for multi-OSD provisioning with minimal interaction.
deactivate - Deactivate OSDs.
create - Create a new OSD from an LVM device.
IBM Storage Ceph 273
trigger - A systemd helper to activate an OSD.
zap - Removes all data and filesystems from a logical volume or partition.
migrate - Migrate BlueFS data from to another LVM device.
new-wal - Allocate new WAL volume for the OSD at specified logical volume.
new-db - Allocate new DB volume for the OSD at specified logical volume.
Note: Using the create subcommand combines the prepare and activate subcommands into one subcommand.
Reference
For more information, see .Creating Ceph OSDs.
Why does ceph-volume replace ceph-disk?
Previous versions of Ceph used the ceph-disk utility to prepare, activate, and create OSDs. Starting with IBM Storage Ceph 5, ceph-disk is replaced by the cephvolume utility that aims to be a single purpose command-line tool to deploy logical volumes as OSDs, while maintaining a similar API to ceph-disk when preparing,
activating, and creating OSDs.
How does ceph-volume work?
The ceph-volume is a modular tool that currently supports two ways of provisioning hardware devices, legacy ceph-disk devices and LVM (Logical Volume Manager)
devices. The ceph-volume lvm command uses the LVM tags to store information about devices specific to Ceph and its relationship with OSDs. It uses these tags to
later re-discover and query devices associated with OSDS so that it can activate them. It supports technologies based on LVM and dm-cache as well.
The ceph-volume utility uses dm-cache transparently and treats it as a logical volume. You might consider the performance gains and losses when using dm-cache,
depending on the specific workload you are handling. Generally, the performance of random and sequential read operations increases at smaller block sizes; while the
performance of random and sequential write operations decreases at larger block sizes. Using ceph-volume does not introduce any significant performance penalties.
Important: The ceph-disk utility is deprecated.
Note: The ceph-volume simple command can handle legacy ceph-disk devices, if these devices are still in use.
How does ceph-disk work?
The ceph-disk utility was required to support many different types of init systems, such as upstart or sysvinit, while being able to discover devices. For this reason,
ceph-disk concentrates only on GUID Partition Table (GPT) partitions. Specifically on GPT GUIDs that label devices in a unique way to answer questions like:
Is this device a journal?
Is this device an encrypted data partition?
Was the device left partially prepared?
To solve these questions, ceph-disk uses UDEV rules to match the GUIDs.
What are disadvantages of using ceph-disk?
Using the UDEV rules to call ceph-disk can lead to a back-and-forth between the ceph-disk systemd unit and the ceph-disk executable. The process is very
unreliable and time consuming and can cause OSDs to not come up at all during the boot process of a node. Moreover, it is hard to debug, or even replicate these problems
given the asynchronous behavior of UDEV.
Because ceph-disk works with GPT partitions exclusively, it cannot support other technologies, such as Logical Volume Manager (LVM) volumes, or similar device
mapper devices.
To ensure the GPT partitions work correctly with the device discovery workflow, ceph-disk requires a large number of special flags to be used. In addition, these
partitions require devices to be exclusively owned by Ceph.
Preparing OSDs
The prepare subcommand prepares an OSD back-end object store and consumes logical volumes (LV) for both the OSD data and journal. It does not modify the logical
volumes, except for adding some extra metadata tags using LVM. These tags make volumes easier to discover, and they also identify the volumes as part of the Ceph
Storage Cluster and the roles of those volumes in the storage cluster.
The BlueStore OSD backend supports the following configurations:
A block device, a block.wal device, and a block.db device
A block device and a block.wal device
A block device and a block.db device
A single block device
The prepare subcommand accepts a whole device or partition, or a logical volume for block.
274 IBM Storage Ceph
Prerequisites
Root-level access to the OSD nodes.
Optionally, create logical volumes. If you provide a path to a physical device, the subcommand turns the device into a logical volume. This approach is simpler, but
you cannot configure or change the way the logical volume is created.
Procedure
1. Extract the Ceph keyring:
Syntax
ceph auth get client.ID -o ceph.client.ID.keyring
Example
[ceph: root@host01 /]# ceph-volume lvm prepare --bluestore --data example_vg/data_lv
2. Prepare the LVM volumes:
Syntax
ceph-volume lvm prepare --bluestore --data VOLUME_GROUP/LOGICAL_VOLUME
Example
[ceph: root@host01 /]# ceph-volume lvm prepare --bluestore --data example_vg/data_lv
a. Optionally, if you want to use a separate device for RocksDB, specify the --block.db and --block.wal options:
Syntax
ceph-volume lvm prepare --bluestore --block.db --block.wal --data VOLUME_GROUP/LOGICAL_VOLUME
Example
[ceph: root@host01 /]# ceph-volume lvm prepare --bluestore --block.db --block.wal --data example_vg/data_lv
b. Optionally, to encrypt data, use the --dmcrypt flag:
Syntax
ceph-volume lvm prepare --bluestore --dmcrypt --data VOLUME_GROUP/LOGICAL_VOLUME
Example
[ceph: root@host01 /]# ceph-volume lvm prepare --bluestore --dmcrypt --data example_vg/data_lv
References
For more information, see:
Activating OSDs
Creating Ceph OSDs
Listing devices
You can use the ceph-volume lvm list subcommand to list logical volumes and devices that are associated with a Ceph cluster. Ensure they contain enough metadata
to allow for that discovery. The OSD ID associated with the devices groups the output. For logical volumes, the devices key is populated with the physical devices that
are associated with the logical volume.
Sometimes, the output of the ceph -s command shows the following error message:
1 devices have fault light turned on
In such cases, you can list the devices with ceph device ls-lights command that gives the details about the lights on the devices. Based on the information, you can turn
off the lights on the devices.
Prerequisites
A running IBM Storage Ceph cluster.
Root-level access to the Ceph OSD node.
Procedure
1. List the devices in the Ceph cluster:
Example
[ceph: root@host01 /]# ceph-volume lvm list
====== osd.6 =======
IBM Storage Ceph 275
[block]
/dev/ceph-83909f70-95e9-4273-880e-5851612cbe53/osd-block-7ce687d9-07e7-4f8f-a34e-d1b0efb89920
block device
d1b0efb89920
block uuid
cephx lockbox secret
cluster fsid
cluster name
crush device class
encrypted
osd fsid
osd id
osdspec affinity
type
vdo
devices
/dev/ceph-83909f70-95e9-4273-880e-5851612cbe53/osd-block-7ce687d9-07e7-4f8f-a34e4d7gzX-Nzxp-UUG0-bNxQ-Jacr-l0mP-IPD8cX
1ca9f6a8-d036-11ec-8263-fa163ee967ad
ceph
None
0
7ce687d9-07e7-4f8f-a34e-d1b0efb89920
6
all-available-devices
block
0
/dev/vdc
2. List the devices in the storage cluster with the lights.
[ceph: root@host01 /]# ceph device ls-lights
{
}
"fault": [
"SEAGATE_ST12000NM002G_ZL2KTGCK0000C149"
],
"ident": []
3. Turn off the lights on the device.
ceph device light off DEVICE_NAME FAULT/INDENT --force
Example
[ceph: root@host01 /]# ceph device light off SEAGATE_ST12000NM002G_ZL2KTGCK0000C149 fault --force
Activating OSDs
The activation process for a Caph OSD enables a systemd unit at boot time, which allows the correct OSD identifier and its UUID to be enabled and mounted.
Prerequisites
A running IBM Storage Ceph cluster.
Root-level access to the Ceph OSD node.
Ceph OSDs prepared by the ceph-volume utility.
Procedure
1. Get the OSD ID and OSD FSID from an OSD node:
Example
[ceph: root@host01 /]# ceph-volume lvm list
2. Activate the OSD:
Syntax
ceph-volume lvm activate --bluestore OSD_ID OSD_FSID
Example
[ceph: root@host01 /]# ceph-volume lvm activate --bluestore 10 7ce687d9-07e7-4f8f-a34e-d1b0efb89920
To activate all OSDs that are prepared for activation, use the --all option:
Example
[ceph: root@host01 /]# ceph-volume lvm activate --all
3. Optionally, you can use the trigger subcommand. This command cannot be used directly, and it is used by systemd so that it proxies input to ceph-volume
lvm activate. This parses the metadata coming from systemd and startup, detecting the UUID and ID associated with an OSD.
Syntax
ceph-volume lvm trigger SYSTEMD_DATA
Here the SYSTEMD_DATA is in OSD_ID-OSD_FSID format.
Example
[ceph: root@host01 /]# ceph-volume lvm trigger 10 7ce687d9-07e7-4f8f-a34e-d1b0efb89920
276 IBM Storage Ceph
Reference
For more information, see:
Preparing OSDs
Creating Ceph OSDs
Deactivating OSDs
You can deactivate the Ceph OSDs using the ceph-volume lvm subcommand. This subcommand removes the volume groups and the logical volume.
Prerequisites
A running IBM Storage Ceph cluster.
Root-level access to the Ceph OSD node.
The Ceph OSDs are activated using the ceph-volume utility.
Procedure
1. Get the OSD ID from the OSD node:
[ceph: root@host01 /]# ceph-volume lvm list
2. Deactivate the OSD:
Syntax
ceph-volume lvm deactivate OSD_ID
Example
[ceph: root@host01 /]# ceph-volume lvm deactivate 16
Reference
For more information, see:
Activating OSDs
Preparing OSDs
Creating Ceph OSDs
Creating Ceph OSDs
The create subcommand calls the prepare subcommand, and then calls the activate subcommand.
Prerequisites
A running IBM Storage Ceph cluster.
Root-level access to the Ceph OSD nodes.
Note: If you prefer to have more control over the creation process, you can use the prepare and activate subcommands separately to create the OSD instead of using
create. You can use the two subcommands to gradually introduce new OSDs into a storage cluster, while avoiding having to rebalance large amounts of data. Both
approaches work the same way, except that using the create subcommand causes the OSD to become up and in immediately after completion.
Procedure
1. To create a new OSD:
Syntax
ceph-volume lvm create --bluestore --data VOLUME_GROUP/LOGICAL_VOLUME
Example
[root@osd ~]# ceph-volume lvm create --bluestore --data example_vg/data_lv
Reference
For more information, see:
IBM Storage Ceph 277
Preparing OSDs
Activating OSDs
Migrating BlueFS data
You can migrate the BlueStore file system (BlueFS) data, that is the RocksDB data, from the source volume to the target volume using the migrate LVM subcommand.
The source volume, except the main one, is removed on success.
LVM volumes are primarily for the target only.
The new volumes are attached to the OSD, replacing one of the source drives.
Following are the placement rules for the LVM volumes:
If source list has DB or WAL volume, then the target device replaces it.
If source list has slow volume only, then explicit allocation using the new-db or new-wal command is needed.
The new-db and new-wal commands attaches the given logical volume to the given OSD as a DB or a WAL volume respectively.
Prerequisites
A running IBM Storage Ceph cluster.
Root-level access to the Ceph OSD node.
Ceph OSDs prepared by the ceph-volume utility.
Volume groups and Logical volumes are created.
Procedure
1. Log in the cephadm shell:
Example
[root@host01 ~]# cephadm shell
2. Stop the OSD to which you have to add the DB or the WAL device:
Example
[ceph: root@host01 /]# ceph orch daemon stop osd.1
3. Mount the new devices to the container:
Example
[root@host01 ~]# cephadm shell --mount /var/lib/ceph/72436d46-ca06-11ec-9809-ac1f6b5635ee/osd.1:/var/lib/ceph/osd/ceph-1
4. Attach the given logical volume to OSD as a DB/WAL device:
Note: This command fails if the OSD has an attached DB.
Syntax
ceph-volume lvm new-db --osd-id OSD_ID --osd-fsid OSD_FSID --target VOLUME_GROUP_NAME/LOGICAL_VOLUME_NAME
Example
[ceph: root@host01 /]# ceph-volume lvm new-db --osd-id 1 --osd-fsid 7ce687d9-07e7-4f8f-a34e-d1b0efb89921 --target
vgname/new_db
[ceph: root@host01 /]# ceph-volume lvm new-wal --osd-id 1 --osd-fsid 7ce687d9-07e7-4f8f-a34e-d1b0efb89921 --target
vgname/new_wal
5. You can migrate BlueFS data in the following ways:
Move BlueFS data from main device to LV that is already attached as DB:
Syntax
ceph-volume lvm migrate --osd-id OSD_ID --osd-fsid OSD_UUID --from data --target
VOLUME_GROUP_NAME/LOGICAL_VOLUME_NAME
Example
[ceph: root@host01 /]# ceph-volume lvm migrate --osd-id 1 --osd-fsid 0263644D-0BF1-4D6D-BC34-28BD98AE3BC8 --from data
--target vgname/db
Move BlueFS data from shared main device to LV which shall be attached as a new DB:
Syntax
ceph-volume lvm migrate --osd-id OSD_ID --osd-fsid OSD_UUID --from data --target
VOLUME_GROUP_NAME/LOGICAL_VOLUME_NAME
278 IBM Storage Ceph
Example
[ceph: root@host01 /]# ceph-volume lvm migrate --osd-id 1 --osd-fsid 0263644D-0BF1-4D6D-BC34-28BD98AE3BC8 --from data
--target vgname/new_db
Move BlueFS data from DB device to new LV, and replace the DB device:
Syntax
ceph-volume lvm migrate --osd-id OSD_ID --osd-fsid OSD_UUID --from db --target VOLUME_GROUP_NAME/LOGICAL_VOLUME_NAME
Example
[ceph: root@host01 /]# ceph-volume lvm migrate --osd-id 1 --osd-fsid 0263644D-0BF1-4D6D-BC34-28BD98AE3BC8 --from db -target vgname/new_db
Move BlueFS data from main and DB devices to new LV, and replace the DB device:
Syntax
ceph-volume lvm migrate --osd-id OSD_ID --osd-fsid OSD_UUID --from data db --target
VOLUME_GROUP_NAME/LOGICAL_VOLUME_NAME
Example
[ceph: root@host01 /]# ceph-volume lvm migrate --osd-id 1 --osd-fsid 0263644D-0BF1-4D6D-BC34-28BD98AE3BC8 --from data
db --target vgname/new_db
Move BlueFS data from main, DB, and WAL devices to new LV, remove the WAL device, and replace the DB device:
Syntax
ceph-volume lvm migrate --osd-id OSD_ID --osd-fsid OSD_UUID --from data db wal --target
VOLUME_GROUP_NAME/LOGICAL_VOLUME_NAME
Example
[ceph: root@host01 /]# ceph-volume lvm migrate --osd-id 1 --osd-fsid 0263644D-0BF1-4D6D-BC34-28BD98AE3BC8 --from data
db --target vgname/new_db
Move BlueFS data from main, DB, and WAL devices to the main device, remove the WAL and DB devices:
Syntax
ceph-volume lvm migrate --osd-id OSD_ID --osd-fsid OSD_UUID --from db wal --target
VOLUME_GROUP_NAME/LOGICAL_VOLUME_NAME
Example
[ceph: root@host01 /]# ceph-volume lvm migrate --osd-id 1 --osd-fsid 0263644D-0BF1-4D6D-BC34-28BD98AE3BC8 --from db
wal --target vgname/data
Expanding BlueFS DB device
Expand the storage of BlueStore File System.
About this task
You can expand the storage of the BlueStore file system (BlueFS) data that is the RocksDB data of ceph-volume created OSDs with the ceph-bluestore tool.
Before you begin
A running IBM Storage Ceph cluster.
Ceph OSDs are prepared by the ceph-volume utility.
Volume groups and Logical volumes are created.
Note: Run these steps on the host where the OSD is deployed.
Procedure
1. Optional: Inside the cephadm shell, list the devices in the IBM Storage Ceph cluster.
ceph-volume lvm list
For example,
[ceph: root@host01 /]# ceph-volume lvm list
====== osd.3 =======
[db]
/dev/db-test/db1
block device
block uuid
/dev/test/lv1
N5zoix-FePe-uExe-UngY-D9YG-BMs0-1tTDyB
IBM Storage Ceph 279
cephx lockbox secret
cluster fsid
cluster name
crush device class
db device
db uuid
encrypted
osd fsid
osd id
osdspec affinity
type
vdo
devices
[block]
1a6112da-ed05-11ee-bacd-525400565cda
ceph
/dev/db-test/db1
1TUaDY-3mEt-fReP-cyB2-JyZ1-oUPa-hKPfo6
0
94ff742c-7bfd-4fb5-8dc4-843d10ac6731
3
None
db
0
/dev/vdh
/dev/test/lv1
block device
block uuid
cephx lockbox secret
cluster fsid
cluster name
crush device class
db device
db uuid
encrypted
osd fsid
osd id
osdspec affinity
type
vdo
devices
/dev/test/lv1
N5zoix-FePe-uExe-UngY-D9YG-BMs0-1tTDyB
1a6112da-ed05-11ee-bacd-525400565cda
ceph
/dev/db-test/db1
1TUaDY-3mEt-fReP-cyB2-JyZ1-oUPa-hKPfo6
0
94ff742c-7bfd-4fb5-8dc4-843d10ac6731
3
None
block
0
/dev/vdg
2. Get the volume group information.
vgs
For example,
[root@host01 ~]# vgs
VG
db-test
test
#PV #LV #SN Attr
VSize
VFree
1
1
0 wz--n- <200.00g <160.00g
1
1
0 wz--n- <200.00g <170.00g
3. Stop the Ceph OSD service.
systemctl stop SERVICE_ID
For example,
[root@host01 ~]# systemctl stop host01a6112da-ed05-11ee-bacd-525400565cda@osd.3.service
4. Resize, shrink, and expand the logical volumes.
lvresize -l 100%FREE PATH_OF_DB_DEVICE
For example,
[root@host01 ~]# lvresize -l 100%FREE /dev/db-test/db1
Size of logical volume db-test/db1 changed from 40.00 GiB (10240 extents) to <160.00 GiB (40959 extents).
Logical volume db-test/db1 successfully resized.
5. Launch the cephadm shell.
cephadm shell -m /var/lib/ceph/CLUSTER_FSID/osd.OSD_ID:/var/lib/ceph/osd/ceph-OSD_ID:z
For example,
[root@host01 ~]# cephadm shell -m /var/lib/ceph/1a6112da-ed05-11ee-bacd-525400565cda/osd.3:/var/lib/ceph/osd/ceph-3:z
The ceph-bluestore-tool needs to access the BlueStore data from within the cephadm shell container, so it must be bind-mounted. Use the -m option to make
the BlueStore data available.
6. Check the size of the Rocks DB before expansion.
ceph-bluestore-tool show-label --path OSD_DIRECTORY_PATH
For example,
[ceph: root@host01 /]# ceph-bluestore-tool show-label --path /var/lib/ceph/osd/ceph-3/
inferring bluefs devices from bluestore path
{
"/var/lib/ceph/osd/ceph-3/block": {
"osd_uuid": "94ff742c-7bfd-4fb5-8dc4-843d10ac6731",
"size": 32212254720,
"btime": "2024-04-03T08:34:12.742848+0000",
"description": "main",
"bfm_blocks": "7864320",
"bfm_blocks_per_key": "128",
"bfm_bytes_per_block": "4096",
"bfm_size": "32212254720",
"bluefs": "1",
"ceph_fsid": "1a6112da-ed05-11ee-bacd-525400565cda",
"ceph_version_when_created": "ceph version 19.0.0-2493-gd82c9aa1 (d82c9aa17f09785fe698d262f9601d87bb79f962) squid
(dev)",
"created_at": "2024-04-03T08:34:15.637253Z",
280 IBM Storage Ceph
"elastic_shared_blobs": "1",
"kv_backend": "rocksdb",
"magic": "ceph osd volume v026",
"mkfs_done": "yes",
"osd_key": "AQCEFA1m9xuwABAAwKEHkASVbgB1GVt5jYC2Sg==",
"osdspec_affinity": "None",
"ready": "ready",
"require_osd_release": "19",
"whoami": "3"
}
},
"/var/lib/ceph/osd/ceph-3/block.db": {
"osd_uuid": "94ff742c-7bfd-4fb5-8dc4-843d10ac6731",
"size": 40794497536,
"btime": "2024-04-03T08:34:12.748816+0000",
"description": "bluefs db"
}
7. Expand the BlueStore device.
ceph-bluestore-tool bluefs-bdev-expand --path OSD_DIRECTORY_PATH
For example,
[ceph: root@host01 /]# ceph-bluestore-tool bluefs-bdev-expand --path /var/lib/ceph/osd/ceph-3/
inferring bluefs devices from bluestore path
1 : device size 0x27ffbfe000 : using 0x2300000(35 MiB)
2 : device size 0x780000000 : using 0x52000(328 KiB)
Expanding DB/WAL...
1 : expanding to 0x171794497536
1 : size label updated to 171794497536
8. Verify the block.db is expanded.
ceph-bluestore-tool show-label --path OSD_DIRECTORY_PATH
For example,
[ceph: root@host01 /]# ceph-bluestore-tool show-label --path /var/lib/ceph/osd/ceph-3/
inferring bluefs devices from bluestore path
{
"/var/lib/ceph/osd/ceph-3/block": {
"osd_uuid": "94ff742c-7bfd-4fb5-8dc4-843d10ac6731",
"size": 32212254720,
"btime": "2024-04-03T08:34:12.742848+0000",
"description": "main",
"bfm_blocks": "7864320",
"bfm_blocks_per_key": "128",
"bfm_bytes_per_block": "4096",
"bfm_size": "32212254720",
"bluefs": "1",
"ceph_fsid": "1a6112da-ed05-11ee-bacd-525400565cda",
"ceph_version_when_created": "ceph version 19.0.0-2493-gd82c9aa1 (d82c9aa17f09785fe698d262f9601d87bb79f962) squid
(dev)",
"created_at": "2024-04-03T08:34:15.637253Z",
"elastic_shared_blobs": "1",
"kv_backend": "rocksdb",
"magic": "ceph osd volume v026",
"mkfs_done": "yes",
"osd_key": "AQCEFA1m9xuwABAAwKEHkASVbgB1GVt5jYC2Sg==",
"osdspec_affinity": "None",
"ready": "ready",
"require_osd_release": "19",
"whoami": "3"
},
"/var/lib/ceph/osd/ceph-3/block.db": {
"osd_uuid": "94ff742c-7bfd-4fb5-8dc4-843d10ac6731",
"size": 171794497536,
"btime": "2024-04-03T08:34:12.748816+0000",
"description": "bluefs db"
}
9. Exit the shell and restart the OSD.
For example,
[root@host01 ~]# systemctl start host01a6112da-ed05-11ee-bacd-525400565cda@osd.3.service
osd.3
host01
running (15s)
0s ago 13m
46.9M
4096M 19.0.0-2493-gd82c9aa1
3714003597ec 02150b3b6877
Using batch mode
The batch subcommand automates the creation of multiple OSDs when single devices are provided.
The ceph-volume command decides the best method to use to create the OSDs, based on drive type. Ceph OSD optimization depends on the available devices:
If all devices are traditional hard drives, batch creates one OSD per device.
If all devices are solid state drives, batch creates two OSDs per device.
IBM Storage Ceph 281
If there is a mix of traditional hard drives and solid state drives, batch uses the traditional hard drives for data, and creates the largest possible journal (block.db)
on the solid state drive.
Note: The batch subcommand does not support the creation of a separate logical volume for the write-ahead-log (block.wal) device.
Prerequisites
A running IBM Storage Ceph cluster.
Root-level access to the Ceph OSD nodes.
Procedure
1. To create OSDs on several drives:
Syntax
ceph-volume lvm batch --bluestore PATH_TO_DEVICE [PATH_TO_DEVICE]
Example
[ceph: root@host01 /]# ceph-volume lvm batch --bluestore /dev/sda /dev/sdb /dev/nvme0n1
Reference
For more information, see Creating Ceph OSDs.
Zapping data
The zap subcommand removes all data and file systems from a logical volume or partition.
You can use the zap subcommand to zap logical volumes, partitions, or raw devices that are used by Ceph OSDs for reuse. Any file system present on the given logical
volume or partition are removed and all data is purged.
Optionally, you can use the --destroy flag for complete removal of a logical volume, partition, or the physical device.
Prerequisites
A running IBM Storage Ceph cluster.
Root-level access to the Ceph OSD node.
Procedure
Zap the logical volume:
Syntax
ceph-volume lvm zap VOLUME_GROUP_NAME/LOGICAL_VOLUME_NAME [--destroy]
Example
[ceph: root@host01 /]# ceph-volume lvm zap osd-vg/data-lv
Zap the partition:
Syntax
ceph-volume lvm zap DEVICE_PATH_PARTITION [--destroy]
Example
[ceph: root@host01 /]# ceph-volume lvm zap /dev/sdc1
Zap the raw device:
Syntax
ceph-volume lvm zap DEVICE_PATH --destroy
Example
[ceph: root@host01 /]# ceph-volume lvm zap /dev/sdc --destroy
Purge multiple devices with the OSD ID:
Syntax
ceph-volume lvm zap --destroy --osd-id OSD_ID
Example
282 IBM Storage Ceph
[ceph: root@host01 /]# ceph-volume lvm zap --destroy --osd-id 16
Note: All the relative devices are zapped.
Purge OSDs with the FSID:
Syntax
ceph-volume lvm zap --destroy --osd-fsid OSD_FSID
Example
[ceph: root@host01 /]# ceph-volume lvm zap --destroy --osd-fsid 65d7b6b1-e41a-4a3c-b363-83ade63cb32b
Note: All the relative devices are zapped.
Ceph performance benchmark
As a storage administrator, you can benchmark performance of the IBM Storage Ceph cluster. The purpose of this section is to give Ceph administrators a basic
understanding of Ceph's native benchmarking tools. These tools will provide some insight into how the Ceph storage cluster is performing. This is not the definitive guide
to Ceph performance benchmarking, nor is it a guide on how to tune Ceph accordingly.
Performance baseline
Benchmarking Ceph performance
Benchmarking Ceph block performance
Benchmarking CephFS performance
Benchmark Ceph File System (CephFS) performance with the FIO tool.
Benchmarking Ceph Object Gateway performance
Benchmark Ceph Object Gateway performance with the s3cmd tool.
Performance baseline
The OSD, including the journal, disks and the network throughput should each have a performance baseline to compare against. You can identify potential tuning
opportunities by comparing the baseline performance data with the data from Ceph’s native tools. Red Hat Enterprise Linux has many built-in tools, along with a plethora
of open source community tools, available to help accomplish these tasks.
Reference
For more details about some of the available tools, see this Knowledge base article.
Benchmarking Ceph performance
Ceph includes the rados bench command to do performance benchmarking on a RADOS storage cluster. The command will execute a write test and two types of read
tests. The --no-cleanup option is important to use when testing both read and write performance. By default the rados bench command will delete the objects it has
written to the storage pool. Leaving behind these objects allows the two read tests to measure sequential and random read performance.
Note: Before running these performance tests, drop all the file system caches by running the following:
[ceph: root@host01 /]# echo 3 | sudo tee /proc/sys/vm/drop_caches && sudo sync
Prerequisites
A running IBM Storage Ceph cluster.
Root-level access to the node.
Procedure
1. Create a new storage pool:
Example
[ceph: root@host01 /]# ceph osd pool create testbench 100 100
2. Run a write test for 10 seconds to the newly created storage pool:
Example
[ceph: root@host01 /]# rados bench -p testbench 10 write --no-cleanup
Maintaining 16 concurrent writes of 4194304 bytes for up to 10 seconds or 0 objects
Object prefix: benchmark_data_cephn1.home.network_10510
sec Cur ops
started finished avg MB/s cur MB/s last lat
avg lat
0
0
0
0
0
0
0
1
16
16
0
0
0
0
2
16
16
0
0
0
0
3
16
16
0
0
0
0
IBM Storage Ceph 283
4
16
5
16
6
16
7
16
8
16
9
16
10
16
11
16
12
16
13
16
14
16
15
16
16
11
17
11
18
11
19
11
Total time run:
Total writes made:
Write size:
Bandwidth (MB/sec):
17
18
18
19
25
25
27
27
27
28
28
28
28
28
28
28
1
2
2
3
9
9
11
11
11
12
12
12
17
17
17
17
19.505620
28
4194304
5.742
0.998879
1.59849
1.33222
1.71239
4.49551
3.99636
4.39632
3.99685
3.66397
3.68975
3.42617
3.19785
4.24726
3.99751
3.77546
3.57683
1
4
0
2
24
0
4
0
0
1.33333
0
0
6.66667
0
0
0
3.19824
4.56163
6.90712
7.75362
9.65085
12.8124
12.5302
-
3.19824
3.87993
3.87993
4.889
6.71216
6.71216
7.18999
7.18999
7.18999
7.65853
7.65853
7.65853
9.27548
9.27548
9.27548
9.27548
Stddev Bandwidth:
5.4617
Max bandwidth (MB/sec): 24
Min bandwidth (MB/sec): 0
Average Latency:
10.4064
Stddev Latency:
3.80038
Max latency:
19.503
Min latency:
3.19824
3. Run a sequential read test for 10 seconds to the storage pool:
Example
[ceph: root@host01 /]# rados bench -p testbench 10 seq
sec Cur ops
started
0
0
0
Total time run:
Total reads made:
Read size:
Bandwidth (MB/sec):
finished
0
0.804869
28
4194304
139.153
Average Latency:
Max latency:
Min latency:
0.420841
0.706133
0.0816332
avg MB/s
0
cur MB/s
0
last lat
-
avg lat
0
4. Run a random read test for 10 seconds to the storage pool:
Example
[ceph: root@host01 /]# rados bench -p testbench 10 rand
sec Cur ops
started
0
0
0
1
16
46
2
16
81
3
16
120
4
15
157
5
16
206
6
16
253
7
16
287
8
16
325
9
16
362
10
16
405
Total time run:
Total reads made:
Read size:
Bandwidth (MB/sec):
finished avg MB/s
0
0
30
119.801
65
129.408
104
138.175
142
141.485
190
151.553
237
157.608
271
154.412
309
154.044
346
153.245
389
155.092
10.302229
405
4194304
157.248
Average Latency:
Max latency:
Min latency:
0.405976
1.00869
0.0378431
cur MB/s last lat
0
120 0.440184
140 0.577359
156 0.597435
152 0.683111
192 0.310578
188 0.0745175
136 0.792774
152 0.314254
148 0.355576
172
0.64734
avg lat
0
0.388125
0.417461
0.409318
0.419964
0.408343
0.387207
0.39043
0.39876
0.406032
0.398372
5. To increase the number of concurrent reads and writes, use the -t option, which the default is 16 threads. Also, the -b parameter can adjust the size of the object
being written. The default object size is 4 MB. A safe maximum object size is 16 MB. IBM recommends running multiple copies of these benchmark tests to different
pools. Doing this shows the changes in performance from multiple clients.
Add the --run-name LABEL option to control the names of the objects that get written during the benchmark test. Multiple rados bench commands might be
ran simultaneously by changing the --run-name label for each running command instance. This prevents potential I/O errors that can occur when multiple clients
are trying to access the same object and allows for different clients to access different objects. The --run-name option is also useful when trying to simulate a real
world workload.
Example
[ceph: root@host01 /]# rados bench -p testbench 10 write -t 4 --run-name client1
Maintaining 4 concurrent writes of 4194304 bytes for up to 10 seconds or 0 objects
Object prefix: benchmark_data_node1_12631
sec Cur ops
started finished avg MB/s cur MB/s last lat
avg lat
0
0
0
0
0
0
0
1
4
4
0
0
0
0
2
4
6
2
3.99099
4
1.94755
1.93361
3
4
8
4
5.32498
8
2.978
2.44034
4
4
8
4
3.99504
0
2.44034
284 IBM Storage Ceph
5
4
6
3
7
4
8
4
9
4
10
4
11
4
12
4
13
4
Total time run:
Total writes made:
Write size:
Bandwidth (MB/sec):
10
6
10
7
12
8
14
10
16
12
17
13
17
13
18
14
18
14
13.123548
18
4194304
5.486
4.79504
4.64471
4.55287
4.9821
5.31621
5.18488
4.71431
4.65486
4.29757
4
4
4
8
8
4
0
2
0
2.92419
3.02498
3.12204
2.55901
2.68769
2.11937
2.4836
-
2.4629
2.5432
2.61555
2.68396
2.68081
2.63763
2.63763
2.62662
2.62662
Stddev Bandwidth:
3.0991
Max bandwidth (MB/sec): 8
Min bandwidth (MB/sec): 0
Average Latency:
2.91578
Stddev Latency:
0.956993
Max latency:
5.72685
Min latency:
1.91967
6. Remove the data created by the rados bench command:
Example
[ceph: root@host01 /]# rados -p testbench cleanup
Benchmarking Ceph block performance
Ceph includes the rbd bench-write command to test sequential writes to the block device measuring throughput and latency. The default byte size is 4096, the default
number of I/O threads is 16, and the default total number of bytes to write is 1 GB. These defaults can be modified by the --io-size, --io-threads and --io-total
options respectively.
Prerequisites
A running IBM Storage Ceph cluster.
Root-level access to the node.
Run the write performance test against the block device:
Example
[root@host01 ~]# rbd bench --io-type write image01 --pool=testbench
bench-write io_size 4096 io_threads 16 bytes 1073741824 pattern seq
SEC
OPS
OPS/SEC
BYTES/SEC
2
11127
5479.59 22444382.79
3
11692
3901.91 15982220.33
4
12372
2953.34 12096895.42
5
12580
2300.05 9421008.60
6
13141
2101.80 8608975.15
7
13195
356.07 1458459.94
8
13820
390.35 1598876.60
9
14124
325.46 1333066.62
..
Reference
For more information about the rbd command, see Ceph Block Devices.
Benchmarking CephFS performance
Benchmark Ceph File System (CephFS) performance with the FIO tool.
About this task
You can use the FIO tool to benchmark Ceph File System (CephFS) performance. This tool can also be used to benchmark Ceph Block Device.
Before you begin
A running IBM Storage Ceph cluster.
Root-level access to the node.
FIO tool is installed on the nodes.
Block Device or the Ceph File System mounted on the node.
IBM Storage Ceph 285
Procedure
1. Navigate to the node or the application where the Block Device or the CephFS is mounted.
[root@host01 ~]# cd /mnt/ceph-block-device
[root@host01 ~]# cd /mnt/ceph-file-system
2. Run FIO command. Start the bs value from 4k and repeat in power of 2 increments (4k, 8k, 16k, 32k ... 128k... 512k, 1 m, 2 m, 4 m) and with different iodepth
settings. Also run tests at your expected workload operation size.
For example, 4 K tests with different iodepth values:
fio --name=randwrite --rw=randwrite --direct=1 --ioengine=libaio --bs=4k --iodepth=32 --size=5G --runtime=60 -group_reporting=1
For example, 8 K tests with different iodepth values:
fio --name=randwrite --rw=randwrite --direct=1 --ioengine=libaio --bs=8k --iodepth=32
group_reporting=1
--size=5G --runtime=60 --
Note: For more information on the usage of fio command, see the fio man page.
Benchmarking Ceph Object Gateway performance
Benchmark Ceph Object Gateway performance with the s3cmd tool.
About this task
You can use the s3cmd tool to benchmark Ceph Object Gateway performance.
Use get and put requests to determine the performance.
Before you begin
A running IBM Storage Ceph cluster.
Root-level access to the node.
s3cmd installed on the nodes.
Procedure
1. Upload a file and measure the speed. The time command measures the duration of upload.
time s3cmd put PATH_OF_SOURCE_FILE PATH_OF_DESTINATION_FILE
For example,
time s3cmd put /path-to-local-file s3://bucket-name/remote/file
Replace /path-to-local-file with the file you want to upload and s3://bucket-name/remote/file with the destination in your S3 bucket.
2. Download a file and measure the speed. The time command measures the duration of download.
time s3cmd get PATH_OF_DESTINATION_FILE DESTINATION_PATH
For example,
time s3cmd get s3://bucket-name/remote/file /path-to-local-destination
Replace s3://bucket-name/remote/file with the S3 object that you want to download and /path-to-local-destination with the local directory where
you want to save the file.
3. List all the objects in the specified bucket and measure response time.
time s3cmd ls s3://BUCKET_NAME
For example,
time s3cmd ls s3://bucket-name
4. Analyze the output to calculate upload and download speed. Measure response time reported by the time command.
Ceph performance counters
As a storage administrator, you can gather performance metrics of the IBM Storage Ceph cluster. The Ceph performance counters are a collection of internal infrastructure
metrics. The collection, aggregation, and graphing of this metric data can be done by an assortment of tools and can be useful for performance analytics.
Access to Ceph performance counters
Display Ceph performance counters
Dump Ceph performance counters
286 IBM Storage Ceph
Average count and sum
The avgcount is the number of operations within this range and the sum is the total latency in seconds. To know approximately how much latency there is per
operation, divide the sum by the avgcount.
Ceph Monitor metrics
Ceph OSD metrics
Ceph Object Gateway metrics
Access to Ceph performance counters
The performance counters are available through a socket interface for the Ceph Monitors and the OSDs. The socket file for each respective daemon is located under
/var/run/ceph, by default. The performance counters are grouped together into collection names. These collections names represent a subsystem or an instance of a
subsystem.
Here is the full list of the Monitor and the OSD collection name categories with a brief description for each:
Monitor Collection Name Categories
Cluster Metrics - Displays information about the storage cluster: Monitors, OSDs, Pools, and PGs
Level Database Metrics - Displays information about the back-end KeyValueStore database
Monitor Metrics - Displays general monitor information
Paxos Metrics - Displays information on cluster quorum management
Throttle Metrics - Displays the statistics on how the monitor is throttling
OSD Collection Name Categories
Write Back Throttle Metrics - Displays the statistics on how the write back throttle is tracking unflushed IO
Level Database Metrics - Displays information about the back-end KeyValueStore database
Objecter Metrics - Displays information on various object-based operations
Read and Write Operations Metrics - Displays information on various read and write operations
Recovery State Metrics - Displays latencies on various recovery states
OSD Throttle Metrics - Display the statistics on how the OSD is throttling
RADOS Gateway Collection Name Categories
Object Gateway Client Metrics - Displays statistics on GET and PUT requests
Objecter Metrics - Displays information on various object-based operations
Object Gateway Throttle Metrics - Display the statistics on how the OSD is throttling
Display Ceph performance counters
The ceph daemon DAEMON_NAME perf schema command outputs the available metrics. Each metric has an associated bit field value type.
Prerequisites
A running IBM Storage Ceph cluster.
Root-level access to the node.
Procedure
1. To view the metric’s schema:
Syntax
ceph daemon DAEMON_NAME perf schema
Note: Run the ceph daemon command from the node running the daemon.
2. Executing ceph daemon DAEMON_NAME perf schema command from the monitor node:
Example
[ceph: root@host01 /]# ceph daemon mon.host01 perf schema
3. Executing the ceph daemon DAEMON_NAME perf schema command from the OSD node:
Example
[ceph: root@host01 /]# ceph daemon osd.11 perf schema
IBM Storage Ceph 287
Table 1. The bit field value definitions
Bit
1
2
4
8
Meaning
Floating point value
Unsigned 64-bit integer value
Average (Sum + Count)
Counter
Each value will have bit 1 or 2 set to indicate the type, either a floating point or an integer value. When bit 4 is set, there will be two values to read, a sum and a count.
When bit 8 is set, the average for the previous interval would be the sum delta, since the previous read, divided by the count delta. Alternatively, dividing the values
outright would provide the lifetime average value. Typically these are used to measure latencies, the number of requests and a sum of request latencies. Some bit values
are combined, for example 5, 6 and 10. A bit value of 5 is a combination of bit 1 and bit 4. This means the average will be a floating point value. A bit value of 6 is a
combination of bit 2 and bit 4. This means the average value will be an integer. A bit value of 10 is a combination of bit 2 and bit 8. This means the counter value will be an
integer value.
Reference
Average count and sum
Dump Ceph performance counters
The ceph daemon .. perf dump command outputs the current values and groups the metrics under the collection name for each subsystem.
Prerequisites
A running IBM Storage Ceph cluster.
Root-level access to the node.
Procedure
1. To view the current metric data:
Syntax
ceph daemon DAEMON_NAME perf dump
Note: run the ceph daemon command from the node running the daemon.
2. Executing ceph daemon .. perf dump command from the Monitor node:
[ceph: root@host01 /]# ceph daemon mon.host01 perf dump
3. Executing the ceph daemon .. perf dump command from the OSD node:
[ceph: root@host01 /]# ceph daemon osd.11 perf dump
Reference
To view a short description of each Monitor metric available, see Ceph Monitor metrics.
Average count and sum
The avgcount is the number of operations within this range and the sum is the total latency in seconds. To know approximately how much latency there is per operation,
divide the sum by the avgcount.
All latency numbers have a bit field value of 5. This field contains floating point values for the average count and sum.
For a short description of each OSD metric available, see Ceph OSD metrics.
Ceph Monitor metrics
Cluster metrics
Level database metrics
General monitor metrics
Paxos metrics
Throttle metrics
Cluster metrics
288 IBM Storage Ceph
Table 1 describes the cluster metrics for the cluster collection.
Table 1. Cluster metrics table
Metric Name
num_mon
2
Bit Field Value
Number of monitors
Short Description
num_mon_quorum
2
Number of monitors in quorum
num_osd
2
Total number of OSD
num_osd_up
2
Number of OSDs that are up
num_osd_in
2
Number of OSDs that are in cluster
osd_epoch
2
Current epoch of OSD map
osd_bytes
2
Total capacity of cluster in bytes
osd_bytes_used
2
Number of used bytes on cluster
osd_bytes_avail
2
Number of available bytes on cluster
num_pool
2
Number of pools
num_pg
2
Total number of placement groups
num_pg_active_clean
2
Number of placement groups in active+clean state
num_pg_active
2
Number of placement groups in active state
num_pg_peering
2
Number of placement groups in peering state
num_object
2
Total number of objects on cluster
num_object_degraded
2
Number of degraded (missing replicas) objects
num_object_misplaced 2
Number of misplaced (wrong location in the cluster) objects
num_object_unfound
2
Number of unfound objects
num_bytes
2
Total number of bytes of all objects
num_mds_up
2
Number of MDSs that are up
num_mds_in
2
Number of MDS that are in cluster
num_mds_failed
2
Number of failed MDS
mds_epoch
2
Current epoch of MDS map
Level database metrics
Table 2 describes the level database metrics for the leveldb collection.
Table 2. Level database metrics table
Metric Name
leveldb_get
10
Bit Field Value
Gets
Short Description
leveldb_transaction
10
Transactions
leveldb_compact
10
Compactions
leveldb_compact_range
10
Compactions by range
leveldb_compact_queue_merge 10
Merging of ranges in compaction queue
leveldb_compact_queue_len
Length of compaction queue
2
General monitor metrics
Table 3 describes the general monitor metrics for the mon collection.
Table 3. General monitor metrics table
Metric Name
num_sessions
2
Bit Field Value
Current number of opened monitor sessions
Short Description
session_add
10
Number of created monitor sessions
session_rm
10
Number of remove_session calls in monitor
session_trim
10
Number of trimmed monitor sessions
num_elections 10
Number of elections monitor took part in
election_call 10
Number of elections started by monitor
election_win
10
Number of elections won by monitor
election_lose 10
Number of elections lost by monitor
Paxos metrics
Table 4 describes the Paxos metrics for the paxos collection.
Table 4. Paxos metrics table
Metric Name
start_leader
Bit Field Value
10
Starts in leader role
Short Description
start_peon
10
Starts in peon role
restart
10
Restarts
refresh
10
Refreshes
refresh_latency
5
Refresh latency
begin
10
Started and handled begins
begin_keys
6
Keys in transaction on begin
IBM Storage Ceph 289
Metric Name
begin_bytes
6
Bit Field Value
Data in transaction on begin
Short Description
begin_latency
5
Latency of begin operation
commit
10
Commits
commit_keys
6
Keys in transaction on commit
commit_bytes
6
Data in transaction on commit
commit_latency
5
Commit latency
collect
10
Peon collects
collect_keys
6
Keys in transaction on peon collect
collect_bytes
6
Data in transaction on peon collect
collect_latency
5
Peon collect latency
collect_uncommitted 10
Uncommitted values in started and handled collects
collect_timeout
10
Collect timeouts
accept_timeout
10
Accept timeouts
lease_ack_timeout
10
Lease acknowledgment timeouts
lease_timeout
10
Lease timeouts
store_state
10
Store a shared state on disk
store_state_keys
6
Keys in transaction in stored state
store_state_bytes
6
Data in transaction in stored state
store_state_latency 5
Storing state latency
share_state
10
Sharing of state
share_state_keys
6
Keys in shared state
share_state_bytes
6
Data in shared state
new_pn
10
New proposal number queries
new_pn_latency
5
New proposal number getting latency
Throttle metrics
Table 5 describes the throttle metrics for the throttle-* collection.
Table 5. Throttle Metrics Table
Metric Name
Bit Field Value
Short Description
val
10
Currently available throttle
max
10
Max value for throttle
get
10
Gets
get_sum
10
Got data
get_or_fail_fail
10
Get blocked during get_or_fail
get_or_fail_success 10
Successful get during get_or_fail
take
10
Takes
take_sum
10
Taken data
put
10
Puts
put_sum
10
Put data
wait
5
Waiting latency
Ceph OSD metrics
Write back throttle metrics
Level database metrics
Objecter metrics
Read and write operations metrics
Recovery state metrics
OSD throttle metrics
Write back throttle metrics
Table 1 describes the write back throttle metrics for the WBThrottle collection.
Table 1. Write back throttle metrics table
Metric Name
bytes_dirtied
2
Bit Field Value
Dirty data
bytes_wb
2
Written data
ios_dirtied
2
Dirty operations
ios_wb
2
Written operations
inodes_dirtied 2
290 IBM Storage Ceph
Short Description
Entries waiting for write
Metric Name
inodes_wb
Bit Field Value
2
Short Description
Written entries
Level database metrics
Table 2 describes the level database metrics for the leveldb collection.
Table 2. Level database metrics table
Collection Name
Metric Name
leveldb
leveldb_get
10
Bit Field Value
Gets
Short Description
leveldb_transaction
10
Transactions
leveldb_compact
10
Compactions
leveldb_compact_range
10
Compactions by range
leveldb_compact_queue_merge 10
Merging of ranges in compaction queue
leveldb_compact_queue_len
Length of compaction queue
2
Objecter metrics
Table 3 describes the objecter metrics for the objecter collection.
Table 3. Objecter metrics table
Metric Name
op_active
Bit Field Value
2
Active operations
Short Description
op_laggy
2
Laggy operations
op_send
10
Sent operations
op_send_bytes
10
Sent data
op_resend
10
Resent operations
op_ack
10
Commit callbacks
op_commit
10
Operation commits
op
10
Operation
op_r
10
Read operations
op_w
10
Write operations
op_rmw
10
Read-modify-write operations
op_pg
10
PG operation
osdop_stat
10
Stat operations
osdop_create
10
Create object operations
osdop_read
10
Read operations
osdop_write
10
Write operations
osdop_writefull
10
Write full object operations
osdop_append
10
Append operation
osdop_zero
10
Set object to zero operations
osdop_truncate
10
Truncate object operations
osdop_delete
10
Delete object operations
osdop_mapext
10
Map extent operations
osdop_sparse_read
10
Sparse read operations
osdop_clonerange
10
Clone range operations
osdop_getxattr
10
Get xattr operations
osdop_setxattr
10
Set xattr operations
osdop_cmpxattr
10
Xattr comparison operations
osdop_rmxattr
10
Remove xattr operations
osdop_resetxattrs
10
Reset xattr operations
osdop_tmap_up
10
TMAP update operations
osdop_tmap_put
10
TMAP put operations
osdop_tmap_get
10
TMAP get operations
osdop_call
10
Call (execute) operations
osdop_watch
10
Watch by object operations
osdop_notify
10
Notify about object operations
osdop_src_cmpxattr 10
Extended attribute comparison in multi operations
osdop_other
10
Other operations
linger_active
2
Active lingering operations
linger_send
10
Sent lingering operations
linger_resend
10
Resent lingering operations
linger_ping
10
Sent pings to lingering operations
poolop_active
2
Active pool operations
poolop_send
10
Sent pool operations
poolop_resend
10
Resent pool operations
poolstat_active
2
Active get pool stat operations
poolstat_send
10
Pool stat operations sent
IBM Storage Ceph 291
Metric Name
poolstat_resend
10
Bit Field Value
Resent pool stats
Short Description
statfs_active
2
Statfs operations
statfs_send
10
Sent FS stats
statfs_resend
10
Resent FS stats
command_active
2
Active commands
command_send
10
Sent commands
command_resend
10
Resent commands
map_epoch
2
OSD map epoch
map_full
10
Full OSD maps received
map_inc
10
Incremental OSD maps received
osd_sessions
2
Open sessions
osd_session_open
10
Sessions opened
osd_session_close
10
Sessions closed
osd_laggy
2
Laggy OSD sessions
Read and write operations metrics
Table 4 describes the read and write operations metrics for the osd collection.
Table 4. Read and write operations metrics table
Metric Name
Bit Field Value
Short Description
op_wip
2
Replication operations currently being processed (primary)
op_in_bytes
10
Client operations total write size
op_out_bytes
10
Client operations total read size
op_latency
5
Latency of client operations (including queue time)
op_process_latency
5
Latency of client operations (excluding queue time)
op_r
10
Client read operations
op_r_out_bytes
10
Client data read
op_r_latency
5
Latency of read operation (including queue time)
op_r_process_latency
5
Latency of read operation (excluding queue time)
op_w
10
Client write operations
op_w_in_bytes
10
Client data written
op_w_rlat
5
Client write operation readable/applied latency
op_w_latency
5
Latency of write operation (including queue time)
op_w_process_latency
5
Latency of write operation (excluding queue time)
op_rw
10
Client read-modify-write operations
op_rw_in_bytes
10
Client read-modify-write operations write in
op_rw_out_bytes
10
Client read-modify-write operations read out
op_rw_rlat
5
Client read-modify-write operation readable/applied latency
op_rw_latency
5
Latency of read-modify-write operation (including queue time)
op_rw_process_latency
5
Latency of read-modify-write operation (excluding queue time)
subop
10
Sub operations
subop_in_bytes
10
Sub operations total size
subop_latency
5
Sub operations latency
subop_w
10
Replicated writes
subop_w_in_bytes
10
Replicated written data size
subop_w_latency
5
Replicated writes latency
subop_pull
10
Sub operations pull requests
subop_pull_latency
5
Sub operations pull latency
subop_push
10
Sub operations push messages
subop_push_in_bytes
10
Sub operations pushed size
subop_push_latency
5
Sub operations push latency
pull
10
Pull requests sent
push
10
Push messages sent
push_out_bytes
10
Pushed size
push_in
10
Inbound push messages
push_in_bytes
10
Inbound pushed size
recovery_ops
10
Started recovery operations
loadavg
2
CPU load
buffer_bytes
2
Total allocated buffer size
numpg
2
Placement groups
numpg_primary
2
Placement groups for which this osd is primary
numpg_replica
2
Placement groups for which this osd is replica
numpg_stray
2
Placement groups ready to be deleted from this osd
heartbeat_to_peers
2
Heartbeat (ping) peers we send to
heartbeat_from_peers
2
Heartbeat (ping) peers we recv from
map_messages
10
OSD map messages
292 IBM Storage Ceph
Metric Name
map_message_epochs
10
Bit Field Value
OSD map epochs
Short Description
map_message_epoch_dups
10
OSD map duplicates
stat_bytes
2
OSD size
stat_bytes_used
2
Used space
stat_bytes_avail
2
Available space
copyfrom
10
RADOS copy-from operations
tier_promote
10
Tier promotions
tier_flush
10
Tier flushes
tier_flush_fail
10
Failed tier flushes
tier_try_flush
10
Tier flush attempts
tier_try_flush_fail
10
Failed tier flush attempts
tier_evict
10
Tier evictions
tier_whiteout
10
Tier whiteouts
tier_dirty
10
Dirty tier flag set
tier_clean
10
Dirty tier flag cleaned
tier_delay
10
Tier delays (agent waiting)
tier_proxy_read
10
Tier proxy reads
agent_wake
10
Tiering agent wake up
agent_skip
10
Objects skipped by agent
agent_flush
10
Tiering agent flushes
agent_evict
10
Tiering agent evictions
object_ctx_cache_hit
10
Object context cache hits
object_ctx_cache_total
10
Object context cache lookups
ceph_cluster_osd_blocklist_count 2
Number of clients blocklisted
Recovery state metrics
Table 5 describes the recovery state metrics for the recoverystate_perf collection.
Table 5. Recovery state metrics table
Metric Name
initial_latency
5
Bit Field Value
Initial recovery state latency
Short Description
started_latency
5
Started recovery state latency
reset_latency
5
Reset recovery state latency
start_latency
5
Start recovery state latency
primary_latency
5
Primary recovery state latency
peering_latency
5
Peering recovery state latency
backfilling_latency
5
Backfilling recovery state latency
waitremotebackfillreserved_latency 5
Wait remote backfill reserved recovery state latency
waitlocalbackfillreserved_latency
5
Wait local backfill reserved recovery state latency
notbackfilling_latency
5
Notbackfilling recovery state latency
repnotrecovering_latency
5
Repnotrecovering recovery state latency
repwaitrecoveryreserved_latency
5
Rep wait recovery reserved recovery state latency
repwaitbackfillreserved_latency
5
Rep wait backfill reserved recovery state latency
RepRecovering_latency
5
RepRecovering recovery state latency
activating_latency
5
Activating recovery state latency
waitlocalrecoveryreserved_latency
5
Wait local recovery reserved recovery state latency
waitremoterecoveryreserved_latency 5
Wait remote recovery reserved recovery state latency
recovering_latency
5
Recovering recovery state latency
recovered_latency
5
Recovered recovery state latency
clean_latency
5
Clean recovery state latency
active_latency
5
Active recovery state latency
replicaactive_latency
5
Replicaactive recovery state latency
stray_latency
5
Stray recovery state latency
getinfo_latency
5
Getinfo recovery state latency
getlog_latency
5
Getlog recovery state latency
waitactingchange_latency
5
Waitactingchange recovery state latency
incomplete_latency
5
Incomplete recovery state latency
getmissing_latency
5
Get missing recovery state latency
waitupthru_latency
5
Waitupthru recovery state latency
OSD throttle metrics
Table 6 describes the OSD throttle metrics for the throttle-* collection.
Table 6. OSD throttle metrics table
Metric Name
Bit Field Value
Short Description
IBM Storage Ceph 293
Metric Name
Bit Field Value
Short Description
val
10
Currently available throttle
max
10
Max value for throttle
get
10
Gets
get_sum
10
Got data
get_or_fail_fail
10
Get blocked during get_or_fail
get_or_fail_success 10
Successful get during get_or_fail
take
10
Takes
take_sum
10
Taken data
put
10
Puts
put_sum
10
Put data
wait
5
Waiting latency
Ceph Object Gateway metrics
Ceph Object Gateway client metrics
Objecter metrics
Ceph Object Gateway throttle metrics
Ceph Object Gateway client metrics
Table 1 describes the Ceph Object Gateway client metrics for the client.rgw.<rgw_node_name> collection.
Table 1. Ceph Object Gateway client metrics table
Metric Name
Bit Field Value
Short Description
req
10
Requests
failed_req
10
Aborted requests
get
10
Gets
get_b
10
Size of gets
get_initial_lat
5
Get latency
put
10
Puts
put_b
10
Size of puts
put_initial_lat
5
Put latency
qlen
2
Queue length
qactive
2
Active requests queue
cache_hit
10
Cache hits
cache_miss
10
Cache miss
keystone_token_cache_hit
10
Keystone token cache hits
keystone_token_cache_miss 10
Keystone token cache miss
Objecter metrics
Table 2 describes the objecter metrics for the objecter collection.
Table 2. Objecter metrics table
Metric Name
op_active
2
Active operations
op_laggy
2
Laggy operations
op_send
10
Sent operations
op_send_bytes
10
Sent data
op_resend
10
Resent operations
op_ack
10
Commit callbacks
op_commit
10
Operation commits
op
10
Operation
op_r
10
Read operations
op_w
10
Write operations
op_rmw
10
Read-modify-write operations
op_pg
10
PG operation
osdop_stat
10
Stat operations
osdop_create
10
Create object operations
osdop_read
10
Read operations
osdop_write
10
Write operations
osdop_writefull
10
Write full object operations
osdop_append
10
Append operation
osdop_zero
10
Set object to zero operations
osdop_truncate
10
Truncate object operations
294 IBM Storage Ceph
Bit Field Value
Short Description
Metric Name
osdop_delete
10
Bit Field Value
Delete object operations
Short Description
osdop_mapext
10
Map extent operations
osdop_sparse_read
10
Sparse read operations
osdop_clonerange
10
Clone range operations
osdop_getxattr
10
Get xattr operations
osdop_setxattr
10
Set xattr operations
osdop_cmpxattr
10
Xattr comparison operations
osdop_rmxattr
10
Remove xattr operations
osdop_resetxattrs
10
Reset xattr operations
osdop_tmap_up
10
TMAP update operations
osdop_tmap_put
10
TMAP put operations
osdop_tmap_get
10
TMAP get operations
osdop_call
10
Call (execute) operations
osdop_watch
10
Watch by object operations
osdop_notify
10
Notify about object operations
osdop_src_cmpxattr 10
Extended attribute comparison in multi operations
osdop_other
10
Other operations
linger_active
2
Active lingering operations
linger_send
10
Sent lingering operations
linger_resend
10
Resent lingering operations
linger_ping
10
Sent pings to lingering operations
poolop_active
2
Active pool operations
poolop_send
10
Sent pool operations
poolop_resend
10
Resent pool operations
poolstat_active
2
Active get pool stat operations
poolstat_send
10
Pool stat operations sent
poolstat_resend
10
Resent pool stats
statfs_active
2
Statfs operations
statfs_send
10
Sent FS stats
statfs_resend
10
Resent FS stats
command_active
2
Active commands
command_send
10
Sent commands
command_resend
10
Resent commands
map_epoch
2
OSD map epoch
map_full
10
Full OSD maps received
map_inc
10
Incremental OSD maps received
osd_sessions
2
Open sessions
osd_session_open
10
Sessions opened
osd_session_close
10
Sessions closed
osd_laggy
2
Laggy OSD sessions
Ceph Object Gateway throttle metrics
Table 3 describes the Ceph Object Gateway throttle metrics for the throttle-* collection.
Table 3. Ceph Object Gateway throttle metrics table
Metric Name
Bit Field Value
Short Description
val
10
Currently available throttle
max
10
Max value for throttle
get
10
Gets
get_sum
10
Got data
get_or_fail_fail
10
Get blocked during get_or_fail
get_or_fail_success 10
Successful get during get_or_fail
take
10
Takes
take_sum
10
Taken data
put
10
Puts
put_sum
10
Put data
wait
5
Waiting latency
mClock OSD scheduler
As a storage administrator, you can implement the IBM Storage Ceph's quality of service (QoS) using mClock queueing scheduler. This is based on an adaptation of the
mClock algorithm called dmClock.
The mClock OSD scheduler provides the desired QoS using configuration profiles to allocate proper reservation, weight, and limit tags to the service types.
IBM Storage Ceph 295
The mClock OSD scheduler performs the QoS calculations for the different device types, that is SSD or HDD, by using the OSD’s IOPS capability (determined automatically)
and maximum sequential bandwidth capability (See osd_mclock_max_sequential_bandwidth_hdd and osd_mclock_max_sequential_bandwidth_ssd in the
mclock configuration options.
Comparison of mClock OSD scheduler with WPQ OSD scheduler
Allocation of input and output resources
Factors impacting mClock operation queues
mClock configuration
mClock clients
mClock profiles
mClock profile types
Changing mClock profiles
Switching between built-in and custom profiles
Switching temporarily between mClock profiles
Degraded and misplaced object recovery rate with mClock profiles
Modifying backfills and recovery options
Ceph OSD capacity determination
Verifying the capacity of an OSD
Manually benchmarking OSDs
Determining BlueStore throttle values
Specifying maximum OSD capacity
mClock configuration options
Comparison of mClock OSD scheduler with WPQ OSD scheduler
The mClock OSD scheduler replaces the Weighted Priority Queue (WPQ) OSD scheduler as a default scheduler in IBM Storage Ceph 6.1.
Important: The mClock scheduler is supported for BlueStore OSDs. For Filestore OSDs the osd_op_queue is set to wpq and it is enforced even if the user attempts to
change it.
The WPQ OSD scheduler features a strict sub-queue, which is de-queued before the normal queue. The WPQ removes operations from a queue in relation to their
priorities to prevent depletion of any queue. This helps in cases where some Ceph OSDs are more overloaded than others.
The mClock OSD scheduler currently features an immediate queue, into which operations that require immediate response are queued. The immediate queue is not
handled by mClock and functions as a simple first in, first out queue and is given the first priority.
Operations, such as OSD replication operations, OSD operation replies, peering, recoveries marked with the highest priority, and so forth, are queued into the immediate
queue. All other operations are enqueued into the mClock queue that works according to the mClock algorithm.
The mClock queue, mclock_scheduler, prioritizes operations based on which bucket they belong to, that is pg recovery, pg scrub, snap
trim, client op, and pg deletion.
With background operations in progress, the average client throughputs, that is the input and output operations per second (IOPS), are significantly higher and latencies
are lower with the mClock profiles when compared to the WPQ scheduler. That is because of mClock’s effective allocation of the QoS parameters.
Reference
For automated OSD capacity determination, see mClock profiles.
Allocation of input and output resources
This section describes how the QoS controls work internally with reservation, limit, and weight allocation. The user is not expected to set these controls as the mClock
profiles automatically set them. Tuning these controls can only be performed using the available mClock profiles.
The dmClock algorithm allocates the input and output (I/O) resources of the Ceph cluster in proportion to weights. It implements the constraints of minimum reservation
and maximum limitation to ensure the services can compete for the resources fairly.
Currently, the mclock_scheduler operation queue divides Ceph services involving I/O resources into following buckets:
client op: the input and output operations per second (IOPS) issued by a client.
pg deletion: the IOPS issued by primary Ceph OSD.
snap trim: the snapshot trimming-related requests.
pg recovery: the recovery-related requests.
pg scrub: the scrub-related requests.
The resources are partitioned using the following three sets of tags, meaning that the share of each type of service is controlled by these three tags:
Reservation
Limit
Weight
Reservation
296 IBM Storage Ceph
The minimum IOPS allocated for the service. The more reservation a service has, the more resources it is guaranteed to possess, as long as it requires so.
For example, a service with the reservation set to 0.1 (or 10%) always has 10% of the OSD’s IOPS capacity allocated for itself. Therefore, even if the clients start to issue
large amounts of I/O requests, they do not exhaust all the I/O resources and the service’s operations are not depleted even in a cluster with high load.
Limit
The maximum IOPS allocated for the service. The service does not get more than the set number of requests per second serviced, even if it requires so and no other
services are competing with it. If a service crosses the enforced limit, the operation remains in the operation queue until the limit is restored.
Note: If the value is set to 0 (disabled), the service is not restricted by the limit setting and it can use all the resources if there is no other competing operation. This is
represented as "MAX" in the mClock profiles.
Note: The reservation and limit parameter allocations are per-shard, based on the type of backing device, that is HDD or SSD, under the Ceph OSD.
Weight
The proportional share of capacity if extra capacity or system is not enough. The service can use a larger portion of the I/O resource, if its weight is higher than its
competitor’s.
Note: The reservation and limit values for a service are specified in terms of a proportion of the total IOPS capacity of the OSD. The proportion is represented as a
percentage in the mClock profiles. The weight does not have a unit. The weights are relative to one another, so if one class of requests has a weight of 9 and another a
weight of 1, then the requests are performed at a 9 to 1 ratio. However, that only happens once the reservations are met and those values include the operations
performed under the reservation phase.
Important: If the weight is set to W, then for a given class of requests the next one that enters has a weight tag of 1/W and the previous weight tag, or the current time,
whichever is larger. That means, if W is too large and thus 1/W is too small, the calculated tag might never be assigned as it gets a value of the current time. Therefore,
values for weight should be always under the number of requests expected to be serviced each second.
Factors impacting mClock operation queues
There are three factors that can reduce the impact of the mClock operation queues within IBM Storage Ceph
The number of shards for client operations.
The number of operations in the operation sequencer.
The usage of distributed system for Ceph OSDs
The number of shards for client operations
Requests to a Ceph OSD are sharded by their placement group identifier. Each shard has its own mClock queue and these queues neither interact, nor share information
amongst them.
The number of shards can be controlled with these configuration options:
osd_op_num_shards
osd_op_num_shards_hdd
osd_op_num_shards_ssd
A lower number of shards increase the impact of the mClock queues, but might have other damaging effects.
Note: Use the default number of shards as defined by the configuration options osd_op_num_shards, osd_op_num_shards_hdd, and osd_op_num_shards_ssd.
The number of operations in the operation sequencer
Requests are transferred from the operation queue to the operation sequencer, in which they are processed. The mClock scheduler is located in the operation queue. It
determines which operation to transfer to the operation sequencer.
The number of operations allowed in the operation sequencer is a complex issue. The aim is to keep enough operations in the operation sequencer so it always works on
some, while it waits for disk and network access to complete other operations.
However, mClock no longer has control over an operation that is transferred to the operation sequencer. Therefore, to maximize the impact of mClock, the goal is also to
keep as few operations in the operation sequencer as possible.
The configuration options that influence the number of operations in the operation sequencer are:
bluestore_throttle_bytes
bluestore_throttle_deferred_bytes
bluestore_throttle_cost_per_io
bluestore_throttle_cost_per_io_hdd
bluestore_throttle_cost_per_io_ssd
Note: Use the default values as defined by the bluestore_throttle_bytes and bluestore_throttle_deferred_bytes options. However, these options can be
determined during the benchmarking phase.
The usage of distributed system for Ceph OSDs
The third factor that affects the impact of the mClock algorithm is the usage of a distributed system, where requests are made to multiple Ceph OSDs, and each Ceph OSD
can have multiple shards.
Note: dmClock is the distributed version of mClock.
IBM Storage Ceph 297
Reference
For more details about `osd_op_num_shards_hdd` and `osd_op_num_shards_ssd` parameters, see Object Storage Daemon (OSD) configuration options.
For more details about BlueStore throttle parameters, see BlueStore configuration options.
Manually benchmarking OSDs
mClock configuration
To make the mClock more user-friendly and intuitive, the mClock configuration profiles are introduced in IBM Storage Ceph 6. The mClock profiles hide the low-level
details from users, making it easier to configure and use mClock.
The following input parameters are required for an mClock profile to configure the quality of service (QoS) related parameters:
The total capacity of input and output operations per second (IOPS) for each Ceph OSD. This is determined automatically.
The maximum sequential bandwidth capacity (MiB/s) of each OS. See osd_mclock_max_sequential_bandwidth_[hdd/ssd] option
An mClock profile type to be enabled. The default is balanced.
Using the settings in the specified profile, a Ceph OSD determines and applies the lower-level mClock and Ceph parameters. The parameters applied by the mClock profile
make it possible to tune the QoS between the client I/O and background operations in the OSD.
Reference
For automated OSD capacity determination, see the Ceph OSD capacity determination .
mClock clients
The mClock scheduler handles requests from different types of Ceph services. Each service is considered by mClock as a type of client. Depending on the type of requests
handled, mClock clients are classified into the buckets:
Client - Handles input and output (I/O) requests issued by external clients of Ceph.
Background recovery - Handles internal recovery requests.
Background best-effort - Handles internal backfill, scrub, snap trim, and placement group (PG) deletion requests.
The mClock scheduler derives the cost of an operation used in the QoS calculations from osd_mclock_max_capacity_iops_hdd | osd_mclock_max_capacity_iops_ssd,
osd_mclock_max_sequential_bandwidth_hdd | osd_mclock_max_sequential_bandwidth_ssd and osd_op_num_shards_hdd | osd_op_num_shards_ssd parameters.
mClock profiles
An mClock profile is a configuration setting. When applied to a running IBM Storage Ceph cluster, it enables the throttling of the IOPS operations belonging to different
client classes, such as background recovery, scrub, snap trim, client op, and pg
deletion.
The mClock profile uses the capacity limits and the mClock profile type selected by the user to determine the low-level mClock resource control configuration parameters
and applies them transparently. Other IBM Storage Ceph configuration parameters are also applied. The low-level mClock resource control parameters are the
reservation, limit, and weight that provide control of the resource shares. The mClock profiles allocate these parameters differently for each client type.
mClock profile types
mClock profiles can be classified into built-in and custom profiles.
If any mClock profile is active, the following IBM Storage Ceph configuration sleep options get disabled, which means they are set to 0:
osd_recovery_sleep
osd_recovery_sleep_hdd
osd_recovery_sleep_ssd
osd_recovery_sleep_hybrid
osd_scrub_sleep
osd_delete_sleep
osd_delete_sleep_hdd
osd_delete_sleep_ssd
osd_delete_sleep_hybrid
298 IBM Storage Ceph
osd_snap_trim_sleep
osd_snap_trim_sleep_hdd
osd_snap_trim_sleep_ssd
osd_snap_trim_sleep_hybrid
It is to ensure that mClock scheduler is able to determine when to pick the next operation from its operation queue and transfer it to the operation sequencer. This results
in the desired QoS being provided across all its clients.
Custom profile
This profile gives users complete control over all the mClock configuration parameters.
Built-in profiles
When a built-in profile is enabled, the mClock scheduler calculates the low-level mClock parameters, that is, reservation, weight, and limit, based on the profile enabled
for each client type.
The mClock parameters are calculated based on the maximum Ceph OSD capacity provided beforehand. Therefore, the following mClock configuration options cannot be
modified when using any of the built-in profiles:
osd_mclock_scheduler_client_res
osd_mclock_scheduler_client_wgt
osd_mclock_scheduler_client_lim
osd_mclock_scheduler_background_recovery_res
osd_mclock_scheduler_background_recovery_wgt
osd_mclock_scheduler_background_recovery_lim
osd_mclock_scheduler_background_best_effort_res
osd_mclock_scheduler_background_best_effort_wgt
osd_mclock_scheduler_background_best_effort_lim
Note: These defaults cannot be changed using any of the config subsystem commands like config set, config daemon or config tell commands. Although
the above command(s) report success, the mclock QoS parameters are reverted to their respective built-in profile defaults.
The following recovery and backfill related Ceph options are overridden to mClock defaults:
Warning: Do not change these options as the built-in profiles are optimized based on them. Changing these defaults can result in unexpected performance outcomes.
osd_max_backfills
osd_recovery_max_active
osd_recovery_max_active_hdd
osd_recovery_max_active_ssd
The following options show the mClock defaults which is same as the current defaults to maximize the performance of the foreground client operations:
osd_max_backfills
Original default
1
mClock default
1
osd_recovery_max_active
Original default
0
mClock default
0
osd_recovery_max_active_hdd
Original default
3
mClock default
3
osd_recovery_max_active_sdd
Original default
10
mClock default
10
Note: The above mClock defaults can be modified, only if necessary, by enabling osd_mclock_override_recovery_settings that is by default set as false. See
Modifying backfill and recovery options to modify these parameters.
IBM Storage Ceph 299
Built-in profiles
Users can choose from the following built-in profile types:
balanced (default)
high_client_ops
high_recovery_ops
Note: The values mentioned in the list below represent the proportion of the total IOPS capacity of the Ceph OSD allocated for the service type.
balanced:
balanced:
The default mClock profile is set to balanced because it represents a compromise between prioritizing client IO or recovery IO. It allocates equal reservation or priority
to client operations and background recovery operations. Background best-effort operations are given lower reservation and therefore take longer to complete when there
are competing operations. This profile meets the normal or steady state requirements of the cluster which is the case when external client performance requirements is
not critical and there are other background operations that still need attention within the OSD.
There might be instances that necessitate giving higher priority to either client operations or recovery operations. To meet such requirements you can choose either the
high_client_ops profile to prioritize client IO or the high_recovery_ops profile to prioritize recovery IO. These profiles are discussed further below.
Service type: client
Reservation
50%
Limit
MAX
Weight
1
Service type: background recovery
Reservation
50%
Limit
MAX
Weight
1
Service type: background best-effort
Reservation
MIN
Limit
90%
Weight
1
high_client_ops
high_client_ops:
This profile optimizes client performance over background activities by allocating more reservation and limit to client operations as compared to background operations in
the Ceph OSD. This profile, for example, can be enabled to provide the needed performance for I/O intensive applications for a sustained period of time at the cost of
slower recoveries. The list below shows the resource control parameters set by the profile:
Service type: client
Reservation
60%
Limit
MAX
Weight
2
Service type: background recovery
Reservation
40%
Limit
MAX
Weight
1
Service type: background best-effort
300 IBM Storage Ceph
Reservation
MIN
Limit
70%
Weight
1
high_recovery_ops
high_recovery_ops:
This profile optimizes background recovery performance as compared to external clients and other background operations within the Ceph OSD.
For example, it could be temporarily enabled by an administrator to accelerate background recoveries during non-peak hours. The list below shows the resource control
parameters set by the profile:
Service type: client
Reservation
30%
Limit
MAX
Weight
1
Service type: background recovery
Reservation
70%
Limit
MAX
Weight
2
Service type: background best-effort
Reservation
MIN
Limit
MAX
Weight
1
Reference
For more information about mClock configuration options, see the mClock configuration options.
Changing mClock profiles
The default mClock profile is set to balanced. The other types of the built-in profile are high_client_ops and high_recovery_ops.
Note: The custom profile is not recommended unless you are an advanced user.
Prerequisites
A running IBM Storage Ceph cluster.
Root-level access to the Ceph Monitor host.
Procedure
1. Log into the Cephadm shell:
Example
[root@host01 ~]# cephadm shell
2. Set the osd_mclock_profile option:
Syntax
ceph config set osd._OSD_ID_ osd_mclock_profile VALUE
Example
IBM Storage Ceph 301
[ceph: root@host01 /]# ceph config set osd.0 osd_mclock_profile high_recovery_ops
This example changes the profile to allow faster recoveries on osd.0.
Note: For optimal performance the profile must be set on all Ceph OSDs by using the following command:
Syntax
ceph config set osd osd_mclock_profile VALUE
Switching between built-in and custom profiles
The following steps describe switching from built-in profile to custom profile and vice-versa.
You might want to switch to the custom profile if you want complete control over all the mClock configuration options. However, it is recommended not to use the custom
profile unless you are an advanced user.
Prerequisites
A running IBM Storage Ceph cluster.
Root-level access to the Ceph Monitor host.
Procedure
Switch from built-in profile to custom profile
1. Log into the Cephadm shell:
Example
[root@host01 ~]# cephadm shell
2. Switch to the custom profile:
Syntax
ceph config set osd.OSD_ID osd_mclock_profile custom
Example
[ceph: root@host01 /]# ceph config set osd.0 osd_mclock_profile custom
Note: For optimal performance the profile must be set on all Ceph OSDs by using the following command:
[ceph: root@host01 /]# ceph config set osd osd_mclock_profile custom
3. Optional: After switching to the custom profile, modify the desired mClock configuration options:
Syntax
ceph config set osd.OSD_ID MCLOCK_CONFIGURATION_OPTION VALUE
Example
[ceph: root@host01 /]# ceph config set osd.0 osd_mclock_scheduler_client_res 0.5
This example changes the client reservation IOPS ratio for a specific OSD osd.0 to 0.5 (50%)
Important: Change the reservations of other services, such as background recovery and background best-effort accordingly to ensure that the sum of the
reservations does not exceed the maximum proportion (1.0) of the IOPS capacity of the OSD.
Switch from custom profile to built-in profile
1. Log into the cephadm shell:
Example
[root@host01 ~]# cephadm shell
2. Set the desired built-in profile:
Syntax
ceph config set osd osd_mclock_profile MCLOCK_PROFILE
Example
[ceph: root@host01 /]# ceph config set osd osd_mclock_profile high_client_ops
This example sets the built-in profile to high_client_ops on all Ceph OSDs.
3. Determine the existing custom mClock configuration settings in the database:
Example
[ceph: root@host01 /]# ceph config dump
302 IBM Storage Ceph
4. Remove the custom mClock configuration settings determined earlier:
Syntax
ceph config rm osd MCLOCK_CONFIGURATION_OPTION
Example
[ceph: root@host01 /]# ceph config rm osd osd_mclock_scheduler_client_res
This example removes the configuration option osd_mclock_scheduler_client_res that was set on all Ceph OSDs.
After all existing custom mClock configuration settings are removed from the central configuration database, the configuration settings related to
high_client_ops are applied.
5. Verify the settings on Ceph OSDs:
Syntax
ceph config show osd.OSD_ID
Example
[ceph: root@host01 /]# ceph config show osd.0
Reference
For the list of the mClock configuration options that cannot be modified with _built-in_ profiles, see mClock profile types.
Switching temporarily between mClock profiles
This section contains steps to temporarily switch between mClock profiles.
Warning: This section is for advanced users or for experimental testing. Do not use the below commands on a running storage cluster as it could have unexpected
outcomes.
Note: The configuration changes on a Ceph OSD using the below commands are temporary and are lost when the Ceph OSD is restarted.
Important: The configuration options that are overridden using the commands described in this section cannot be modified further using the ceph config set
osd.OSD_ID command. The changes do not take effect until a given Ceph OSD is restarted. This is intentional, as per the configuration subsystem design. However, any
further modifications can still be made temporarily using these commands.
Prerequisites
A running IBM Storage Ceph cluster.
Root-level access to the Ceph Monitor host.
Procedure
1. Log into the Cephadm shell:
Example
[root@host01 ~]# cephadm shell
2. Run the following command to override the mClock settings:
Syntax
ceph tell osd.OSD_ID injectargs '--MCLOCK_CONFIGURATION_OPTION=VALUE'
Example
[ceph: root@host01 /]# ceph tell osd.0 injectargs '--osd_mclock_profile=high_recovery_ops'
This example overrides the osd_mclock_profile option on osd.0.
3. Optional: You can use the alternative to the previous ceph tell osd.OSD_ID
injectargs command:
Syntax
ceph daemon osd.OSD_ID config set MCLOCK_CONFIGURATION_OPTION VALUE
Example
[ceph: root@host01 /]# ceph daemon osd.0 config set osd_mclock_profile high_recovery_ops
Note: The individual QoS related configuration options for the custom profile can also be modified temporarily using the above commands.
Degraded and misplaced object recovery rate with mClock profiles
IBM Storage Ceph 303
Degraded object recovery is categorized into the background recovery bucket. Across all mClock profiles, degraded object recovery is given higher priority when compared
to misplaced object recovery because degraded objects present a data safety issue not present with objects that are merely misplaced.
Backfill or the misplaced object recovery operation is categorized into the background best-effort bucket. According to the balanced and high_client_ops mClock
profiles, background best-effort client is not constrained by reservation (set to zero) but is limited to use a fraction of the participating OSD’s capacity if there are no other
competing services.
Therefore, with the balanced or high_client_ops profile and with other background competing services active, backfilling rates are expected to be slower when
compared to the previous WeightedPriorityQueue (WPQ) scheduler.
If higher backfill rates are desired, please follow the steps mentioned in the section below.
Improving backfilling rates
For faster backfilling rate when using either balanced or high_client_ops profile, follow the below steps:
Switch to the 'high_recovery_ops' mClock profile for the duration of the backfills. See changing an clock profile to achieve this. Once the backfilling phase is
complete, switch the mClock profile to the previously active profile. In case there is no significant improvement in the backfilling rate with the
high_recovery_ops profile, continue to the next step.
Switch the mClock profile back to the previously active profile.
Modify osd_max_backfills to a higher value, for example, 3. See Modifying backfills and recovery options.
Once the backfilling is complete, osd_max_backfills can be reset to the default value of 1 by following the same procedure mentioned in step 3.
Warning: Please note that modifying osd_max_backfills may result in other operations, for example, client operations may experience higher latency during the
backfilling phase. Therefore, users are recommended to increase osd_max_backfills in small increments to minimize performance impact to other operations in the
cluster.
Modifying backfills and recovery options
Modify the backfills and recovery options with the ceph
config set command.
The backfill or recovery options that can be modified are listed in mClock profile types.
Warning: This section is for advanced users or for experimental testing. Do not use the below commands on a running storage cluster as it could have unexpected
outcomes.
Modify the values only for experimental testing, or if the cluster is unable to handle the values or it shows poor performance with the default settings.
Important: The modification of the mClock default backfill or recovery options is restricted by the osd_mclock_override_recovery_settings option, which is set to
false by default.
If you attempt to modify any default backfill or recovery options without setting osd_mclock_override_recovery_settings to true, it resets the options back to
the mClock defaults along with a warning message logged in the cluster log.
Prerequisites
A running IBM Storage Ceph cluster.
Root-level access to the Ceph Monitor host.
Procedure
1. Log into the Cephadm shell:
Example
[root@host01 ~]# cephadm shell
2. Set the osd_mclock_override_recovery_settings configuration option to true on all Ceph OSDs:
Example
[ceph: root@host01 /]# ceph config set osd osd_mclock_override_recovery_settings true
3. Set the desired backfills or recovery option:
Syntax
ceph config set osd OPTION VALUE
Example
[ceph: root@host01 /]# ceph config set osd osd_max_backfills 5
4. Wait a few seconds and verify the configuration for the specific OSD:
Syntax
ceph config show osd.OSD_ID | grep OPTION
Example
304 IBM Storage Ceph
[ceph: root@host01 /]# ceph config show osd.0 | grep osd_max_backfills
5. Reset the osd_mclock_override_recovery_settings configuration option to false on all OSDs:
Example
[ceph: root@host01 /]# ceph config set osd osd_mclock_override_recovery_settings false
Ceph OSD capacity determination
The Ceph OSD capacity in terms of total IOPS is determined automatically during the Ceph OSD initialization. This is achieved by running the Ceph OSD bench tool and
overriding the default value of osd_mclock_max_capacity_iops_[hdd, ssd] option depending on the device type. No other action or input is expected from the
user to set the Ceph OSD capacity.
Mitigation of unrealistic Ceph OSD capacity from the automated procedure
In certain conditions, the Ceph OSD bench tool might show unrealistic or inflated results depending on the drive configuration and other environment related conditions.
To mitigate the performance impact due to this unrealistic capacity, a couple of threshold configuration options depending on the OSD device type are defined and used:
osd_mclock_iops_capacity_threshold_hdd = 500
osd_mclock_iops_capacity_threshold_ssd = 80000
The following automated step is performed:
Fallback to using default OSD capacity
If the Ceph OSD bench tool reports a measurement that exceeds the above threshold values, the fallback mechanism reverts to the default value of
osd_mclock_max_capacity_iops_hdd or osd_mclock_max_capacity_iops_ssd. The threshold configuration options can be reconfigured based on the type of
drive used.
A cluster warning is logged in case the measurement exceeds the threshold:
Example
2022-10-27T15:30:23.270+0000 7f9b5dbe95c0 0 log_channel(cluster) log [WRN]
: OSD bench result of 39546.479392 IOPS exceeded the threshold limit of 25000.000000 IOPS for osd.1. IOPS capacity is unchanged
at 21500.000000 IOPS. The recommendation is to establish the osd's IOPS capacity using other benchmark tools (e.g. Fio) and
then override osd_mclock_max_capacity_iops_[hdd|ssd].
Important: If the default capacity does not accurately represent the Ceph OSD capacity, it is highly recommended to run a custom benchmark using the preferred tool, for
example Fio, on the drive and then override the osd_mclock_max_capacity_iops_[hdd, ssd] option as described in specifying maximum OSD capacity
Reference
To manually benchmark Ceph OSDs or manually tune the BlueStore throttle parameters, see Manually benchmarking OSDs.
For more information about the osd_mclock_max_capacity_iops_[hdd, ssd] and osd_mclock_iops_capacity_threshold_[hdd, ssd] options, see
the mClock configuration options.
Verifying the capacity of an OSD
You can verify the capacity of a Ceph OSD after setting up the storage cluster.
Prerequisites
A running IBM Storage Ceph cluster.
Root-level access to the Ceph Monitor host.
Procedure
1. Log into the Cephadm shell:
Example
[root@host01 ~]# cephadm shell
2. Verify the capacity of a Ceph OSD:
Syntax
ceph config show osd.OSD_ID osd_mclock_max_capacity_iops_[hdd, ssd]
Example
[ceph: root@host01 /]# ceph config show osd.0 osd_mclock_max_capacity_iops_ssd
21500.000000
IBM Storage Ceph 305
Manually benchmarking OSDs
To manually benchmark a Ceph OSD, any existing benchmarking tool, for example Fio, can be used. Regardless of the tool or command used, the steps below remain the
same.
Important: The number of shards and BlueStore throttle parameters have an impact on the mClock operation queues. Therefore, it is critical to set these values carefully
in order to maximize the impact of the mclock scheduler. For more information on these values, see factors that impact mClock operation queues.
Note: The steps in this section are only necessary if you want to override the Ceph OSD capacity determined automatically during the OSD initialization.
If you have already determined the benchmark data and wish to manually override the maximum OSD capacity for a Ceph OSD, see Specifying maximum OSD capacity.
Prerequisites
A running IBM Storage Ceph cluster.
Root-level access to the Ceph Monitor host.
Procedure
1. Log into the Cephadm shell:
Example
[root@host01 ~]# cephadm shell
2. Benchmark a Ceph OSD:
Syntax
ceph tell osd.OSD_ID bench [TOTAL_BYTES] [BYTES_PER_WRITE] [OBJ_SIZE] [NUM_OBJS]
where:
TOTAL_BYTES: Total number of bytes to write.
BYTES_PER_WRITE: Block size per write.
OBJ_SIZE: Bytes per object.
NUM_OBJS: Number of objects to write.
Example
[ceph: root@host01 /]# ceph tell osd.0 bench 12288000 4096 4194304 100
{
"bytes_written": 12288000,
"blocksize": 4096,
"elapsed_sec": 1.3718913019999999,
"bytes_per_sec": 8956977.8466311768,
"iops": 2186.7621695876896
}
Determining BlueStore throttle values
This optional section details the steps used to determine the correct BlueStore throttle values. The steps use the default shards.
Important: Before running the test, clear the caches to get an accurate measurement. Clear the OSD caches between each benchmark run using the following command:
Syntax
ceph tell osd.OSD_ID cache drop
Example
[ceph: root@host01 /]# ceph tell osd.0 cache drop
Prerequisites
A running IBM Storage Ceph cluster.
Root-level access to the Ceph Monitor node hosting the OSDs that you wish to benchmark.
Procedure
1. Log into the Cephadm shell:
Example
[root@host01 ~]# cephadm shell
2. Run a simple 4KiB random write workload on an OSD:
306 IBM Storage Ceph
Syntax
ceph tell osd.OSD_ID bench 12288000 4096 4194304 100
Example
[ceph: root@host01 /]# ceph tell osd.0 bench 12288000 4096 4194304 100
{
"bytes_written": 12288000,
"blocksize": 4096,
"elapsed_sec": 1.3718913019999999,
"bytes_per_sec": 8956977.8466311768,
"iops": 2186.7621695876896 1
}
a. The overall throughput obtained from the output of the osd bench command. This value is the baseline throughput, when the default BlueStore throttle
options are in effect.
3. Note the overall throughput, that is IOPS, obtained from the output of the previous command.
4. If the intent is to determine the BlueStore throttle values for your environment, set bluestore_throttle_bytes and
bluestore_throttle_deferred_bytes options to 32 KiB, that is, 32768 Bytes:
Syntax
ceph config set osd.OSD_ID bluestore_throttle_bytes 32768
ceph config set osd.OSD_ID bluestore_throttle_deferred_bytes 32768
Example
[ceph: root@host01 /]# ceph config set osd.0 bluestore_throttle_bytes 32768
[ceph: root@host01 /]# ceph config set osd.0 bluestore_throttle_deferred_bytes 32768
Otherwise, you can skip to the next section, specifying maximum OSD capacity.
5. Run the 4KiB random write test as before using an OSD bench command:
Example
[ceph: root@host01 /]# ceph tell osd.0 bench 12288000 4096 4194304 100
6. Notice the overall throughput from the output and compare the value against the baseline throughput recorded earlier.
7. If the throughput does not match with the baseline, increase the BlueStore throttle options by multiplying by 2.
8. Repeat the steps by running the 4KiB random write test, comparing the value against the baseline throughput, and increasing the BlueStore throttle options by
multiplying by 2, until the obtained throughput is very close to the baseline value.
Note: For example, during benchmarking on a machine with NVMe SSDs, a value of 256 KiB for both BlueStore throttle and deferred bytes was determined to maximize
the impact of mClock. For HDDs, the corresponding value was 40 MiB, where the overall throughput was roughly equal to the baseline throughput.
In general for HDDs, the BlueStore throttle values are expected to be higher when compared to SSDs.
Specifying maximum OSD capacity
You can override the maximum Ceph OSD capacity automatically set during OSD initialization.
These steps are optional. Perform the following steps if the default capacity does not accurately represent the Ceph OSD capacity.
Note: Ensure that you determine the benchmark data first, as described in manually benchmarking OSDs.
Prerequisites
A running IBM Storage Ceph cluster.
Root-level access to the Ceph Monitor host.
Procedure
Procedure
1. Log into the Cephadm shell:
Example
[root@host01 ~]# cephadm shell
2. Set osd_mclock_max_capacity_iops_[hdd, ssd] option for an OSD:
Syntax
ceph config set osd.OSD_ID osd_mclock_max_capacity_iops_[hdd,ssd] VALUE
Example
[ceph: root@host01 /]# ceph config set osd.0 osd_mclock_max_capacity_iops_hdd 350
This example sets the maximum capacity for osd.0, where an underlying device type is HDD, to 350 IOPS.
IBM Storage Ceph 307
mClock configuration options
This section contains the list of mClock configuration options:
osd_mclock_profile
Description
It sets the type of mClock profile to use for providing the quality of service (QoS) based on operations belonging to different classes, such as background
recovery, backfill, pg scrub, snap trim, client op, and pg
deletion.
Once a built-in profile is enabled, the lower-level mClock resource control parameters, that is reservation, weight, and limit, and some Ceph configuration
parameters are set transparently. This does not apply for the custom profile. Type;; String Default;; balanced Valid choices;; balanced,
high_recovery_ops, high_client_ops, custom
osd_mclock_max_capacity_iops_hdd
Description
It sets a maximum random write IOPS capacity, at 4 KiB block size, to consider per OSD for rotational media. Contributes in QoS calculations when enabling
a dmclock profile. It is only considered for osd_op_queue = mclock_scheduler
Type
Float
Default
315.0
osd_mclock_max_capacity_iops_ssd
Description
It sets a maximum random write IOPS capacity, at 4 KiB block size, to consider per OSD for solid state media.
Type
Float
Default
21500.0
osd_mclock_cost_per_byte_usec_ssd
Description
Indicates cost per byte in microseconds to consider per OSD for SDDs.Contributes in QoS calculations when enabling a dmclock profile. It is only considered
for osd_op_queue =
mclock_scheduler
Type
Float
Default
0.011
osd_mclock_max_sequential_bandwidth_hdd
Description
Indicates the maximum sequential bandwidth in bytes to consider for an OSD whose underlying device type is rotational media. This is considered by the
mclock scheduler to derive the cost factor to be used in QoS calculations. Only considered for osd_op_queue =
mclock_scheduler
Type
Size
Default
150_M
osd_mclock_max_sequential_bandwidth_ssd
Description
Indicates the maximum sequential bandwidth in bytes to consider for an OSD whose underlying device type is solid state media. This is considered by the
mclock scheduler to derive the cost factor to be used in QoS calculations. Only considered for osd_op_queue =
mclock_scheduler
Type
Size
Default
1200_M
osd_mclock_force_run_benchmark_on_init
Description
This force-runs the OSD benchmark on OSD initialization or boot-up.
Type
Boolean
Default
False
308 IBM Storage Ceph
See also
osd_mclock_max_capacity_iops_hdd, osd_mclock_max_capacity_iops_ssd
osd_mclock_skip_benchmark
Description
Setting this option skips the OSD benchmosd_mclock_max_capacity_iops_hddark on OSD initialization or boot-up.
Type
Boolean
Default
False
See also
osd_mclock_max_capacity_iops_hdd, osd_mclock_max_capacity_iops_ssd
osd_mclock_override_recovery_settings
Description
Setting this option enables the override of the recovery or backfill limits for the mClock scheduler as defined by the osd_recovery_max_active_hdd,
osd_recovery_max_active_ssd, and osd_max_backfills options.
Type
Boolean
Default
False
See also
osd_recovery_max_active_hdd, osd_recovery_max_active_ssd, osd_max_backfills
osd_mclock_iops_capacity_threshold_hdd
Description
It indicates the threshold IOPS capacity, at 4KiB block size, beyond which to ignore the Ceph OSD bench results for an OSD for HDDs.
Type
Float
Default
500.0
osd_mclock_iops_capacity_threshold_ssd
Description
It indicates the threshold IOPS capacity, at 4KiB block size, beyond which to ignore the Ceph OSD bench results for an OSD for SSDs.
Type
Float
Default
80000.0
osd_mclock_scheduler_client_res
Description
It is the default I/O proportion reserved for each client. The default value of 0 specifies the lowest possible reservation. Any value greater than 0 and up to
1.0 specifies the minimum IO proportion to reserve for each client in terms of a fraction of the OSD’s maximum IOPS capacity.
Type
float
Default
0
min
0
max
1.0
osd_mclock_scheduler_client_wgt
Description
It is the default I/O share for each client over reservation.
Type
Unsigned integer
Default
1
osd_mclock_scheduler_client_lim
Description
It is the default I/O limit for each client over reservation. The default value of 0 specifies no limit enforcement, which means each client can use the
maximum possible IOPS capacity of the OSD. Any value greater than 0 and up to 1.0 specifies the upper IO limit over reservation that each client receives in
terms of a fraction of the OSD’s maximum IOPS capacity.
IBM Storage Ceph 309
Type
float
Default
0
min
0
max
1.0
osd_mclock_scheduler_background_recovery_res
Description
It is the default I/O proportion reserved for background recovery. The default value of 0 specifies the lowest possible reservation. Any value greater than 0
and up to 1.0 specifies the minimum IO proportion to reserve for background recovery operations in terms of a fraction of the OSD’s maximum IOPS capacity.
Type
float
Default
0
min
0
max
1.0
osd_mclock_scheduler_background_recovery_wgt
Description
It indicates the I/O share for each background recovery over reservation.
Type
Unsigned integer
Default
1
osd_mclock_scheduler_background_recovery_lim
Description
It indicates the I/O limit for background recovery over reservation. The default value of 0 specifies no limit enforcement, which means background recovery
operation can use the maximum possible IOPS capacity of the OSD. Any value greater than 0 and up to 1.0 specifies the upper IO limit over reservation that
background recovery operation receives in terms of a fraction of the OSD’s maximum IOPS capacity.
Type
float
Default
0
min
0
max
1.0
osd_mclock_scheduler_background_best_effort_res
Description
It indicates the default I/O proportion reserved for background best_effort. The default value of 0 specifies the lowest possible reservation. Any value
greater than 0 and up to 1.0 specifies the minimum IO proportion to reserve for background best_effort operations in terms of a fraction of the OSD’s
maximum IOPS capacity.
Type
float
Default
0
min
0
max
1.0
osd_mclock_scheduler_background_best_effort_wgt
Description
It indicates the I/O share for each background best_effort over reservation.
Type
Unsigned integer
Default
1
310 IBM Storage Ceph
osd_mclock_scheduler_background_best_effort_lim
Description
It indicates the I/O limit for background best_effort over reservation. The default value of 0 specifies no limit enforcement, which means background
best_effort operation can use the maximum possible IOPS capacity of the OSD. Any value greater than 0 and up to 1.0 specifies the upper IO limit over
reservation that background best_effort operation receives in terms of a fraction of the OSD’s maximum IOPS capacity.
Type
float
Default
0
min
0
max
1.0
Reference
For more details about osd_op_queue option, see Object Storage Daemon (OSD) configuration options .
BlueStore
BlueStore is the back-end object store for the OSD daemons and puts objects directly on the block device.
Important: BlueStore provides a high-performance backend for OSD daemons in a production environment. By default, BlueStore is configured to be self-tuning. If you
determine that your environment performs better with BlueStore tuned manually, contact IBM Support and share the details of your configuration to help IBM improve the
auto-tuning capability. IBM looks forward to your feedback and appreciates your recommendations.
BlueStore features
BlueStore devices
BlueStore caching
The BlueStore cache is a collection of buffers that, depending on configuration, can be populated with data as the OSD daemon does reading from or writing to the
disk.
Sizing considerations
When mixing traditional and solid state drives using BlueStore OSDs, it is important to size the RocksDB logical volume (block.db) appropriately.
Tuning BlueStore
Use this procedure for new or freshly deployed OSDs.
Resharding RocksDB database
BlueStore fragmentation tool
Ceph BlueStore BlueFS
BlueStore block database stores metadata as key-value pairs in a RocksDB database. The block database resides on a small BlueFS partition on the storage device.
BlueFS is a minimal file system that is designed to hold the RocksDB files.
BlueStore features
The following are some of the main features of using BlueStore:
Direct management of storage devices
BlueStore consumes raw block devices or partitions. This avoids any intervening layers of abstraction, such as local file systems like XFS, that might limit
performance or add complexity.
Metadata management with RocksDB
BlueStore uses the RocksDB key-value database to manage internal metadata, such as the mapping from object names to block locations on a disk.
Full data and metadata checksumming
By default all data and metadata written to BlueStore is protected by one or more checksums. No data or metadata are read from disk or returned to the user
without verification.
Efficient copy-on-write
The Ceph Block Device and Ceph File System snapshots rely on a copy-on-write clone mechanism that is implemented efficiently in BlueStore. This results in
efficient I/O both for regular snapshots and for erasure coded pools which rely on cloning to implement efficient two-phase commits.
No large double-writes
BlueStore first writes any new data to unallocated space on a block device, and then commits a RocksDB transaction that updates the object metadata to reference
the new region of the disk. Only when the write operation is below a configurable size threshold, it falls back to a write-ahead journaling scheme.
Multi-device support
BlueStore can use multiple block devices for storing different data. For example: Hard Disk Drive (HDD) for the data, Solid-state Drive (SSD) for metadata, Nonvolatile Memory (NVM) or Non-volatile random-access memory (NVRAM) or persistent memory for the RocksDB write-ahead log (WAL). For more information, see
BlueStore devices.
Efficient block device usage
Because BlueStore does not use any file system, it minimizes the need to clear the storage device cache.
BlueStore devices
IBM Storage Ceph 311
BlueStore manages either one, two, or three storage devices in the backend.
Primary
WAL
DB
In the simplest case, BlueStore consumes a single primary storage device. The storage device is partitioned into two parts that contain:
OSD metadata: A small partition formatted with XFS that contains basic metadata for the OSD. This data directory includes information about the OSD, such as its
identifier, which cluster it belongs to, and its private keyring.
Data: A large partition occupying the rest of the device that is managed directly by BlueStore and that contains all of the OSD data. This primary device is identified
by a block symbolic link in the data directory.
You can also use two additional devices:
A WAL (write-ahead-log) device: A device that stores BlueStore internal journal or write-ahead log. It is identified by the block.wal symbolic link in the data
directory. Consider using a WAL device only if the device is faster than the primary device. For example, when the WAL device uses an SSD disk and the primary
device uses an HDD disk.
A DB device: A device that stores BlueStore internal metadata. The embedded RocksDB database puts as much metadata as it can on the DB device instead of on
the primary device to improve performance. If the DB device is full, it starts adding metadata to the primary device. Consider using a DB device only if the device is
faster than the primary device.
Warning: If you have only less than a gigabyte storage available on fast devices, IBM recommends using it as a WAL device. If you have more fast devices available,
consider using it as a DB device. The BlueStore journal is always placed on the fastest device, so using a DB device provides the same benefit that the WAL device provides
while also allowing for storing additional metadata.
BlueStore caching
The BlueStore cache is a collection of buffers that, depending on configuration, can be populated with data as the OSD daemon does reading from or writing to the disk.
By default in IBM Storage Ceph, BlueStore will cache on reads, but not writes. This is because the bluestore_default_buffered_write option is set to false to
avoid potential overhead associated with cache eviction.
If the bluestore_default_buffered_write option is set to true, data is written to the buffer first, and then committed to disk. Afterward, a write acknowledgment
is sent to the client, allowing subsequent reads faster access to the data already in cache, until that data is evicted.
Read-heavy workloads will not see an immediate benefit from BlueStore caching. As more reading is done, the cache will grow over time and subsequent reads will see an
improvement in performance. How fast the cache populates depends on the BlueStore block and database disk type, and the client’s workload requirements.
Important: Before enabling the bluestore_default_buffered_write option, contact IBM Support.
Sizing considerations
When mixing traditional and solid state drives using BlueStore OSDs, it is important to size the RocksDB logical volume (block.db) appropriately.
IBM recommends that the RocksDB logical volume be no less than 4% of the block size with object, file and mixed workloads. IBM supports 1% of the BlueStore block
size with RocksDB and OpenStack block workloads. For example, if the block size is 1 TB for an object workload, then at a minimum, create a 40 GB RocksDB logical
volume.
When not mixing drive types, there is no requirement to have a separate RocksDB logical volume. BlueStore will automatically manage the sizing of RocksDB.
BlueStore’s cache memory is used for the key-value pair metadata for RocksDB, BlueStore metadata, and object data.
Note: The BlueStore cache memory values are in addition to the memory footprint already being consumed by the OSD.
Tuning BlueStore
Use this procedure for new or freshly deployed OSDs.
In BlueStore, the raw partition is allocated and managed in chunks of bluestore_min_alloc_size. By default, bluestore_min_alloc_size is 4096, equivalent to
4 KiB for HDDs and SSDs. The unwritten area in each chunk is filled with zeroes when it is written to the raw partition. This can lead to wasted unused space when not
properly sized for your workload, for example when writing small objects.
It is best practice to set bluestore_min_alloc_size to match the smallest write so this write amplification penalty can be avoided.
Important: Changing the value of bluestore_min_alloc_size is not recommended. For any assistance, contact IBM Support.
Note: The settings bluestore_min_alloc_size_ssd and bluestore_min_alloc_size_hdd are specific to SSDs and HDDs, respectively. However, setting them is
not necessary because setting bluestore_min_alloc_size overrides them.
Prerequisites
A running IBM Storage Ceph cluster.
Ceph monitors and managers are deployed in the cluster.
312 IBM Storage Ceph
Servers or nodes that can be freshly provisioned as OSD nodes
The admin keyring for the Ceph Monitor node, if you are redeploying an existing Ceph OSD node.
Procedure
1. On the bootstrapped node, change the value of bluestore_min_alloc_size parameter:
Syntax
ceph config set osd.OSD_ID bluestore_min_alloc_size_DEVICE_NAME VALUE
Example
[ceph: root@host01 /]# ceph config set osd.4 bluestore_min_alloc_size_hdd 8192
You can see bluestore_min_alloc_size is set to 8192 bytes, which is equivalent to 8 KiB.
Note: The selected values should be power of 2 aligned.
2. Restart the OSD’s service.
Syntax
systemctl restart SERVICE_ID
Example
[ceph: root@host01 /]# systemctl restart ceph-499829b4-832f-11eb-8d6d-001a4a000635@osd.4.service
Verification
Verify the setting using the ceph daemon command:
Syntax
ceph daemon osd.OSD_ID config get bluestore_min_alloc_size_DEVICE
Example
[ceph: root@host01 /]# ceph daemon osd.4 config get bluestore_min_alloc_size_hdd
ceph daemon osd.4 config get bluestore_min_alloc_size
{
"bluestore_min_alloc_size": "8192"
}
Reference
For OSD removal and addition information, see Managing OSDs.
Note: For already deployed OSDs, you cannot modify the bluestore_min_alloc_size parameter so you have to remove the OSDs and freshly deploy them again.
Resharding RocksDB database
You can reshard the database with the BlueStore admin tool. It transforms BlueStore’s RocksDB database from one shape to another into several column families without
redeploying the OSDs. Column families have the same features as the whole database, but allows users to operate on smaller data sets and apply different options. It
leverages the different expected lifetime of keys stored. The keys are moved during the transformation without creating new keys or deleting existing keys.
There are two ways to reshard the OSD:
1. Use the rocksdb-resharding.yml playbook.
2. Manually reshard the OSDs.
Prerequisites
A running IBM Storage Ceph cluster.
The object store configured as BlueStore.
OSD nodes deployed on the hosts.
Root level access to the all the hosts.
The ceph-common and cephadm packages instaled on all the hosts.
Use the rocksdb-resharding.yml playbook
1. As a root user, on the administration node, navigate to the cephadm folder where the playbook is installed.
Example
[root@host01 ~]# cd /usr/share/cephadm-ansible
2. Run the playbook.
IBM Storage Ceph 313
Syntax
ansible-playbook -i hosts rocksdb-resharding.yml -e osd_id=OSD_ID -e admin_node=HOST_NAME
For example,
[root@host01 ~]# ansible-playbook -i hosts rocksdb-resharding.yml -e osd_id=7 -e admin_node=host03
...............
TASK [stop the osd]
**************************************************************************************************************************
*********************************************************************
Wednesday 29 November 2023 11:25:18 +0000 (0:00:00.037)
0:00:03.864 ****
changed: [localhost -> host03]
TASK [set_fact ceph_cmd]
**************************************************************************************************************************
****************************************************************
Wednesday 29 November 2023 11:25:32 +0000 (0:00:14.128)
0:00:17.992 ****
ok: [localhost -> host03]
TASK [check fs consistency with fsck before resharding]
**************************************************************************************************************************
*********************************
Wednesday 29 November 2023 11:25:32 +0000 (0:00:00.041)
0:00:18.034 ****
ok: [localhost -> host03]
TASK [show current sharding]
**************************************************************************************************************************
************************************************************
Wednesday 29 November 2023 11:25:43 +0000 (0:00:11.053)
0:00:29.088 ****
ok: [localhost -> host03]
TASK [reshard]
**************************************************************************************************************************
**************************************************************************
Wednesday 29 November 2023 11:25:45 +0000 (0:00:01.446)
0:00:30.534 ****
ok: [localhost -> host03]
TASK [check fs consistency with fsck after resharding]
**************************************************************************************************************************
**********************************
Wednesday 29 November 2023 11:25:46 +0000 (0:00:01.479)
0:00:32.014 ****
ok: [localhost -> host03]
TASK [restart the osd]
**************************************************************************************************************************
******************************************************************
Wednesday 29 November 2023 11:25:57 +0000 (0:00:10.699)
0:00:42.714 ****
changed: [localhost -> host03]
3. Verify that the resharding is complete.
a. Stop the OSD that is resharded.
[ceph: root@host01 /]# ceph orch daemon stop osd.7
b. Enter the OSD container.
[root@host03 ~]# cephadm shell --name osd.7
c. Check for resharding.
[ceph: root@host03 /]# ceph-bluestore-tool --path /var/lib/ceph/osd/ceph-7/ show-sharding
m(3) p(3,0-12) O(3,0-13) L P
d. Start the OSD.
[ceph: root@host01 /]# ceph orch daemon start osd.7
Manually resharding the OSDs
1. Log into the cephadm shell:
Example
[root@host01 ~]# cephadm shell
2. Fetch the OSD_ID and the host details from the administration node:
Example
[ceph: root@host01 /]# ceph orch ps
3. Log into the respective host as a root user and stop the OSD:
Syntax
cephadm unit --name OSD_ID stop
Example
[root@host02 ~]# cephadm unit --name osd.0 stop
4. Enter into the stopped OSD daemon container:
314 IBM Storage Ceph
Syntax
cephadm shell --name OSD_ID
Example
[root@host02 ~]# cephadm shell --name osd.0
5. Log into the cephadm shell and check the file system consistency:
Syntax
ceph-bluestore-tool --path/var/lib/ceph/osd/ceph-OSD_ID/ fsck
Example
[ceph: root@host02 /]# ceph-bluestore-tool --path /var/lib/ceph/osd/ceph-0/ fsck
fsck success
6. Check the sharding status of the OSD node:
Syntax
ceph-bluestore-tool --path /var/lib/ceph/osd/ceph-OSD_ID/ show-sharding
Example
[ceph: root@host02 /]# ceph-bluestore-tool --path /var/lib/ceph/osd/ceph-6/ show-sharding
m(3) p(3,0-12) O(3,0-13) L P
7. Run the ceph-bluestore-tool command to reshard. IBM recommends to use the parameters as given in the command:
Syntax
ceph-bluestore-tool --log-level 10 -l log.txt --path /var/lib/ceph/osd/ceph-OSD_ID/ --sharding="m(3) p(3,0-12) O(3,013)=block_cache={type=binned_lru} L P" reshard
Example
[ceph: root@host02 /]# ceph-bluestore-tool --path /var/lib/ceph/osd/ceph-6/ --sharding="m(3) p(3,0-12) O(3,013)=block_cache={type=binned_lru} L P" reshard
reshard success
8. To check the sharding status of the OSD node, run the show-sharding command:
Syntax
ceph-bluestore-tool --path /var/lib/ceph/osd/ceph-OSD_ID/ show-sharding
Example
[ceph: root@host02 /]# ceph-bluestore-tool --path /var/lib/ceph/osd/ceph-6/ show-sharding
m(3) p(3,0-12) O(3,0-13)=block_cache={type=binned_lru} L P
9. Exit from the cephadm shell:
[ceph: root@host02 /]# exit
10. Log into the respective host as a root user and start the OSD:
Syntax
cephadm unit --name OSD_ID start
Example
[root@host02 ~]# cephadm unit --name osd.0 start
Reference
For more information, see Initial installation.
BlueStore fragmentation tool
As a storage administrator, you will want to periodically check the fragmentation level of your BlueStore OSDs. You can check fragmentation levels with one simple
command for offline or online OSDs.
For BlueStore OSDs, the free space gets fragmented over time on the underlying storage device. Some fragmentation is normal, but when there is excessive fragmentation
this causes poor performance.
The BlueStore fragmentation tool generates a score on the fragmentation level of the BlueStore OSD. This fragmentation score is given as a range, 0 through 1. A score of
0 means no fragmentation, and a score of 1 means severe fragmentation.
Table 1. Fragmentation scores' meaning
IBM Storage Ceph 315
Score
Fragmentation Amount
0.0 - 0.4
None to tiny fragmentation.
0.4 - 0.7
Small and acceptable fragmentation.
0.7 - 0.9
Considerable, but safe fragmentation.
0.9 - 1.0
Severe fragmentation and that causes performance issues.
Important: For help resolving severe fragmentation, contact IBM Support.
Checking for fragmentation
Check the fragmentation level of BlueStore OSDs while the BlueStore OSD process is either online or offline.
Checking for fragmentation
Check the fragmentation level of BlueStore OSDs while the BlueStore OSD process is either online or offline.
Before you begin
A running IBM Storage Ceph cluster.
BlueStore OSDs.
Getting an online BlueStore fragmentation score
Procedure
Inspect a running BlueStore OSD process by getting either a simple or more detailed report.
Get a simple report.
ceph daemon OSD_ID bluestore allocator score block
For example,
[ceph: root@host01 /]# ceph daemon osd.123 bluestore allocator score block
Get a detailed report.
ceph daemon OSD_ID bluestore allocator dump block
For example,
[ceph: root@host01 /]# ceph daemon osd.123 bluestore allocator dump block
Getting an offline BlueStore fragmentation score
Procedure
1. Reshard to check the offline BlueStore OSD.
For more information about resharding, see Resharding RocksDB database.
[root@host01 ~]# cephadm shell --name osd.ID
For example,
[root@host01 ~]# cephadm shell --name osd.2
Inferring fsid 110bad0a-bc57-11ee-8138-fa163eb9ffc2
Inferring config /var/lib/ceph/110bad0a-bc57-11ee-8138-fa163eb9ffc2/osd.2/config
Using ceph image with id `17334f841482` and tag `cp.stg.icr.io/cp/ibm-ceph/ceph-6-rhel9:7-22/ibmceph@sha256:09fc3e5baf198614d70669a106eb87dbebee16d4e91484375778d4adbccadacd
2. Inspect a non-running BlueStore OSD process.
Inspect a running BlueStore OSD process by getting either a simple or more detailed report.
Get a simple report.
ceph-bluestore-tool --path PATH_TO_OSD_DATA_DIRECTORY --allocator block free-score
For example,
[root@host01 /]# ceph-bluestore-tool --path /var/lib/ceph/osd/ceph-123 --allocator block free-score
Get a detailed report.
ceph-bluestore-tool --path PATH_TO_OSD_DATA_DIRECTORY --allocator block free-dump
block:
{
"fragmentation_rating": 0.018290238194701977
}
For example,
[root@host01 /]# ceph-bluestore-tool --path /var/lib/ceph/osd/ceph-123 --allocator block free-dump
block:
{
"capacity": 21470642176,
"alloc_unit": 4096,
316 IBM Storage Ceph
"alloc_type": "hybrid",
"alloc_name": "block",
"extents": [
{
"offset": "0x370000",
"length": "0x20000"
},
{
"offset": "0x3a0000",
"length": "0x10000"
},
{
"offset": "0x3f0000",
"length": "0x20000"
},
{
"offset": "0x460000",
"length": "0x10000"
},
Ceph BlueStore BlueFS
BlueStore block database stores metadata as key-value pairs in a RocksDB database. The block database resides on a small BlueFS partition on the storage device.
BlueFS is a minimal file system that is designed to hold the RocksDB files.
Viewing the bluefs_buffered_io setting
As a storage administrator, you can view the current setting for the bluefs_buffered_io parameter.
Viewing Ceph BlueFS statistics for Ceph OSDs
View the BluesFS related information about collocated and non-collocated Ceph OSDs with the bluefs stats command.
BlueFS files
There are three types of files that RocksDB produces.
Control files, for example CURRENT, IDENTITY, and MANIFEST-00011.
Database (DB) table files, for example 004112.sst.
Write ahead logs (WAL), for example 00038.log.
There is also an internal, hidden file that serves as BlueFS replay log, ino
1, that works as directory structure, file mapping, and operations log.
Fallback hierarchy
With BlueFS it is possible to put any file on any device. Parts of file can even reside on different devices, that is WAL, DB, and SLOW. There is an order to where BlueFS puts
files. File is put to secondary storage only when primary storage is exhausted, and tertiary only when secondary is exhausted.
The order for the specific files is as follows, for each device type.
Write ahead logs
WAL, DB, SLOW
Replay log ino 1
DB, SLOW
Control and DB files
DB, SLOW
Control and DB file order when running out of space: SLOW
Important: There is an exception to control and DB file order. When RocksDB detects that you are running out of space on DB file, it directly notifies you to put file to
SLOW device.
Viewing the bluefs_buffered_io setting
As a storage administrator, you can view the current setting for the bluefs_buffered_io parameter.
About this task
The option bluefs_buffered_io is set to True by default for IBM Storage Ceph. This option enable BlueFS to perform buffered reads in some cases, and enables the kernel
page cache to act as a secondary cache for reads like RocksDB block reads.
Important: Changing the value of bluefs_buffered_io is not recommended. Before changing the bluefs_buffered_io parameter, contact your IBM Support account team.
Before you begin
A running IBM Storage Ceph cluster.
Root-level access to the Ceph Monitor node.
Log into the Cephadm shell, using the cephadm shell command.
Procedure
IBM Storage Ceph 317
View the current value of the bluefs_buffered_io parameter, using one of the following procedures.
View the value stored in the configuration database.
ceph config get osd bluefs_buffered_io
For example,
[ceph: root@host01 /]# ceph config get osd bluefs_buffered_io
View the value stored in the configuration database for a specific OSD.
ceph config get OSD_ID bluefs_buffered_io
For example,
[ceph: root@host01 /]# ceph config get osd.2 bluefs_buffered_io
View the running value for an OSD where the running value is different from the value stored in the configuration database.
ceph config show OSD_ID bluefs_buffered_io
For example,
[ceph: root@host01 /]# ceph config show osd.3 bluefs_buffered_io
Viewing Ceph BlueFS statistics for Ceph OSDs
View the BluesFS related information about collocated and non-collocated Ceph OSDs with the bluefs stats command.
About this task
For more information about BlueStore devices, see BlueStore devices.
Before you begin
A running IBM Storage Ceph cluster.
Root-level access to the OSD node.
The object store configured as BlueStore.
Log into the Cephadm shell, using the cephadm shell command.
Procedure
View the BlueStore OSD statistics.
ceph daemon osd.OSD_ID bluefs stats
Example for collocated OSDs
[ceph: root@host01 /]# ceph daemon osd.1 bluefs stats
1 : device size 0x3bfc00000 : using 0x1a428000(420 MiB)
wal_total:0, db_total:15296836403, slow_total:0
Example for non-collocated OSDs
[ceph: root@host01 /]# ceph daemon osd.1 bluefs stats
0 :
1 : device size 0x1dfbfe000 : using 0x1100000(17 MiB)
2 : device size 0x27fc00000 : using 0x248000(2.3 MiB)
RocksDBBlueFSVolumeSelector: wal_total:0, db_total:7646425907, slow_total:10196562739, db_avail:935539507
Usage matrix:
DEV/LEV
WAL
DB
SLOW
*
*
REAL
FILES
LOG
0 B
4 MiB
0 B
0 B
0 B
756 KiB
1
WAL
0 B
4 MiB
0 B
0 B
0 B
3.3 MiB
1
DB
0 B
9 MiB
0 B
0 B
0 B
76 KiB
10
SLOW
0 B
0 B
0 B
0 B
0 B
0 B
0
TOTALS
0 B
17 MiB
0 B
0 B
0 B
0 B
12
MAXIMUMS:
LOG
0 B
4 MiB
0 B
0 B
0 B
756 KiB
WAL
0 B
4 MiB
0 B
0 B
0 B
3.3 MiB
DB
0 B
11 MiB
0 B
0 B
0 B
112 KiB
SLOW
0 B
0 B
0 B
0 B
0 B
0 B
TOTALS
0 B
17 MiB
0 B
0 B
0 B
0 B
In this example,
0
1
2
Refers to the dedicated WAL device, which is block.wal.
Refers to the dedicated DB device, which is block.db.
Refers to the main block device, which is block or slow.
device size
Represents an actual size of the device.
318 IBM Storage Ceph
using
Represents the total usage. It is not restricted to BlueFS.
Note: DB and WAL devices are used only by BlueFS. For a main device, usage from stored BlueStore data is also included. In this example, 2.3 MiB is the data
from BlueStore.
wal_total, db_total, and slow_total
Values that reiterate the device values previously stated.
db_avail
Represents how many bytes can be taken from the SLOW device, if necessary.
Usage matrix:
Rows WAL, DB, and SLOW
Describes where the specific file was intended to be put.
Row LOG
Describes the BlueFS replay log ino 1.
Columns WAL, DB, and SLOW
Describes where data is actually put. The values are in allocation units. WAL and DB have bigger allocation units for performance reasons.
Columns *
Relate to virtual devices new-db and new-wal that are used for ceph-bluestore-tool. It should always show 0 B.
Column REAL
Shows actual usage in bytes.
Column FILES
Shows count of files.
MAXIMUMS
This table captures the maximum value of each entry from the usage matrix.
Cephadm troubleshooting
As a storage administrator, you can troubleshoot the IBM Storage Ceph cluster. Sometimes there is a need to investigate why a Cephadm command failed or why a specific
service does not run properly.
Pause or disable cephadm
Per service and per daemon event
Check cephadm logs
Gather log files
Gather log files to help troubleshoot Cephadm.
Collect systemd status
List all downloaded container images
Manually run containers
CIDR network error
Access admin socket
Manually deploying a mgr daemon
Prerequisites
A running IBM Storage Ceph cluster.
Pause or disable cephadm
If Cephadm does not behave as expected, you can pause most of the background activity with the following commands:
Example
[ceph: root@host01 /]# ceph orch pause
This stops any changes, but Cephadm periodically checks hosts to refresh it’s inventory of daemons and devices.
If you want to disable Cephadm completely, run the following commands:
Example
[ceph: root@host01 /]# ceph orch set backend ''
[ceph: root@host01 /]# ceph mgr module disable cephadm
Note that previously deployed daemon containers continue to exist and start as they did before.
To re-enable Cephadm in the cluster, run the following commands:
Example
[ceph: root@host01 /]# ceph mgr module enable cephadm
[ceph: root@host01 /]# ceph orch set backend cephadm
Per service and per daemon event
IBM Storage Ceph 319
Cephadm stores events per service and per daemon in order to aid in debugging failed daemon deployments. These events often contain relevant information:
Per service
Syntax
ceph orch ls --service_name SERVICE_NAME --format yaml
Example
[ceph: root@host01 /]# ceph orch ls --service_name alertmanager --format yaml
service_type: alertmanager
service_name: alertmanager
placement:
hosts:
- unknown_host
status:
...
running: 1
size: 1
events:
- 2021-02-01T08:58:02.741162 service:alertmanager [INFO] "service was created"
- '2021-02-01T12:09:25.264584 service:alertmanager [ERROR] "Failed to apply: Cannot
place <AlertManagerSpec for service_name=alertmanager> on unknown_host: Unknown hosts"'
Per daemon
Syntax
ceph orch ps --service-name SERVICE_NAME --daemon-id DAEMON_ID --format yaml
Example
[ceph: root@host01 /]# ceph orch ps --service-name mds --daemon-id cephfs.hostname.ppdhsz --format yaml
daemon_type: mds
daemon_id: cephfs.hostname.ppdhsz
hostname: hostname
status_desc: running
...
events:
- 2021-02-01T08:59:43.845866 daemon:mds.cephfs.hostname.ppdhsz [INFO] "Reconfigured
mds.cephfs.hostname.ppdhsz on host 'hostname'"
Check cephadm logs
You can monitor the Cephadm log in real time with the following command:
Example
[ceph: root@host01 /]# ceph -W cephadm
You can see the last few messages with the following command:
Example
[ceph: root@host01 /]# ceph log last cephadm
If you have enabled logging to files, you can see a Cephadm log file called ceph.cephadm.log on the monitor hosts.
Gather log files
Gather log files to help troubleshoot Cephadm.
Note:
Be sure to run all log file commands outside the cephadm shell.
By default, Cephadm stores logs in journald.
Table 1. Gathering Cephadm ogs
Action needed
Gather log files
for all daemons
Command
Example
Notes
journalctl
Read the log file cephadm logs --name DAEMON_NAME
of a specific
daemon
[root@host01 ~]# cephadm logs --name
cephfs.hostname.ppdhsz
This command
works when run
on the same hosts
where the daemon
is running.
Read the log file cephadm logs --fsid FSID --name
DAEMON_NAME
of a specific
daemon running
on a different
host
[root@host01 ~]# cephadm logs --fsid 2d2fd136-6df111ea-ae74-002590e526e8 --name cephfs.hostname.ppdhsz
FSID is the cluster
ID provided by the
ceph status
command.
320 IBM Storage Ceph
Action needed
Fetch all log
files of all the
daemons on a
given host
Command
for name in $(cephadm ls | python3 -c
"import sys, json; [print(i[name]) for i
in json.load(sys.stdin)]") ; do cephadm
logs --fsid FSID_OF_CLUSTER --name "$name"
> $name; done
Example
[root@host01 ~]# for name in $(cephadm ls | python3
-c "import sys, json; [print(i['name']) for i in
json.load(sys.stdin)]") ; do cephadm logs --fsid
57bddb48-ee04-11eb-9962-001a4a000672 --name "$name"
> $name; done
Notes
Collect systemd status
To print the state of a systemd unit, run the following command:
Example
[root@host01 ~]$ systemctl status ceph-a538d494-fb2a-48e4-82c8-b91c37bb0684@mon.host01.service
List all downloaded container images
To list all the container images that are downloaded on a host, run the following command:
Example
[ceph: root@host01 /]# podman ps -a --format json | jq '.[].Image'
"docker.io/library/rhel8"
"cp.icr.io/cp/ibm-ceph/ceph-6-rhel9:latest"
Manually run containers
Cephadm writes small wrappers that runs a container. Refer to /var/lib/ceph/CLUSTER_FSID/SERVICE_NAME/unit` to run the container execution command.
Analyzing SSH errors
If you get the following error:
Example
execnet.gateway_bootstrap.HostNotFound: -F /tmp/cephadm-conf-73z09u6g -i /tmp/cephadm-identity-ky7ahp_5 root@10.10.1.2
...
raise OrchestratorError(msg) from e
orchestrator._interface.OrchestratorError: Failed to connect to 10.10.1.2 (10.10.1.2).
Please make sure that the host is reachable and accepts connections using the cephadm SSH key
Try the following options to troubleshoot the issue:
To ensure Cephadm has a SSH identity key, run the following command:
Example
[ceph: root@host01 /]# ceph config-key get mgr/cephadm/ssh_identity_key > ~/cephadm_private_key
INFO:cephadm:Inferring fsid f8edc08a-7f17-11ea-8707-000c2915dd98
INFO:cephadm:Using recent ceph image docker.io/ceph/ceph:v15 obtained 'mgr/cephadm/ssh_identity_key'
[root@mon1 ~] # chmod 0600 ~/cephadm_private_key
If the above command fails, Cephadm does not have a key. To generate a SSH key, run the following command:
Example
[ceph: root@host01 /]# chmod 0600 ~/cephadm_private_key
Or
Example
[ceph: root@host01 /]# cat ~/cephadm_private_key | ceph cephadm set-ssk-key -iTo ensure that the SSH configuration is correct, run the following command:
Example
[ceph: root@host01 /]# ceph cephadm get-ssh-config
To verify the connection to the host, run the following command:
Example
[ceph: root@host01 /]# ssh -F config -i ~/cephadm_private_key root@host01
Verify public key is in authorized_keys.
To verify that the public key is in the authorized_keys file, run the following commands:
IBM Storage Ceph 321
Example
[ceph: root@host01 /]# ceph cephadm get-pub-key
[ceph: root@host01 /]# grep "`cat ~/ceph.pub`" /root/.ssh/authorized_keys
CIDR network error
Classless inter domain routing (CIDR) also known as supernetting, is a method of assigning Internet Protocol (IP) addresses, the Cephadm log entries shows the current
state that improves the efficiency of address distribution and replaces the previous system based on Class A, Class B and Class C networks. If you see one of the following
errors:
ERROR: Failed to infer CIDR network for mon ip ***; pass --skip-mon-network to configure it
later
Or
Must set public_network config option or specify a CIDR network, ceph addrvec, or plain
IP
You need to run the following command:
Example
[ceph: root@host01 /]# ceph config set host public_network hostnetwork
Access admin socket
Each Ceph daemon provides an admin socket that bypasses the MONs.
To access the admin socket, enter the daemon container on the host:
Example
[ceph: root@host01 /]# cephadm enter --name cephfs.hostname.ppdhsz
[ceph: root@mon1 /]# ceph --admin-daemon /var/run/ceph/ceph-cephfs.hostname.ppdhsz.asok config show
Manually deploying a mgr daemon
Cephadm requires a mgr daemon in order to manage the IBM Storage Ceph cluster. In case the last mgr daemon of an IBM Storage Ceph cluster was removed, you can
manually deploy a mgr daemon, on a random host of the IBM Storage Ceph cluster.
Prerequisites
A running IBM Storage Ceph cluster.
Root-level access to all the nodes.
Hosts are added to the cluster.
Procedure
1. Log into the Cephadm shell:
Example
[root@host01 ~]# cephadm shell
2. Disable the Cephadm scheduler to prevent Cephadm from removing the new MGR daemon, with the following command:
Example
[ceph: root@host01 /]# ceph config-key set mgr/cephadm/pause true
3. Get or create the auth entry for the new MGR daemon:
Example
[ceph: root@host01 /]# ceph auth get-or-create mgr.host01.smfvfd1 mon "profile mgr" osd "allow *" mds "allow *"
[mgr.host01.smfvfd1]
key = AQDhcORgW8toCRAAlMzlqWXnh3cGRjqYEa9ikw==
4. Open ceph.conf file:
Example
[ceph: root@host01 /]# ceph config generate-minimal-conf
# minimal ceph.conf for 8c9b0072-67ca-11eb-af06-001a4a0002a0
[global]
322 IBM Storage Ceph
fsid = 8c9b0072-67ca-11eb-af06-001a4a0002a0
mon_host = [v2:10.10.200.10:3300/0,v1:10.10.200.10:6789/0] [v2:10.10.10.100:3300/0,v1:10.10.200.100:6789/0]
5. Get the container image:
Example
[ceph: root@host01 /]# ceph config get "mgr.host01.smfvfd1" container_image
6. Create a config-json.json file and add the following:
Note: Use the values from the output of the ceph config generate-minimal-conf command.
Example
{
{
"config": "# minimal ceph.conf for 8c9b0072-67ca-11eb-af06-001a4a0002a0\n[global]\n\tfsid = 8c9b0072-67ca-11eb-af06001a4a0002a0\n\tmon_host = [v2:10.10.200.10:3300/0,v1:10.10.200.10:6789/0]
[v2:10.10.10.100:3300/0,v1:10.10.200.100:6789/0]\n",
"keyring": "[mgr.Ceph5-2.smfvfd1]\n\tkey = AQDhcORgW8toCRAAlMzlqWXnh3cGRjqYEa9ikw==\n"
}
}
7. Exit from the Cephadm shell:
Example
[ceph: root@host01 /]# exit
8. Deploy the MGR daemon:
Example
[root@host01 ~]# cephadm --image cp.icr.io/cp/ibm-ceph/ceph-6-rhel9:latest deploy --fsid
001a4a0002a0 --name mgr.host01.smfvfd1 --config-json config-json.json
8c9b0072-67ca-11eb-af06-
Verification
In the Cephadm shell, run the following command:
Example
[ceph: root@host01 /]# ceph -s
You can see a new mgr daemon has been added.
Cephadm operations
As a storage administrator, you can carry out Cephadm operations in the IBM Storage Ceph cluster.
Monitor cephadm log messages
Ceph daemon logs
Data location
Cephadm health checks
Prerequisites
A running IBM Storage Ceph cluster.
Monitor cephadm log messages
Cephadm logs to the cephadm cluster log channel so you can monitor progress in real time.
To monitor progress in real time, run the following command:
Example
[ceph: root@host01 /]# ceph -W cephadm
Example
2022-11-02T17:51:36.335728+0000 mgr.Ceph5-1.nqikfh [INF] refreshing Ceph5-adm facts
2022-11-02T17:51:37.170982+0000 mgr.Ceph5-1.nqikfh [INF] deploying 1 monitor(s) instead of 2 so monitors may achieve
consensus
2022-11-02T17:51:37.173487+0000 mgr.Ceph5-1.nqikfh [ERR] It is NOT safe to stop ['mon.Ceph5-adm']: not enough monitors
would be available (Ceph5-2) after stopping mons [Ceph5-adm]
2022-11-02T17:51:37.179658+0000 mgr.Ceph5-1.nqikfh [INF] Found osd claims -> {}
2022-11-02T17:51:37.180116+0000 mgr.Ceph5-1.nqikfh [INF] Found osd claims for drivegroup all-available-devices -> {}
2022-11-02T17:51:37.182138+0000 mgr.Ceph5-1.nqikfh [INF] Applying all-available-devices on host Ceph5-adm...
2022-11-02T17:51:37.182987+0000 mgr.Ceph5-1.nqikfh [INF] Applying all-available-devices on host Ceph5-1...
2022-11-02T17:51:37.183395+0000 mgr.Ceph5-1.nqikfh [INF] Applying all-available-devices on host Ceph5-2...
2022-11-02T17:51:43.373570+0000 mgr.Ceph5-1.nqikfh [INF] Reconfiguring node-exporter.Ceph5-1 (unknown last config time)...
2022-11-02T17:51:43.373840+0000 mgr.Ceph5-1.nqikfh [INF] Reconfiguring daemon node-exporter.Ceph5-1 on Ceph5-1
IBM Storage Ceph 323
By default, the log displays info-level events and above. To see the debug-level messages, run the following commands:
Example
[ceph: root@host01 /]# ceph config set mgr mgr/cephadm/log_to_cluster_level debug
[ceph: root@host01 /]# ceph -W cephadm --watch-debug
[ceph: root@host01 /]# ceph -W cephadm --verbose
Return debugging level to default info:
Example
[ceph: root@host01 /]# ceph config set mgr mgr/cephadm/log_to_cluster_level info
To see the recent events, run the following command:
Example
[ceph: root@host01 /]# ceph log last cephadm
Theses events are also logged to ceph.cephadm.log file on the monitor hosts and to the monitor daemon’s stderr.
Ceph daemon logs
You can view the Ceph daemon logs through stderr or files.
Logging to stdout
Traditionally, Ceph daemons have logged to /var/log/ceph. By default, Cephadm daemons log to stderr and the logs are captured by the container runtime
environment. For most systems, by default, these logs are sent to journald and accessible through the journalctl command.
For example, to view the logs for the daemon on host01 for a storage cluster with ID 5c5a50ae-272a-455d-99e9-32c6a013e694:
Example
[ceph: root@host01 /]# journalctl -u ceph-5c5a50ae-272a-455d-99e9-32c6a013e694@host01
This works well for normal Cephadm operations when logging levels are low.
To disable logging to stderr, set the following values:
Example
[ceph: root@host01 /]# ceph config set global log_to_stderr false
[ceph: root@host01 /]# ceph config set global mon_cluster_log_to_stderr false
Logging to files
You can also configure Ceph daemons to log to files instead of stderr. When logging to files, Ceph logs are located in /var/log/ceph/CLUSTER_FSID.
To enable logging to files, set the follwing values:
Example
[ceph: root@host01 /]# ceph config set global log_to_file true
[ceph: root@host01 /]# ceph config set global mon_cluster_log_to_file true
Important: Log rotation to a non-default path is not currently supported.
Note: Disable logging to stderr to avoid double logs.
By default, Cephadm sets up log rotation on each host to rotate these files. You can configure the logging retention schedule by modifying
/etc/logrotate.d/ceph.CLUSTER_FSID.
Data location
Cephadm daemon data and logs are located in slightly different locations than the older versions of Ceph:
/var/log/ceph/CLUSTER_FSID contains all the storage cluster logs. Note
that by default Cephadm logs through stderr` and the container runtime, so these logs are usually not present.
/var/lib/ceph/CLUSTER_FSID contains all the cluster daemon data, besides logs.
/var/lib/ceph/CLUSTER_FSID/DAEMON_NAME contains all the data for a specific daemon.
/var/lib/ceph/CLUSTER_FSID/crash` contains the crash reports for the storage cluster.
/var/lib/ceph/CLUSTER_FSID/removed` contains old daemon data directories for the stateful daemons, for example monitor or Prometheus, that have been
removed by Cephadm.
Disk usage
A few Ceph daemons may store a significant amount of data in /var/lib/ceph, notably the monitors and Prometheus daemon, hence move this directory to its own disk,
partition, or logical volume so that the root file system is not filled up.
324 IBM Storage Ceph
Cephadm health checks
As a storage administrator, you can monitor the IBM Storage Ceph cluster with the additional health checks provided by the Cephadm module. This is supplementary to
the default health checks provided by the storage cluster.
Cephadm operations health checks
Cephadm configuration health checks
Cephadm operations health checks
Health checks are executed when the Cephadm module is active. You can get the following health warnings:
CEPHADM_PAUSED
Cephadm background work is paused with the ceph orch pause command. Cephadm continues to perform passive monitoring activities such as checking the host and
daemon status, but it does not make any changes like deploying or removing daemons. You can resume Cephadm work with the ceph orch resume command.
CEPHADM_STRAY_HOST
One or more hosts have running Ceph daemons but are not registered as hosts managed by the Cephadm module. This means that those services are not currently
managed by Cephadm, for example, a restart and upgrade that is included in the ceph orch ps command. You can manage the host(s) with the ceph orch host add
HOST_NAME command but ensure that SSH access to the remote hosts is configured. Alternatively, you can manually connect to the host and ensure that services on that
host are removed or migrated to a host that is managed by Cephadm. You can also disable this warning with the setting ceph config set mgr
mgr/cephadm/warn_on_stray_hosts false
CEPHADM_STRAY_DAEMON
One or more Ceph daemons are running but are not managed by the Cephadm module. This might be because they were deployed using a different tool, or because they
were started manually. Those services are not currently managed by Cephadm, for example, a restart and upgrade that is included in the ceph orch ps command.
If the daemon is a stateful one that is a monitor or OSD daemon, these daemons should be adopted by Cephadm. For stateless daemons, you can provision a new daemon
with the ceph orch
apply command and then stop the unmanaged daemon.
You can disable this health warning with the setting ceph config set mgr
mgr/cephadm/warn_on_stray_daemons false.
CEPHADM_HOST_CHECK_FAILED
One or more hosts have failed the basic Cephadm host check, which verifies that:name: value
The host is reachable and you can execute Cephadm.
The host meets the basic prerequisites, like a working container runtime that is Podman , and working time synchronization. If this test fails, Cephadm wont be able
to manage the services on that host.
You can manually run this check with the ceph cephadm check-host HOST_NAME command. You can remove a broken host from management with the ceph orch
host rm
HOST_NAME command. You can disable this health warning with the setting ceph config
set mgr mgr/cephadm/warn_on_failed_host_check false.
Cephadm configuration health checks
Cephadm periodically scans each of the hosts in the storage cluster, to understand the state of the OS, disks, and NICs . These facts are analyzed for consistency across
the hosts in the storage cluster to identify any configuration anomalies. The configuration checks are an optional feature.
You can enable this feature with the following command:
Example
[ceph: root@host01 /]# ceph config set mgr mgr/cephadm/config_checks_enabled true
The configuration checks are triggered after each host scan, which is for a duration of one minute.
The ceph -W cephadm command shows log entries of the current state and outcome of the configuration checks as follows:
Disabled state
Example
ALL cephadm checks are disabled, use ceph config set mgr
mgr/cephadm/config_checks_enabled true command to enable
Enabled state
Example
CEPHADM 8/8 checks enabled and executed (0 bypassed, 0 disabled). No issues detected
The configuration checks themselves are managed through several cephadm subcommands.
IBM Storage Ceph 325
To determine whether the configuration checks are enabled, run the following command:
Example
[ceph: root@host01 /]# ceph cephadm config-check status
This command returns the status of the configuration checker as either Enabled or Disabled.
To list all the configuration checks and their current state, run the following command:
Example
[ceph: root@host01 /]# ceph cephadm config-check ls
NAME
HEALTHCHECK
STATUS
kernel_security CEPHADM_CHECK_KERNEL_LSM
enabled
hosts
os_subscription CEPHADM_CHECK_SUBSCRIPTION
enabled
public_network
CEPHADM_CHECK_PUBLIC_MEMBERSHIP enabled
osd_mtu_size
CEPHADM_CHECK_MTU
enabled
osd_linkspeed
CEPHADM_CHECK_LINKSPEED
enabled
network_missing CEPHADM_CHECK_NETWORK_MISSING
enabled
Ceph hosts
ceph_release
CEPHADM_CHECK_CEPH_RELEASE
enabled
the same release (unless upgrade is active)
kernel_version
CEPHADM_CHECK_KERNEL_VERSION
enabled
consistent
DESCRIPTION
checks SELINUX/Apparmor profiles are consistent across cluster
checks subscription states are consistent for all cluster hosts
check that all hosts have a NIC on the Ceph public_netork
check that OSD hosts share a common MTU setting
check that OSD hosts share a common linkspeed
checks that the cluster/public networks defined exist on the
check for Ceph version consistency - ceph daemons should be on
checks that the MAJ.MIN of the kernel on Ceph hosts is
Each configuration check is described as follows:
CEPHADM_CHECK_KERNEL_LSM
Each host within the storage cluster is expected to operate within the same Linux Security Module (LSM) state. For example, if the majority of the hosts are running with
SELINUX in enforcing mode, any host not running in this mode would be flagged as an anomaly and a healthcheck with a warning state is raised.
CEPHADM_CHECK_SUBSCRIPTION
This check relates to the status of the vendor subscription. This check is only performed for hosts using Red Hat Enterprise Linux, but helps to confirm that all the hosts
are covered by an active subscription so that patches and updates are available.
CEPHADM_CHECK_PUBLIC_MEMBERSHIP
All members of the cluster should have NICs configured on at least one of the public network subnets. Hosts that are not on the public network will rely on routing which
may affect performance.
CEPHADM_CHECK_MTU
The maximum transmission unit (MTU) of the NICs on OSDs can be a key factor in consistent performance. This check examines hosts that are running OSD services to
ensure that the MTU is configured consistently within the cluster. This is determined by establishing the MTU setting that the majority of hosts are using, with any
anomalies resulting in a Ceph health check.
CEPHADM_CHECK_LINKSPEED
Similar to the MTU check, linkspeed consistency is also a factor in consistent cluster performance. This check determines the linkspeed shared by the majority of the OSD
hosts, resulting in a healthcheck for any hosts that are set at a lower linkspeed rate.
CEPHADM_CHECK_NETWORK_MISSING
The public_network and cluster_network settings support subnet definitions for IPv4 and IPv6. If these settings are not found on any host in the storage cluster a
healthcheck is raised.
CEPHADM_CHECK_CEPH_RELEASE Under normal operations, the Ceph cluster should be running daemons under the same Ceph release, for example all IBM Storage
Ceph cluster releases. This check looks at the active release for each daemon, and reports any anomalies as a healthcheck. This check is bypassed if an upgrade process
is active within the cluster.
CEPHADM_CHECK_KERNEL_VERSION
The OS kernel version is checked for consistency across the hosts. Once again, the majority of the hosts is used as the basis of identifying anomalies.
Using cephadm-ansible modules
Use cephadm-ansible modules in Ansible playbooks to administer your IBM Storage Ceph cluster.
The cephadm-ansible package provides several modules that wrap cephadm calls to let you write your own unique Ansible playbooks to administer your cluster.
Note: cephadm-ansible modules only support the most important tasks. Any operation not covered by cephadm-ansible modules must be completed using either
the command or shell Ansible modules in your playbooks.
cephadm-ansible module options
Bootstrapping storage cluster
Adding or removing hosts
Add and remove hosts in your storage cluster by using the ceph_orch_host module in your Ansible playbook.
Setting configuration options
Applying a service specification
Managing Ceph daemon states
cephadm-ansible modules
326 IBM Storage Ceph
The cephadm-ansible modules are a collection of modules that simplify writing Ansible playbooks by providing a wrapper around cephadm and ceph
orch commands. You can use the modules to write your own unique Ansible playbooks to administer your cluster using one or more of the modules.
The cephadm-ansible package includes the following modules:
cephadm_bootstrap
ceph_orch_host
ceph_config
ceph_orch_apply
ceph_orch_daemon
cephadm_registry_login
cephadm-ansible module options
The following tables list the available options for the cephadm-ansible modules. Options listed as required need to be set when using the modules in your Ansible
playbooks. Options listed with a default value of true indicate that the option is automatically set when using the modules and you do not need to specify it in your
playbook. For example, for the cephadm_bootstrap module, the Ceph Dashboard is installed unless you set dashboard: false.
Table 1. Available options for the cephadm_bootstrap module
Description
cephadm_bootstrap
Required
Default
mon_ip
Ceph Monitor IP address.
true
image
Ceph container image.
false
docker
Use docker instead of podman.
false
fsid
Define the Ceph FSID.
false
pull
Pull the Ceph container image.
false
true
dashboard
Deploy the Ceph Dashboard.
false
true
dashboard_user
Specify a specific Ceph Dashboard user.
false
dashboard_password
Ceph Dashboard password.
false
monitoring
Deploy the monitoring stack.
false
true
firewalld
Manage firewall rules with firewalld.
false
true
allow_overwrite
Allow overwrite of existing --output-config, --output-keyring, or --output-pub-ssh-key files. false
false
registry_url
URL for custom registry.
false
registry_username
Username for custom registry.
false
registry_password
Password for custom registry.
false
registry_json
JSON file with custom registry login information.
false
ssh_user
SSH user to use for cephadm ssh to hosts.
false
ssh_config
SSH config file path for cephadm SSH client.
false
allow_fqdn_hostname Allow hostname that is a fully-qualified domain name (FQDN).
false
cluster_network
false
Subnet to use for cluster replication, recovery and heartbeats.
false
Table 2. Available options for the ceph_orch_host module
ceph_orch_ho
st
fsid
The FSID of the Ceph cluster to interact with.
image
The Ceph container image to use.
Description
Required
Default
false
false
name
Name of the host to add, remove, or update.
true
address
IP address of the host.
true when
state is
present.
set_admin_la Set the _admin label on the specified host.
bel
labels
The list of labels to apply to the host.
state
If set to present, it ensures the name specified in name is present. If set to absent, it removes the host
false
false
false
[]
false
present
specified in name. If set to drain, it schedules to remove all daemons from the host specified in name.
Table 3. Available options for the ceph_config module
Description
ceph_config
Required
fsid
The FSID of the Ceph cluster to interact with.
false
image
The Ceph container image to use.
false
action
Whether to set or get the parameter specified in option. false
who
Which daemon to set the configuration to.
true
option
Name of the parameter to set or get.
true
value
Value of the parameter to set.
true if action is set
Default
set
Table 4. Available options for the ceph_orch_apply module
ceph_orch_apply
Description
Required
fsid
The FSID of the Ceph cluster to interact with. false
image
The Ceph container image to use.
false
spec
The service specification to apply.
true
IBM Storage Ceph 327
Table 5. Available options for the ceph_orch_daemon module
Description
ceph_orch_daemon
Required
fsid
The FSID of the Ceph cluster to interact with.
false
image
The Ceph container image to use.
false
state
The desired state of the service specified in name. true
If started, it ensures the service is started.
If stopped, it ensures the service is stopped.
If restarted, it will restart the service.
daemon_id
The ID of the service.
daemon_type
The type of service.
true
true
Table 6. Available options for the cephadm_registry_login module
cephadm_registry_
login
state
Login or logout of a registry.
docker
Use docker instead of podman.
Description
registry_url
Required
false
login
false
The URL for custom registry.
registry_username Username for custom registry.
false
registry_password Password for custom registry.
true when state is
login.
registry_json
Default
true when state is
login.
The path to a JSON file. This file must be present on remote hosts prior to running this task. This
option is currently not supported.
Bootstrapping storage cluster
As a storage administrator, you can bootstrap a storage cluster using Ansible by using the cephadm_bootstrap and cephadm_registry_login modules in your
Ansible playbook.
Prerequisites
An IP address for the first Ceph Monitor container, which is also the IP address for the first node in the storage cluster.
Login access to cp.icr.io/cp.
A minimum of 10 GB of free space for /var/lib/containers/.
Installation of the cephadm-ansible package on the Ansible administration node.
Passwordless SSH is set up on all hosts in the storage cluster.
For the latest supported Red Hat Enterprise Linux versions, see Compatibility matrix.
Procedure
1. Log in to the Ansible administration node.
2. Navigate to the /usr/share/cephadm-ansible directory on the Ansible administration node:
Example
[ansible@admin ~]$ cd /usr/share/cephadm-ansible
3. Create the hosts file and add hosts, labels, and monitor IP address of the first host in the storage cluster:
Syntax
sudo vi INVENTORY_FILE
HOST1 labels="['LABEL1', 'LABEL2']"
HOST2 labels="['LABEL1', 'LABEL2']"
HOST3 labels="['LABEL1']"
[admin]
ADMIN_HOST monitor_address=MONITOR_IP_ADDRESS labels="['ADMIN_LABEL', 'LABEL1', 'LABEL2']"
Example
[ansible@admin cephadm-ansible]$ sudo vi hosts
host02 labels="['mon', 'mgr']"
host03 labels="['mon', 'mgr']"
host04 labels="['osd']"
host05 labels="['osd']"
host06 labels="['osd']"
[admin]
host01 monitor_address=10.10.128.68 labels="['_admin', 'mon', 'mgr']"
328 IBM Storage Ceph
4. Run the preflight playbook:
Syntax
ansible-playbook -i INVENTORY_FILE cephadm-preflight.yml --extra-vars "ceph_origin=ibm"
Example
[ansible@admin cephadm-ansible]$ ansible-playbook -i hosts cephadm-preflight.yml --extra-vars "ceph_origin=ibm"
5. Create a playbook to bootstrap your cluster:
Syntax
sudo vi PLAYBOOK_FILENAME.yml
--- name: NAME_OF_PLAY
hosts: BOOTSTRAP_HOST
become: USE_ELEVATED_PRIVILEGES
gather_facts: GATHER_FACTS_ABOUT_REMOTE_HOSTS
tasks:
-name: NAME_OF_TASK
cephadm_registry_login:
state: STATE
registry_url: REGISTRY_URL
registry_username: REGISTRY_USER_NAME
registry_password: REGISTRY_PASSWORD
- name: NAME_OF_TASK
cephadm_bootstrap:
mon_ip: "{{ monitor_address }}"
dashboard_user: DASHBOARD_USER
dashboard_password: DASHBOARD_PASSWORD
allow_fqdn_hostname: ALLOW_FQDN_HOSTNAME
cluster_network: NETWORK_CIDR
Example
[ansible@admin cephadm-ansible]$ sudo vi bootstrap.yml
--- name: bootstrap the cluster
hosts: host01
become: true
gather_facts: false
tasks:
- name: login to registry
cephadm_registry_login:
state: login
registry_url: cp.icr.io/cp
registry_username: user1
registry_password: mypassword1
- name: bootstrap initial cluster
cephadm_bootstrap:
mon_ip: "{{ monitor_address }}"
dashboard_user: mydashboarduser
dashboard_password: mydashboardpassword
allow_fqdn_hostname: true
cluster_network: 10.10.128.0/28
6. Run the playbook:
Syntax
ansible-playbook -i INVENTORY_FILE PLAYBOOK_FILENAME.yml -vvv
Example
ansible@admin cephadm-ansible]$ ansible-playbook -i hosts bootstrap.yml -vvv
Verification
Review the Ansible output after running the playbook.
Adding or removing hosts
Add and remove hosts in your storage cluster by using the ceph_orch_host module in your Ansible playbook.
Prerequisites
A running IBM Storage Ceph cluster.
Register the nodes to the CDN and attach subscriptions.
Ansible user with sudo and passwordless SSH access to all nodes in the storage cluster.
IBM Storage Ceph 329
Installation of the cephadm-ansible package on the Ansible administration node.
New hosts have the storage cluster’s public SSH key.
For more information about copying the storage cluster's public SSH keys to new hosts, see Adding hosts.
Procedure
1. Use the following procedure to add new hosts t
0
You can add this document to your study collection(s)
Sign in Available only to authorized usersYou can add this document to your saved list
Sign in Available only to authorized users(For complaints, use another form )