-
Notifications
You must be signed in to change notification settings - Fork 80
Hare provisioning for single node deployment
Note: In case of Mini-Provisioning, consul will not be started/stopped by Hare so it is pre-requisite that consul should be running on all the nodes present
/opt/seagate/cortx/hare/bin/hare_setup --help# /opt/seagate/cortx/hare/bin/hare_setup --help
usage: hare_setup [-h]
{post_install,config,init,test,support_bundle,reset,cleanup,prepare,pre-upgrade,post-upgrade}
...
Configure hare settings
positional arguments:
{post_install,config,init,test,support_bundle,reset,cleanup,prepare,pre-upgrade,post-upgrade}
post_install Validates installation
config Configures Hare
init Initializes Hare
test Tests Hare component
support_bundle Generates support bundle
reset Resets temporary Hare data and configuration
cleanup Resets Hare configuration, logs & formats Motr
metadata
prepare Validates configuration pre-requisites
pre-upgrade Performs the Hare rpm pre-upgrade tasks
post-upgrade Performs the Hare rpm post-upgrade tasks
optional arguments:
-h, --help show this help message and exit
Hare setup utility (hare_setup) comprises of sub commands (as below) that help to provision cortx-hare and configure cortx-motr such that the corresponding hare and motr services are ready to start. This utility can be used to provision cortx-hare in virtual as well as physical cluster environments. Few points to note,
- hare_setup utility does not install the above mention pre requisites but assumes they are already installed and performs required validation checks.
- Although every sub command is idempotent, it has pre requisites of its own which are implemented in the form of validation checks.
- Due to the nature of operations performed, some of the sub commands have a limitation that they cannot be executed in parallel on multiple nodes in a cluster.
- Uses conf-store module provided by cortx-py-utils.
- Hare single node provisioning templates are available that can be used to generate required conf-store files.
- User needs to make sure that total number of data devices in a pool(storage set) are >=
cluster>CLUSTER_ID>storage_set[0]>durability>sns>data+cluster>CLUSTER_ID>storage_set[0]>durability>sns>parity+cluster>CLUSTER_ID>storage_set[0]>durability>sns>spare.
A typical sequence of execution is as mentioned in hare setup.yaml.
/opt/seagate/cortx/hare/bin/hare_setup post_install --config 'json:///root/hare.post_install.conf.tmpl.1-node'
- Validates if mandatory pre requisites are met.
- Reports any unavailable features for a given setup. This requires cortx-prvsnr to be installed as this sub command uses provisioner cli to fetch the setup (virtual or physical) information.
- Installs Hare logrotate configuration files.
-
--configparameter is effectively ignored but left because of an explicit requirement of Mini Provisioner.
/opt/seagate/cortx/hare/bin/hare_setup prepare --config 'json:///root/hare.prepare.conf.tmpl.1-node'
[root@ssc-vm-1623:root] /opt/seagate/cortx/hare/bin/hare_setup upgrade --config 'json:///root/hare.rpm'
Prerequisites:
- cortx-py-utils package must be installed.
- Motr must be installed.
- S3server must be installed.
- Data interfaces must be registered with lnet. Refer this
- lnet kernel module inserted and lnet service must be running.
- Add/set 'srvnode-1.data.private' to entry containing 'private IP' in /etc/hosts file
- Motr data and metadata volumes must be configured. Refer this
- Follow the mini-provisioning post-install, prepare and config steps of cortx-motr. This generates the shared motr-hare confstore file at path
/opt/seagate/cortx/motr/conf/motr_hare_keys.json. - Copy the template file hare.config.conf.tmpl.1-node in /root/. and populate following keys appropriately for nodes in a storage set.
TMPL_CLUSTER_ID
TMPL_STORAGESET_COUNT
TMPL_DATA_UNITS_COUNT (4)
TMPL_PARITY_UNITS_COUNT (2)
TMPL_SPARE_UNITS_COUNT (2)
TMPL_STORAGESET_NAME
TMPL_MACHINE_ID (cat /etc/machine-id)
TMPL_HOSTNAME
TMPL_SERVER_NODE_NAME
TMPL_DATA_INTERFACE_TYPE (tcp)
TMPL_PRIVATE_FQDN
TMPL_PRIVATE_DATA_INTERFACE_1 (eth0)
TMPL_PRIVATE_DATA_INTERFACE_2 (optional)
TMPL_S3SERVER_INSTANCES_COUNT
Note: Total number of data devices in a storage set must be >= TMPL_DATA_UNITS_COUNT + TMPL_PARITY_UNITS_COUNT + TMPL_SPARE_UNITS_COUNT.
TMPL_DATA_DEVICE_1
TMPL_DATA_DEVICE_2
TMPL_METADATA_DEVICE
{
"cluster": {
"my-cluster": {
"site": {
"storage_set_count": "1"
},
"storage_set": [
{
"durability": {
"sns": {
"data": "1",
"parity": "0",
"spare": "0"
}
},
"name": "storage1",
"server_nodes": [
"7650c0e8d27ffb4f44b8ac3c26f29a45"
]
}
]
}
},
"server_node": {
"7650c0e8d27ffb4f44b8ac3c26f29a45": {
"cluster_id": "my-cluster",
"hostname": "ssc-vm-3353.colo.seagate.com",
"name": "srvnode-1",
"network": {
"data": {
"interface_type": "tcp",
"private_fqdn": "srvnode-1.data.private",
"private_interfaces": [
"eth0",
"eno2"
]
}
},
"s3_instances": "1",
"storage": {
"cvg": [
{
"data_devices": [
"/dev/loop0",
"/dev/loop1"
],
"metadata_devices": [
"/dev/loop2"
]
}
]
}
}
}
}
Generates hare and motr cluster description file using ConfStore.
-
--configis the URL for ConfStore file. - --file Where the Cluster Description File (CDF) must be written to as a result of this operation.
/opt/seagate/cortx/hare/bin/hare_setup config --config 'json:///root/hare.config.conf.tmpl.1-node' --file '/var/lib/hare/cluster.yaml'
ConfStore keys:
cluster>CLUSTER_ID>site>storage_set_count
cluster>CLUSTER_ID>storage_set[0]>durability>sns>data
cluster>CLUSTER_ID>storage_set[0]>durability>sns>parity
cluster>CLUSTER_ID>storage_set[0]>durability>sns>spare
cluster>CLUSTER_ID>storage_set[0]>name
cluster>CLUSTER_ID>storage_set[0]>server_nodes
server_node>MACH_ID>cluster_id
server_node>MACH_ID>storage>cvg[0]>data_devices
server_node>MACH_ID>storage>cvg[0]>metadata_devices
server_node>MACH_ID>hostname
server_node>MACH_ID>network>data>interface_type
server_node>MACH_ID>network>data>private_fdqn
server_node>MACH_ID>network>data>private_interfaces
server_node>MACH_ID>s3_instances
Sample ConfStore json:
Example of data read from Conf-store can be found here
Prerequisites:
- cortx-py-utils package must be installed.
- Consul must be installed.
- Motr must be installed.
- S3server must be installed.
- Data interfaces must be registered with lnet. Refer this
- lnet kernel module inserted and lnet service must be running.
- Motr data and metadata volumes must be configured. Refer this
- Copy the template file hare.init.conf.tmpl.1-node in /root/. and populate following keys appropriately for nodes in a storage set.
TMPL_CLUSTER_ID
TMPL_STORAGESET_COUNT
TMPL_DATA_UNITS_COUNT (4)
TMPL_PARITY_UNITS_COUNT (2)
TMPL_SPARE_UNITS_COUNT (2)
TMPL_STORAGESET_NAME
TMPL_MACHINE_ID (/etc/machine-id)
TMPL_HOSTNAME
TMPL_SERVER_NODE_NAME
TMPL_DATA_INTERFACE_TYPE (tcp)
TMPL_PRIVATE_DATA_INTERFACE_1 (eth0)
TMPL_PRIVATE_DATA_INTERFACE_2 (optional)
TMPL_S3SERVER_INSTANCES_COUNT
Note: Total number of data devices in a storage set must be >= TMPL_DATA_UNITS_COUNT + TMPL_PARITY_UNITS_COUNT + TMPL_SPARE_UNITS_COUNT.
TMPL_DATA_DEVICE_1
TMPL_DATA_DEVICE_2
TMPL_METADATA_DEVICE
Conf-store keys used for validation
server_node>{machine_id}>cluster_id
server_node>{machine_id}>hostname
cluster>{cluster_id}>storage_set[0]>server_nodes
[root@ssc-vm-1623:cortx-hare-1] /opt/seagate/cortx/hare/bin/hare_setup init --config 'json:///root/hare.init.conf.tmpl.1-node' --file '/var/lib/hare/cluster.yaml'
Cluster is not running
2021-02-02 13:14:34: Generating cluster configuration... OK
2021-02-02 13:14:35: Starting Consul server on this node........ OK
2021-02-02 13:14:40: Importing configuration into the KV store... OK
2021-02-02 13:14:40: Starting Consul on other nodes... OK
2021-02-02 13:14:41: Updating Consul configuraton from the KV store... OK
2021-02-02 13:14:43: Waiting for the RC Leader to get elected...... OK
2021-02-02 13:14:46: Starting Motr (phase1, mkfs)... OK
2021-02-02 13:14:53: Starting Motr (phase1, m0d)... OK
2021-02-02 13:15:05: Starting Motr (phase2, mkfs)... OK
2021-02-02 13:15:12: Starting Motr (phase2, m0d)... OK
2021-02-02 13:15:25: Starting S3 servers (phase3)... OK
2021-02-02 13:15:28: Checking health of services... OK
Stopping s3server@0x7200000000000001:0x1d at ssc-vm-1623.colo.seagate.com...
Stopping s3server@0x7200000000000001:0x20 at ssc-vm-1623.colo.seagate.com...
Stopping s3server@0x7200000000000001:0x23 at ssc-vm-1623.colo.seagate.com...
Stopping s3server@0x7200000000000001:0x26 at ssc-vm-1623.colo.seagate.com...
Stopping s3server@0x7200000000000001:0x29 at ssc-vm-1623.colo.seagate.com...
Stopping s3server@0x7200000000000001:0x2c at ssc-vm-1623.colo.seagate.com...
Stopping s3server@0x7200000000000001:0x2f at ssc-vm-1623.colo.seagate.com...
Stopping s3server@0x7200000000000001:0x32 at ssc-vm-1623.colo.seagate.com...
Stopped s3server@0x7200000000000001:0x1d at ssc-vm-1623.colo.seagate.com
Stopping s3server@0x7200000000000001:0x35 at ssc-vm-1623.colo.seagate.com...
Stopped s3server@0x7200000000000001:0x20 at ssc-vm-1623.colo.seagate.com
Stopping s3server@0x7200000000000001:0x38 at ssc-vm-1623.colo.seagate.com...
Stopped s3server@0x7200000000000001:0x2f at ssc-vm-1623.colo.seagate.com
Stopping s3server@0x7200000000000001:0x3b at ssc-vm-1623.colo.seagate.com...
Stopped s3server@0x7200000000000001:0x38 at ssc-vm-1623.colo.seagate.com
Stopped s3server@0x7200000000000001:0x2c at ssc-vm-1623.colo.seagate.com
Stopped s3server@0x7200000000000001:0x35 at ssc-vm-1623.colo.seagate.com
Stopped s3server@0x7200000000000001:0x32 at ssc-vm-1623.colo.seagate.com
Stopped s3server@0x7200000000000001:0x3b at ssc-vm-1623.colo.seagate.com
Stopped s3server@0x7200000000000001:0x26 at ssc-vm-1623.colo.seagate.com
Stopped s3server@0x7200000000000001:0x23 at ssc-vm-1623.colo.seagate.com
Stopped s3server@0x7200000000000001:0x29 at ssc-vm-1623.colo.seagate.com
Stopping m0d@0x7200000000000001:0xc (ios) at ssc-vm-1623.colo.seagate.com...
Stopped m0d@0x7200000000000001:0xc (ios) at ssc-vm-1623.colo.seagate.com
Stopping m0d@0x7200000000000001:0x9 (confd) at ssc-vm-1623.colo.seagate.com...
Stopped m0d@0x7200000000000001:0x9 (confd) at ssc-vm-1623.colo.seagate.com
Stopping hare-hax at ssc-vm-1623.colo.seagate.com...
Stopped hare-hax at ssc-vm-1623.colo.seagate.com
Stopping hare-consul-agent at ssc-vm-1623.colo.seagate.com...
Stopped hare-consul-agent at ssc-vm-1623.colo.seagate.com
Shutting down RC Leader at ssc-vm-1623.colo.seagate.com...
**ERROR**
Cluster is not running
- Checks if cluster is already running, if yes, shuts down the running cluster.
- Executes hctl bootstrap to,
- Generate hare configuration files and initialize hare Consul KV.
- Configure and initialize cortx-motr.
- Shuts down the hare and motr cluster once initialization is complete.
Note: Only one node in a cluster can execute the init section at any given instance because hctl bootstrap is a cluster wide operation.
Prerequisites:
- cortx-py-utils package must be installed.
- Consul must be installed.
- Motr must be installed.
- S3server must be installed.
- Data interfaces must be registered with lnet. Refer this
- lnet kernel module inserted and lnet service must be running.
- Motr data and metadata volumes must be configured. Refer this
- Copy the template file hare.test.conf.tmpl.1-nodein /root/. and populate following keys appropriately for nodes in a storage set.
TMPL_CLUSTER_ID
TMPL_STORAGESET_COUNT
TMPL_DATA_UNITS_COUNT (4)
TMPL_PARITY_UNITS_COUNT (2)
TMPL_SPARE_UNITS_COUNT (2)
TMPL_STORAGESET_NAME
TMPL_MACHINE_ID (/etc/machine-id)
TMPL_HOSTNAME
TMPL_SERVER_NODE_NAME
TMPL_DATA_INTERFACE_TYPE (tcp)
TMPL_PRIVATE_DATA_INTERFACE_1 (eth0)
TMPL_PRIVATE_DATA_INTERFACE_2 (optional)
TMPL_S3SERVER_INSTANCES_COUNT
Note: Total number of data devices in a storage set must be >= TMPL_DATA_UNITS_COUNT + TMPL_PARITY_UNITS_COUNT + TMPL_SPARE_UNITS_COUNT.
TMPL_DATA_DEVICE_1
TMPL_DATA_DEVICE_2
TMPL_METADATA_DEVICE
Conf-store keys used
server_node>{machine_id}>cluster_id
server_node>{machine_id}>hostname
cluster>{cluster_id}>storage_set[0]>server_nodes
/opt/seagate/cortx/hare/bin/hare_setup test --config 'json:///root/hare.test.conf.tmpl.1-node'
Runs hare sanity tests to validate the cluster configuration.
To be implemented
To be implemented
/opt/seagate/cortx/hare/bin/hare_setup reset --config 'json:///root/hare.reset.conf.tmpl.1-node'
- Resets hare and motr configuration.
- Deletes existing hare log files.
To be implemented
/opt/seagate/cortx/hare/bin/hare_setup cleanup --config 'json:///root/hare.cleanup.conf.tmpl.1-node'
- Deletes test data and unwanted log files.
To be implemented
To be implemented
Without bundle-id and destination directory path options, hare_setup support_bundle command generates Hare support bundle at /tmp/hare.
[root@ssc-vm-1623:root] /opt/seagate/cortx/hare/bin/hare_setup support_bundle
[root@ssc-vm-1623:root] ls /tmp/hare
hare_ssc-vm-1623.tar.gz
With bundle-id and destination directory arguments to hare_setup support_bundle.
[root@ssc-vm-1623:root] /opt/seagate/cortx/hare/bin/hare_setup support_bundle SB12345 /root
[root@ssc-vm-1623:root] ls hare
hare_SB12345.tar.gz
Runs hctl reportbug command to generate hare support bundle.
- Any hare_setup command generates
No such file or directory: 'provisioner'error
[root@ssc-vm-1623:root] /opt/seagate/cortx/hare/bin/hare_setup post_install --config 'json:///tmp/exampleV2.json'
Failed to fetch data from provisioner ([Errno 2] No such file or directory: 'provisioner': 'provisioner')
Presently by default hare_setup tries to configure hare logrotate which invokes provisioner api. Thus if provisioner cli is not installed, hare_setup throws this error. But this does not mean that the corresponding hare_setup command failed. So either cortx-prvsnr cli needs to be installed or the error can be ignored by the user.
This problem is being worked upon in https://github.com/Seagate/cortx-hare/pull/1489, so further builds may not have this issue.
- Invalid device paths
[root@ssc-vm-1623:cortx-hare-1] hctl bootstrap --mkfs /var/lib/hare/cluster.yaml
2021-02-03 11:36:10: Generating cluster configuration...Traceback (most recent call last):
File "<string>", line 13, in <module>
File "<string>", line 13, in <listcomp>
File "<string>", line 7, in blockdev_size
FileNotFoundError: [Errno 2] No such file or directory: '/dev/sdx'
Traceback (most recent call last):
File "/opt/seagate/cortx/hare/bin/../bin/cfgen", line 1757, in <module>
main()
File "/opt/seagate/cortx/hare/bin/../bin/cfgen", line 132, in main
enrich_cluster_desc(cluster_desc, opts.mock)
File "/opt/seagate/cortx/hare/bin/../bin/cfgen", line 226, in enrich_cluster_desc
m0d['io_disks']['data'])
File "/opt/seagate/cortx/hare/bin/../bin/cfgen", line 463, in get_disks
'sudo', 'python3', '-c', code))]
File "/opt/seagate/cortx/hare/bin/../bin/cfgen", line 408, in run_command
timeout=15).decode()
File "/usr/lib64/python3.6/subprocess.py", line 356, in check_output
**kwargs).stdout
File "/usr/lib64/python3.6/subprocess.py", line 438, in run
output=stdout, stderr=stderr)
subprocess.CalledProcessError: Command '('sudo', 'python3', '-c', "import io\nimport os\n\n# os.path.getsize() and os.stat().st_size don't work well with loop devices,\n# they always return 0.\ndef blockdev_size(path):\n with open(path, 'rb') as f:\n return f.seek(0, io.SEEK_END)\n\nprint([dict(path=path,\n size=blockdev_size(path),\n blksize=os.stat(path).st_blksize)\n for path in ['/dev/sdx', '/dev/sdy', '/dev/sdz']])\n")' returned non-zero exit status 1.
This can happen if the data or metadata device paths are wrong in conf-store. Hare fetches device information from conf-store, which is validated during configuration generation during hctl bootstrap. Thus if the device path does not exists hare init may fail. Kindly make sure that the device paths in conf-store are correct.
- Wrong network interfaces
[root@ssc-vm-1623:cortx-hare-1] hctl bootstrap --mkfs /var/lib/hare/cluster.yaml
2021-02-03 11:58:13: Generating cluster configuration...Traceback (most recent call last):
File "/opt/seagate/cortx/hare/bin/../bin/cfgen", line 1757, in <module>
main()
File "/opt/seagate/cortx/hare/bin/../bin/cfgen", line 133, in main
validate_cluster_desc(cluster_desc)
File "/opt/seagate/cortx/hare/bin/../bin/cfgen", line 235, in validate_cluster_desc
validate_nodes_desc(desc['nodes'])
File "/opt/seagate/cortx/hare/bin/../bin/cfgen", line 262, in validate_nodes_desc
Make sure the value of data_iface in the CDF is correct."""
AssertionError: ssc-vm-1623.colo.seagate.com: 'eth3' interface has no IP address
Make sure the value of data_iface in the CDF is correct.
Make sure the public data network interfaces mentioned in conf-store are correct. Hare validates the same while generating configuration and generate the above error in case of invalid network interfaces.
- Wrong motr configuration
Feb 03 07:05:32 ssc-vm-2302.colo.seagate.com systemd[1]: Starting Motr mkfs helper for 0x7200000000000001:0xc service...
Feb 03 07:05:32 ssc-vm-2302.colo.seagate.com motr-mkfs[23588]: ios seg size is set to LVM size:$MOTR_M0D_IOS_BESEG_SIZE
Feb 03 07:05:32 ssc-vm-2302.colo.seagate.com motr-mkfs[23588]: /sys/kernel/mm/transparent_hugepage/defrag: always madvise [never]
Feb 03 07:05:32 ssc-vm-2302.colo.seagate.com motr-mkfs[23588]: MOTR_M0D_EP: 192.168.83.155@tcp:12345:2:2
Feb 03 07:05:32 ssc-vm-2302.colo.seagate.com motr-mkfs[23588]: MOTR_PROCESS_FID: 0x7200000000000001:0xc
Feb 03 07:05:32 ssc-vm-2302.colo.seagate.com motr-mkfs[23588]: MOTR_HA_EP: 192.168.83.155@tcp:12345:1:1
Feb 03 07:05:32 ssc-vm-2302.colo.seagate.com motr-mkfs[23588]: vm.max_map_count = 30000000
Feb 03 07:05:32 ssc-vm-2302.colo.seagate.com motr-mkfs[23588]: + exec /usr/sbin/m0mkfs -e lnet:192.168.83.155@tcp:12345:2:2 -A linuxstob:/var/log/seagate/motr/addb/m0d-0x7200000000000001:0xc/
Feb 03 07:05:39 ssc-vm-2302.colo.seagate.com motr-mkfs[23588]: motr[23590]: c800 WARN [stob/ioq.c:443:ioq_io_error] IO error: stob_id={<100000000000000:bef11f>,<100000000000000:2a>} conf
Feb 03 07:05:39 ssc-vm-2302.colo.seagate.com motr-mkfs[23588]: motr[23590]: c5d0 FATAL [lib/assert.c:50:m0_panic] panic: sio->si_rc == 0 at be_io_cb() (be/io.c:575) [git: sage-base-1.0-1
Feb 03 07:05:39 ssc-vm-2302.colo.seagate.com motr-mkfs[23588]: Motr panic: sio->si_rc == 0 at be_io_cb() be/io.c:575 (errno: 0) (last failed: none) [git: sage-base-1.0-190-ga936745] pid: 2359
Feb 03 07:05:39 ssc-vm-2302.colo.seagate.com motr-mkfs[23588]: Motr panic reason: stob I/O operation failed: bio = 0x7ffca55b2d30, sio = 0x2a83e50, sio->si_rc = -28
....
Feb 03 07:05:42 ssc-vm-2302.colo.seagate.com motr-mkfs[23588]: /usr/libexec/cortx-motr/motr-mkfs: line 40: 23590 Aborted (core dumped) $motr_exec_dir/motr-server $service m0mk
Feb 03 07:05:42 ssc-vm-2302.colo.seagate.com motr-mkfs[23588]: propagating error to parent shell
Feb 03 07:05:42 ssc-vm-2302.colo.seagate.com motr-mkfs[23588]: got sub-shell error, terminating..
Feb 03 07:05:42 ssc-vm-2302.colo.seagate.com systemd[1]: motr-mkfs@0x7200000000000001:0xc.service: main process exited, code=exited, status=42/n/a
Feb 03 07:05:42 ssc-vm-2302.colo.seagate.com systemd[1]: Failed to start Motr mkfs helper for 0x7200000000000001:0xc service.
Feb 03 07:05:42 ssc-vm-2302.colo.seagate.com systemd[1]: Unit motr-mkfs@0x7200000000000001:0xc.service entered failed state.
Feb 03 07:05:42 ssc-vm-2302.colo.seagate.com systemd[1]: motr-mkfs@0x7200000000000001:0xc.service failed.
Above issue can be seen during hare_setup init if /etc/sysconfig/motr configuration does not match the actual setup. In above case MOTR_M0D_IOS_BESEG_SIZE=5497558138880 value is too large for a VM setup and for a given meta data device.