Skip to content
This repository was archived by the owner on Feb 8, 2024. It is now read-only.

Hare provisioning for single node deployment

pavankrishnat edited this page Sep 14, 2021 · 18 revisions

Pre requisites

Note: In case of Mini-Provisioning, consul will not be started/stopped by Hare so it is pre-requisite that consul should be running on all the nodes present

Setup utility

/opt/seagate/cortx/hare/bin/hare_setup --help
# /opt/seagate/cortx/hare/bin/hare_setup --help
usage: hare_setup [-h]
                  {post_install,config,init,test,support_bundle,reset,cleanup,prepare,pre-upgrade,post-upgrade}
                  ...

Configure hare settings

positional arguments:
  {post_install,config,init,test,support_bundle,reset,cleanup,prepare,pre-upgrade,post-upgrade}
    post_install        Validates installation
    config              Configures Hare
    init                Initializes Hare
    test                Tests Hare component
    support_bundle      Generates support bundle
    reset               Resets temporary Hare data and configuration
    cleanup             Resets Hare configuration, logs & formats Motr
                        metadata
    prepare             Validates configuration pre-requisites
    pre-upgrade         Performs the Hare rpm pre-upgrade tasks
    post-upgrade        Performs the Hare rpm post-upgrade tasks

optional arguments:
  -h, --help            show this help message and exit

Hare setup utility (hare_setup) comprises of sub commands (as below) that help to provision cortx-hare and configure cortx-motr such that the corresponding hare and motr services are ready to start. This utility can be used to provision cortx-hare in virtual as well as physical cluster environments. Few points to note,

  • hare_setup utility does not install the above mention pre requisites but assumes they are already installed and performs required validation checks.
  • Although every sub command is idempotent, it has pre requisites of its own which are implemented in the form of validation checks.
  • Due to the nature of operations performed, some of the sub commands have a limitation that they cannot be executed in parallel on multiple nodes in a cluster.
  • Uses conf-store module provided by cortx-py-utils.
  • Hare single node provisioning templates are available that can be used to generate required conf-store files.
  • User needs to make sure that total number of data devices in a pool(storage set) are >= cluster>CLUSTER_ID>storage_set[0]>durability>sns>data + cluster>CLUSTER_ID>storage_set[0]>durability>sns>parity + cluster>CLUSTER_ID>storage_set[0]>durability>sns>spare.

Provisioning

A typical sequence of execution is as mentioned in hare setup.yaml.

Post-install

/opt/seagate/cortx/hare/bin/hare_setup post_install --config 'json:///root/hare.post_install.conf.tmpl.1-node'
  • Validates if mandatory pre requisites are met.
  • Reports any unavailable features for a given setup. This requires cortx-prvsnr to be installed as this sub command uses provisioner cli to fetch the setup (virtual or physical) information.
  • Installs Hare logrotate configuration files.
  • --config parameter is effectively ignored but left because of an explicit requirement of Mini Provisioner.

Prepare

/opt/seagate/cortx/hare/bin/hare_setup prepare --config 'json:///root/hare.prepare.conf.tmpl.1-node'

Upgrade

[root@ssc-vm-1623:root] /opt/seagate/cortx/hare/bin/hare_setup upgrade --config 'json:///root/hare.rpm'

Config

Prerequisites:

  • cortx-py-utils package must be installed.
  • Motr must be installed.
  • S3server must be installed.
  • Data interfaces must be registered with lnet. Refer this
  • lnet kernel module inserted and lnet service must be running.
  • Add/set 'srvnode-1.data.private' to entry containing 'private IP' in /etc/hosts file
  • Motr data and metadata volumes must be configured. Refer this
  • Follow the mini-provisioning post-install, prepare and config steps of cortx-motr. This generates the shared motr-hare confstore file at path /opt/seagate/cortx/motr/conf/motr_hare_keys.json.
  • Copy the template file hare.config.conf.tmpl.1-node in /root/. and populate following keys appropriately for nodes in a storage set.
TMPL_CLUSTER_ID
TMPL_STORAGESET_COUNT
TMPL_DATA_UNITS_COUNT (4)
TMPL_PARITY_UNITS_COUNT (2)
TMPL_SPARE_UNITS_COUNT (2)
TMPL_STORAGESET_NAME
TMPL_MACHINE_ID (cat /etc/machine-id)
TMPL_HOSTNAME
TMPL_SERVER_NODE_NAME
TMPL_DATA_INTERFACE_TYPE (tcp)
TMPL_PRIVATE_FQDN
TMPL_PRIVATE_DATA_INTERFACE_1 (eth0)
TMPL_PRIVATE_DATA_INTERFACE_2 (optional)
TMPL_S3SERVER_INSTANCES_COUNT

Note: Total number of data devices in a storage set must be >= TMPL_DATA_UNITS_COUNT + TMPL_PARITY_UNITS_COUNT + TMPL_SPARE_UNITS_COUNT.
TMPL_DATA_DEVICE_1
TMPL_DATA_DEVICE_2
TMPL_METADATA_DEVICE

Example file

{
  "cluster": {
    "my-cluster": {
      "site": {
        "storage_set_count": "1"
      },
      "storage_set": [
        {
          "durability": {
            "sns": {
              "data": "1",
              "parity": "0",
              "spare": "0"
            }
          },
          "name": "storage1",
          "server_nodes": [
            "7650c0e8d27ffb4f44b8ac3c26f29a45"
          ]
        }
      ]
    }
  },
  "server_node": {
    "7650c0e8d27ffb4f44b8ac3c26f29a45": {
      "cluster_id": "my-cluster",
      "hostname": "ssc-vm-3353.colo.seagate.com",
      "name": "srvnode-1",
      "network": {
        "data": {
          "interface_type": "tcp",
          "private_fqdn": "srvnode-1.data.private",
          "private_interfaces": [
            "eth0",
            "eno2"
          ]
        }
      },
      "s3_instances": "1",
      "storage": {
        "cvg": [
          {
            "data_devices": [
              "/dev/loop0",
              "/dev/loop1"
            ],
            "metadata_devices": [
              "/dev/loop2"
            ]
          }
        ]
      }
    }
  }
}

Generates hare and motr cluster description file using ConfStore.

  • --config is the URL for ConfStore file.
  • --file Where the Cluster Description File (CDF) must be written to as a result of this operation.
/opt/seagate/cortx/hare/bin/hare_setup config --config 'json:///root/hare.config.conf.tmpl.1-node' --file '/var/lib/hare/cluster.yaml'

ConfStore keys:

cluster>CLUSTER_ID>site>storage_set_count
cluster>CLUSTER_ID>storage_set[0]>durability>sns>data
cluster>CLUSTER_ID>storage_set[0]>durability>sns>parity
cluster>CLUSTER_ID>storage_set[0]>durability>sns>spare
cluster>CLUSTER_ID>storage_set[0]>name
cluster>CLUSTER_ID>storage_set[0]>server_nodes
server_node>MACH_ID>cluster_id
server_node>MACH_ID>storage>cvg[0]>data_devices
server_node>MACH_ID>storage>cvg[0]>metadata_devices
server_node>MACH_ID>hostname
server_node>MACH_ID>network>data>interface_type
server_node>MACH_ID>network>data>private_fdqn
server_node>MACH_ID>network>data>private_interfaces
server_node>MACH_ID>s3_instances

Sample ConfStore json:

Example of data read from Conf-store can be found here

Init

Prerequisites:

  • cortx-py-utils package must be installed.
  • Consul must be installed.
  • Motr must be installed.
  • S3server must be installed.
  • Data interfaces must be registered with lnet. Refer this
  • lnet kernel module inserted and lnet service must be running.
  • Motr data and metadata volumes must be configured. Refer this
  • Copy the template file hare.init.conf.tmpl.1-node in /root/. and populate following keys appropriately for nodes in a storage set.
TMPL_CLUSTER_ID
TMPL_STORAGESET_COUNT
TMPL_DATA_UNITS_COUNT (4)
TMPL_PARITY_UNITS_COUNT (2)
TMPL_SPARE_UNITS_COUNT (2)
TMPL_STORAGESET_NAME
TMPL_MACHINE_ID (/etc/machine-id)
TMPL_HOSTNAME
TMPL_SERVER_NODE_NAME
TMPL_DATA_INTERFACE_TYPE (tcp)
TMPL_PRIVATE_DATA_INTERFACE_1 (eth0)
TMPL_PRIVATE_DATA_INTERFACE_2 (optional)
TMPL_S3SERVER_INSTANCES_COUNT

Note: Total number of data devices in a storage set must be >= TMPL_DATA_UNITS_COUNT + TMPL_PARITY_UNITS_COUNT + TMPL_SPARE_UNITS_COUNT.
TMPL_DATA_DEVICE_1
TMPL_DATA_DEVICE_2
TMPL_METADATA_DEVICE

Conf-store keys used for validation

server_node>{machine_id}>cluster_id
server_node>{machine_id}>hostname
cluster>{cluster_id}>storage_set[0]>server_nodes
[root@ssc-vm-1623:cortx-hare-1] /opt/seagate/cortx/hare/bin/hare_setup init --config 'json:///root/hare.init.conf.tmpl.1-node' --file '/var/lib/hare/cluster.yaml'
Cluster is not running
2021-02-02 13:14:34: Generating cluster configuration... OK
2021-02-02 13:14:35: Starting Consul server on this node........ OK
2021-02-02 13:14:40: Importing configuration into the KV store... OK
2021-02-02 13:14:40: Starting Consul on other nodes... OK
2021-02-02 13:14:41: Updating Consul configuraton from the KV store... OK
2021-02-02 13:14:43: Waiting for the RC Leader to get elected...... OK
2021-02-02 13:14:46: Starting Motr (phase1, mkfs)... OK
2021-02-02 13:14:53: Starting Motr (phase1, m0d)... OK
2021-02-02 13:15:05: Starting Motr (phase2, mkfs)... OK
2021-02-02 13:15:12: Starting Motr (phase2, m0d)... OK
2021-02-02 13:15:25: Starting S3 servers (phase3)... OK
2021-02-02 13:15:28: Checking health of services... OK
Stopping s3server@0x7200000000000001:0x1d at ssc-vm-1623.colo.seagate.com...
Stopping s3server@0x7200000000000001:0x20 at ssc-vm-1623.colo.seagate.com...
Stopping s3server@0x7200000000000001:0x23 at ssc-vm-1623.colo.seagate.com...
Stopping s3server@0x7200000000000001:0x26 at ssc-vm-1623.colo.seagate.com...
Stopping s3server@0x7200000000000001:0x29 at ssc-vm-1623.colo.seagate.com...
Stopping s3server@0x7200000000000001:0x2c at ssc-vm-1623.colo.seagate.com...
Stopping s3server@0x7200000000000001:0x2f at ssc-vm-1623.colo.seagate.com...
Stopping s3server@0x7200000000000001:0x32 at ssc-vm-1623.colo.seagate.com...
Stopped s3server@0x7200000000000001:0x1d at ssc-vm-1623.colo.seagate.com
Stopping s3server@0x7200000000000001:0x35 at ssc-vm-1623.colo.seagate.com...
Stopped s3server@0x7200000000000001:0x20 at ssc-vm-1623.colo.seagate.com
Stopping s3server@0x7200000000000001:0x38 at ssc-vm-1623.colo.seagate.com...
Stopped s3server@0x7200000000000001:0x2f at ssc-vm-1623.colo.seagate.com
Stopping s3server@0x7200000000000001:0x3b at ssc-vm-1623.colo.seagate.com...
Stopped s3server@0x7200000000000001:0x38 at ssc-vm-1623.colo.seagate.com
Stopped s3server@0x7200000000000001:0x2c at ssc-vm-1623.colo.seagate.com
Stopped s3server@0x7200000000000001:0x35 at ssc-vm-1623.colo.seagate.com
Stopped s3server@0x7200000000000001:0x32 at ssc-vm-1623.colo.seagate.com
Stopped s3server@0x7200000000000001:0x3b at ssc-vm-1623.colo.seagate.com
Stopped s3server@0x7200000000000001:0x26 at ssc-vm-1623.colo.seagate.com
Stopped s3server@0x7200000000000001:0x23 at ssc-vm-1623.colo.seagate.com
Stopped s3server@0x7200000000000001:0x29 at ssc-vm-1623.colo.seagate.com
Stopping m0d@0x7200000000000001:0xc (ios) at ssc-vm-1623.colo.seagate.com...
Stopped m0d@0x7200000000000001:0xc (ios) at ssc-vm-1623.colo.seagate.com
Stopping m0d@0x7200000000000001:0x9 (confd) at ssc-vm-1623.colo.seagate.com...
Stopped m0d@0x7200000000000001:0x9 (confd) at ssc-vm-1623.colo.seagate.com
Stopping hare-hax at ssc-vm-1623.colo.seagate.com...
Stopped hare-hax at ssc-vm-1623.colo.seagate.com
Stopping hare-consul-agent at ssc-vm-1623.colo.seagate.com...
Stopped hare-consul-agent at ssc-vm-1623.colo.seagate.com
Shutting down RC Leader at ssc-vm-1623.colo.seagate.com...
**ERROR**
Cluster is not running
  • Checks if cluster is already running, if yes, shuts down the running cluster.
  • Executes hctl bootstrap to,
  • Shuts down the hare and motr cluster once initialization is complete.

Note: Only one node in a cluster can execute the init section at any given instance because hctl bootstrap is a cluster wide operation.

Test

Prerequisites:

  • cortx-py-utils package must be installed.
  • Consul must be installed.
  • Motr must be installed.
  • S3server must be installed.
  • Data interfaces must be registered with lnet. Refer this
  • lnet kernel module inserted and lnet service must be running.
  • Motr data and metadata volumes must be configured. Refer this
  • Copy the template file hare.test.conf.tmpl.1-nodein /root/. and populate following keys appropriately for nodes in a storage set.
TMPL_CLUSTER_ID
TMPL_STORAGESET_COUNT
TMPL_DATA_UNITS_COUNT (4)
TMPL_PARITY_UNITS_COUNT (2)
TMPL_SPARE_UNITS_COUNT (2)
TMPL_STORAGESET_NAME
TMPL_MACHINE_ID (/etc/machine-id)
TMPL_HOSTNAME
TMPL_SERVER_NODE_NAME
TMPL_DATA_INTERFACE_TYPE (tcp)
TMPL_PRIVATE_DATA_INTERFACE_1 (eth0)
TMPL_PRIVATE_DATA_INTERFACE_2 (optional)
TMPL_S3SERVER_INSTANCES_COUNT

Note: Total number of data devices in a storage set must be >= TMPL_DATA_UNITS_COUNT + TMPL_PARITY_UNITS_COUNT + TMPL_SPARE_UNITS_COUNT.
TMPL_DATA_DEVICE_1
TMPL_DATA_DEVICE_2
TMPL_METADATA_DEVICE

Conf-store keys used

server_node>{machine_id}>cluster_id
server_node>{machine_id}>hostname
cluster>{cluster_id}>storage_set[0]>server_nodes
/opt/seagate/cortx/hare/bin/hare_setup test --config 'json:///root/hare.test.conf.tmpl.1-node'

Runs hare sanity tests to validate the cluster configuration.

Upgrade

To be implemented

Reset

To be implemented

/opt/seagate/cortx/hare/bin/hare_setup reset --config 'json:///root/hare.reset.conf.tmpl.1-node'
  • Resets hare and motr configuration.
  • Deletes existing hare log files.

Cleanup

To be implemented

/opt/seagate/cortx/hare/bin/hare_setup cleanup --config 'json:///root/hare.cleanup.conf.tmpl.1-node'
  • Deletes test data and unwanted log files.

Backup

To be implemented

Restore

To be implemented

Support Bundle

Without bundle-id and destination directory path options, hare_setup support_bundle command generates Hare support bundle at /tmp/hare.

[root@ssc-vm-1623:root] /opt/seagate/cortx/hare/bin/hare_setup support_bundle
[root@ssc-vm-1623:root] ls /tmp/hare
hare_ssc-vm-1623.tar.gz

With bundle-id and destination directory arguments to hare_setup support_bundle.

[root@ssc-vm-1623:root] /opt/seagate/cortx/hare/bin/hare_setup support_bundle SB12345 /root
[root@ssc-vm-1623:root] ls hare
hare_SB12345.tar.gz

Runs hctl reportbug command to generate hare support bundle.

Trouble shooting

  1. Any hare_setup command generates No such file or directory: 'provisioner' error
[root@ssc-vm-1623:root] /opt/seagate/cortx/hare/bin/hare_setup post_install --config 'json:///tmp/exampleV2.json'
Failed to fetch data from provisioner ([Errno 2] No such file or directory: 'provisioner': 'provisioner')

Presently by default hare_setup tries to configure hare logrotate which invokes provisioner api. Thus if provisioner cli is not installed, hare_setup throws this error. But this does not mean that the corresponding hare_setup command failed. So either cortx-prvsnr cli needs to be installed or the error can be ignored by the user. This problem is being worked upon in https://github.com/Seagate/cortx-hare/pull/1489, so further builds may not have this issue.

  1. Invalid device paths
[root@ssc-vm-1623:cortx-hare-1] hctl bootstrap --mkfs /var/lib/hare/cluster.yaml
2021-02-03 11:36:10: Generating cluster configuration...Traceback (most recent call last):
  File "<string>", line 13, in <module>
  File "<string>", line 13, in <listcomp>
  File "<string>", line 7, in blockdev_size
FileNotFoundError: [Errno 2] No such file or directory: '/dev/sdx'
Traceback (most recent call last):
  File "/opt/seagate/cortx/hare/bin/../bin/cfgen", line 1757, in <module>
    main()
  File "/opt/seagate/cortx/hare/bin/../bin/cfgen", line 132, in main
    enrich_cluster_desc(cluster_desc, opts.mock)
  File "/opt/seagate/cortx/hare/bin/../bin/cfgen", line 226, in enrich_cluster_desc
    m0d['io_disks']['data'])
  File "/opt/seagate/cortx/hare/bin/../bin/cfgen", line 463, in get_disks
    'sudo', 'python3', '-c', code))]
  File "/opt/seagate/cortx/hare/bin/../bin/cfgen", line 408, in run_command
    timeout=15).decode()
  File "/usr/lib64/python3.6/subprocess.py", line 356, in check_output
    **kwargs).stdout
  File "/usr/lib64/python3.6/subprocess.py", line 438, in run
    output=stdout, stderr=stderr)
subprocess.CalledProcessError: Command '('sudo', 'python3', '-c', "import io\nimport os\n\n# os.path.getsize() and os.stat().st_size don't work well with loop devices,\n# they always return 0.\ndef blockdev_size(path):\n    with open(path, 'rb') as f:\n        return f.seek(0, io.SEEK_END)\n\nprint([dict(path=path,\n            size=blockdev_size(path),\n            blksize=os.stat(path).st_blksize)\n       for path in ['/dev/sdx', '/dev/sdy', '/dev/sdz']])\n")' returned non-zero exit status 1.

This can happen if the data or metadata device paths are wrong in conf-store. Hare fetches device information from conf-store, which is validated during configuration generation during hctl bootstrap. Thus if the device path does not exists hare init may fail. Kindly make sure that the device paths in conf-store are correct.

  1. Wrong network interfaces
[root@ssc-vm-1623:cortx-hare-1] hctl bootstrap --mkfs /var/lib/hare/cluster.yaml
2021-02-03 11:58:13: Generating cluster configuration...Traceback (most recent call last):
  File "/opt/seagate/cortx/hare/bin/../bin/cfgen", line 1757, in <module>
    main()
  File "/opt/seagate/cortx/hare/bin/../bin/cfgen", line 133, in main
    validate_cluster_desc(cluster_desc)
  File "/opt/seagate/cortx/hare/bin/../bin/cfgen", line 235, in validate_cluster_desc
    validate_nodes_desc(desc['nodes'])
  File "/opt/seagate/cortx/hare/bin/../bin/cfgen", line 262, in validate_nodes_desc
    Make sure the value of data_iface in the CDF is correct."""
AssertionError: ssc-vm-1623.colo.seagate.com: 'eth3' interface has no IP address
Make sure the value of data_iface in the CDF is correct.

Make sure the public data network interfaces mentioned in conf-store are correct. Hare validates the same while generating configuration and generate the above error in case of invalid network interfaces.

  1. Wrong motr configuration
Feb 03 07:05:32 ssc-vm-2302.colo.seagate.com systemd[1]: Starting Motr mkfs helper for 0x7200000000000001:0xc service...                                                                       
Feb 03 07:05:32 ssc-vm-2302.colo.seagate.com motr-mkfs[23588]: ios seg size is set to LVM size:$MOTR_M0D_IOS_BESEG_SIZE                                                                        
Feb 03 07:05:32 ssc-vm-2302.colo.seagate.com motr-mkfs[23588]: /sys/kernel/mm/transparent_hugepage/defrag: always madvise [never]                                                              
Feb 03 07:05:32 ssc-vm-2302.colo.seagate.com motr-mkfs[23588]: MOTR_M0D_EP: 192.168.83.155@tcp:12345:2:2                                                                                       
Feb 03 07:05:32 ssc-vm-2302.colo.seagate.com motr-mkfs[23588]: MOTR_PROCESS_FID: 0x7200000000000001:0xc                                                                                        
Feb 03 07:05:32 ssc-vm-2302.colo.seagate.com motr-mkfs[23588]: MOTR_HA_EP: 192.168.83.155@tcp:12345:1:1                                                                                        
Feb 03 07:05:32 ssc-vm-2302.colo.seagate.com motr-mkfs[23588]: vm.max_map_count = 30000000                                                                                                     
Feb 03 07:05:32 ssc-vm-2302.colo.seagate.com motr-mkfs[23588]: + exec /usr/sbin/m0mkfs -e lnet:192.168.83.155@tcp:12345:2:2 -A linuxstob:/var/log/seagate/motr/addb/m0d-0x7200000000000001:0xc/
Feb 03 07:05:39 ssc-vm-2302.colo.seagate.com motr-mkfs[23588]: motr[23590]:  c800   WARN  [stob/ioq.c:443:ioq_io_error]  IO error: stob_id={<100000000000000:bef11f>,<100000000000000:2a>} conf
Feb 03 07:05:39 ssc-vm-2302.colo.seagate.com motr-mkfs[23588]: motr[23590]:  c5d0  FATAL  [lib/assert.c:50:m0_panic]  panic: sio->si_rc == 0 at be_io_cb() (be/io.c:575)  [git: sage-base-1.0-1
Feb 03 07:05:39 ssc-vm-2302.colo.seagate.com motr-mkfs[23588]: Motr panic: sio->si_rc == 0 at be_io_cb() be/io.c:575 (errno: 0) (last failed: none) [git: sage-base-1.0-190-ga936745] pid: 2359
Feb 03 07:05:39 ssc-vm-2302.colo.seagate.com motr-mkfs[23588]: Motr panic reason: stob I/O operation failed: bio = 0x7ffca55b2d30, sio = 0x2a83e50, sio->si_rc = -28                           
....                                  
Feb 03 07:05:42 ssc-vm-2302.colo.seagate.com motr-mkfs[23588]: /usr/libexec/cortx-motr/motr-mkfs: line 40: 23590 Aborted                 (core dumped) $motr_exec_dir/motr-server $service m0mk
Feb 03 07:05:42 ssc-vm-2302.colo.seagate.com motr-mkfs[23588]: propagating error to parent shell                                                                                               
Feb 03 07:05:42 ssc-vm-2302.colo.seagate.com motr-mkfs[23588]: got sub-shell error, terminating..                                                                                              
Feb 03 07:05:42 ssc-vm-2302.colo.seagate.com systemd[1]: motr-mkfs@0x7200000000000001:0xc.service: main process exited, code=exited, status=42/n/a                                             
Feb 03 07:05:42 ssc-vm-2302.colo.seagate.com systemd[1]: Failed to start Motr mkfs helper for 0x7200000000000001:0xc service.                                                                  
Feb 03 07:05:42 ssc-vm-2302.colo.seagate.com systemd[1]: Unit motr-mkfs@0x7200000000000001:0xc.service entered failed state.                                                                   
Feb 03 07:05:42 ssc-vm-2302.colo.seagate.com systemd[1]: motr-mkfs@0x7200000000000001:0xc.service failed.

Above issue can be seen during hare_setup init if /etc/sysconfig/motr configuration does not match the actual setup. In above case MOTR_M0D_IOS_BESEG_SIZE=5497558138880 value is too large for a VM setup and for a given meta data device.

References:

Mini provisioner specifications

Hare mini provisioning rfc

📽️ Watch the demo

Clone this wiki locally