raid(1M)
raid - RAID subsystem maintenance utility
Showing IRIX 6.5.30 (default release). Unchanged since IRIX 6.5.
NAME raid - RAID subsystem maintenance utility SYNOPSIS /usr/sbin/raid [-L] -p [-v] [path] /usr/sbin/raid [-L] -c [-m] [-f] [path] /usr/sbin/raid [-L] -i [path] /usr/sbin/raid [-L] -d driveno [-f] path /usr/sbin/raid [-L] -r path /usr/sbin/raid [-L] -3 [-s depth] [-S totalsize] [-f] path /usr/sbin/raid [-L] -z [-f] path /usr/sbin/raid [-L] -l firmware [-f] path DESCRIPTION raid is an administrative and maintenance utility used to query and configure RAID drives. RAID is an acronym for "Redundant Arrays of Inexpensive Disks". See the RAID System Administration Guide for more detailed information about the RAID product, and for additional background on the RAID concept. RAID drives are attached to the system via standard SCSI channels. They look just like regular SCSI disks to IRIX in that they respond to read and write requests just the way a traditional SCSI disk would. The major differences are in formatting procedure, error handling capabilities, and preventive maintenance operations. OPTIONS -3 [-s stripe] [-S totalsize] [-f] path Format a RAID for RAID level 3 with the indicated stripe depth and totalsize. Note that the stripe depth is the number of consecutive sectors on one physical disk that are part of a single stripe. The more traditional term, "stripe size", would be 4 times the stripe depth because there are 4 data disks in one of these RAIDs. The default stripe depth is 32 sectors of 512 bytes each; that results in a stripe size of 64 KB. The default totalsize is the total available data space on the disks, minus some sectors reserved for internal management and rounded to the next lower multiple of the stripe size. If -f is not specified, a yes/no confirmation will be required from the standard input. See the section on formatting for more information. -c [-m] [-f] [path] Check a RAID for down disks or inconsistent configuration. If no path is specified, all RAIDs connected to the system will be checked. Conditions checked for include: a down disk, a disk that was replaced while the RAID was inactive, and a controller that was replaced while the RAID was inactive. If the -m option is given, attempt to correct simple configuration errors. If a disk was replaced while the RAID was inactive, that disk will be marked as down. If a controller was replaced while the RAID was inactive, the controller will be reprogrammed to match the disks in the RAID. If the -m option is specified, changes are required, and the -f option is not given, a yes/no confirmation will be required from the standard input. See the sections on formatting, disk replacement, and preventive maintenance for more information. -d driveno [-f] path Force the disk in slot driveno on the indicated RAID to be marked as down. This can be used on a disk that has begun to fail, but has not yet been marked as down, so that the disk can easily be identified and removed with a reduced risk of system interruption. If -f is not specified, a yes/no confirmation will be required from the standard input. CAUTION: Two or more disks can be marked as down with this option. If a second disk is inadvertently marked as down, the entire contents of the RAID will be lost. The RAID must be reformatted and the contents restored from a back-up tape. See the section on disk replacement for more information. -i [path] Check the integrity of the parity information for a RAID. This involves reading all of the data on the RAID and comparing it with the stored parity information for that data. Any inconsistencies will be corrected by generating new parity from the existing data sectors and overwriting the old or mis-matched parity information. If a path is specified, only that RAID will be checked, otherwise all RAIDs connected to the system will be checked in sequence. See the section on preventive maintenance for more information. -l firmware [-f] path Download a new release of firmware for the RAID controller into the flash memory on the RAID controller. firmware is the pathname of the file containing the firmware image. If -f is not specified, a yes/no confirmation will be required from the standard input. -p [-v] [path] Print configuration and status information for a RAID. The pathname of the RAID, the configured RAID level, the stripe size, the total size of the RAID, and some status information will be printed to stderr. The status information includes indications of when a disk is down, and when the RAID needs initializing or formatting. The -v option will additionally print the SCSI inquiry string information (vendor name, product name, and firmware revision level) for the firmware in the controller and for each of the disks in the RAID. If a path is specified, only that RAID will be shown, otherwise all RAIDs connected to the system will be shown. -r path Rebuild a failed disk drive in a RAID. The RAID must have a slot that has been marked as down (failed) and that disk must have been replaced by a known good drive. The contents of the existing data and parity sectors are used to reconstruct the data that was stored in the corresponding sectors on the failed drive. The reconstructed data is then written to the replacement disk. When that is complete, the new drive is marked as being up and the RAID is fully functional again. See the section on disk replacement for more information. -L Send any error messages produced by raid to /var/adm/SYSLOG as well as stderr. This option is useful when invoking raid from within a shell script. Some error messages from raid, and all messages from IRIX , are unconditionally copied to /var/adm/SYSLOG. FORMATTING RAID level 3 is optimized for support of a few very large serial data requests. It is best suited for support of raw disk access with large transfers or for filesystems containing a smaller number of very large files that will be accessed via very large transfers. The process of formatting a RAID involves programming the RAID controller, zeroing out the contents of the drives so that we know that parity and data match, and writing configuration information to reserved areas on each disk. A signature is generated that is unique to both this RAID and this format operation. The controller has a Non-Volatile RAM (NVRAM) that is used to store configuration information and some information required for error recovery. The format signature is also stored in the NVRAM. There are a small number of sectors on each of the disks in a RAID that are reserved for administrative use, and are not user accessible. At format time, the configuration parameters and the format signature are written into those reserved sectors. Since the signature is stored in the controller and on each disk, it allows us to examine a RAID and determine if any of the hardware has been replaced. Since fx doesn't know about the reserved sectors or programming the RAID controller, the raid program must be used to format a RAID device. The fx command to format a device will only zero the data and parity disks in the RAID, it cannot be used to change the configuration parameters of the RAID. DISK REPLACEMENT A feature of the disks in a RAID is that they will try to provide an early warning that they are about to fail. If a disk fails outright, or is predicting that it will fail shortly, it should be replaced as soon as possible. A message to that effect will be sent to the console and to /var/adm/SYSLOG in either case. During the time that a disk is marked down, the RAID cannot protect the system from data loss due to hard, or even transient, disk errors. When a disk fails outright, it will be marked as down and the corresponding failed LED will be lit on the controller and in the front bezel of the disk. When a disk is predicting its own failure, it will not yet be marked as down. It may fail shortly, but it has not yet failed. As a result, the failed LED has not been lit and the task of removing the correct disk can be error prone. In this case, it may be advantageous to use the -d option to force the failing disk down. That will light the failed LED, and make finding the correct disk easier. It will also eliminate the system delays associated with the RAID controller determining that a disk has been pulled out of the RAID. CAUTION: Note that removing a disk other than the failed disk will completely invalidate all data in the RAID. The RAID will need to be reformatted and the contents reloaded from backup tape. Once the correct disk has been identified, simply pull it out of the RAID enclosure, whether the RAID is active or not. Plug in the replacement disk in the same manner. After the new disk has spun up, which may take up to 1 minute, execute the raid -r path command to rebuild the data on the new disk. This will take some time, and will be impacted by heavy data requests to the RAID. Once the rebuild completes, the disk will be marked as up and the RAID will again be ready to ensure data availability. PREVENTIVE MAINTENANCE Preventive maintenance should be performed on all RAIDs at regular intervals in order to ensure data integrity in the event of a disk failure. In this case, preventive maintenance means running a check that the parity information in the RAID matches the data in the RAID. That check can be invoked via executing the raid -i command. If any discrepancies are found between the data and the corresponding parity information, the parity is regenerated from the existing data and a message is sent to the console and to /var/adm/SYSLOG. The -i option will check each RAID in sequence. By invoking several copies of raid, each with the pathname of a RAID, several RAIDs can be checked in parallel. Note that checking the integrity of a RAID will impact the available bandwidth from the RAID, and heavy I/O requests to the RAID will impact the time to perform the integrity check. The system performs a raid -cmf command at every boot and system shutdown. ERROR HANDLING Occasionally, a SCSI bus may be reset in order to clear an error on one of the devices attached to that bus. A bus reset effectively resets all of the devices on the bus, including the RAID controller. If the RAID controller was in the middle of writing to the RAID, it is possible for data and the corresponding parity information to become inconsistent. The usraid driver checks for this after all bus resets and will regenerate the parity information for all stripes left inconsistent for that reason. A message will be sent to the console and to /var/adm/SYSLOG detailing which stripes had to be fixed. The system will then retry any operations that were in progress at the time of the reset. If the controller ever reports an internal hardware error, the system will invoke the controller's internal diagnostics in an attempt to either verify the error or show that it was transient. All hardware errors are detailed in messages sent to the console and to /var/adm/SYSLOG. DEVICE NODES The MAKEDEV script will create device nodes for all possible disk drives on all configured SCSI busses. After a RAID device has been configured and formatted, it responds to read and write commands just like a normal SCSI disk. In fact, the dksc (standard SCSI disk) driver could be used to access the RAID. Unfortunately, the maintenance and error recovery operations detailed above will not be performed by the dksc driver, while they would be performed by the usraid driver. Without those maintenance and error recovery operations, there is a likelihood of data loss. As a result, the MAKEDEV script will create device nodes for all currently connected RAID devices, but not all possible RAID devices, and will remove the dksc nodes for the SCSI bus and target ID corresponding to those RAIDs. FILES /dev/dsk/rad*, /dev/rdsk/rad* SEE ALSO usraid(7M), fx(1M), hinv(1M), mknod(1M), mount(1M), dvhtool(1M), MAKEDEV(1M), Add_disk(1), and vh(7M).