raid(1M)

raid - RAID subsystem maintenance utility

As shipped in IRIX 6.5. First release of IRIX 6.5.

NAME
     raid - RAID subsystem maintenance utility

SYNOPSIS
     /usr/sbin/raid [-L] -p [-v] [path]
     /usr/sbin/raid [-L] -c [-m] [-f] [path]
     /usr/sbin/raid [-L] -i [path]
     /usr/sbin/raid [-L] -d driveno [-f] path
     /usr/sbin/raid [-L] -r path
     /usr/sbin/raid [-L] -3 [-s depth] [-S totalsize] [-f] path
     /usr/sbin/raid [-L] -z [-f] path
     /usr/sbin/raid [-L] -l firmware [-f] path

DESCRIPTION
     raid is an administrative and maintenance utility used to query and
     configure RAID drives.  RAID is an acronym for "Redundant Arrays of
     Inexpensive Disks".  See the RAID System Administration Guide for more
     detailed information about the RAID product, and for additional
     background on the RAID concept.

     RAID drives are attached to the system via standard SCSI channels.  They
     look just like regular SCSI disks to IRIX in that they respond to read
     and write requests just the way a traditional SCSI disk would.  The major
     differences are in formatting procedure, error handling capabilities, and
     preventive maintenance operations.

OPTIONS
     -3 [-s stripe] [-S totalsize] [-f] path
          Format a RAID for RAID level 3 with the indicated stripe depth and
          totalsize.  Note that the stripe depth is the number of consecutive
          sectors on one physical disk that are part of a single stripe.  The
          more traditional term, "stripe size", would be 4 times the stripe
          depth because there are 4 data disks in one of these RAIDs.

          The default stripe depth is 32 sectors of 512 bytes each; that
          results in a stripe size of 64 KB.  The default totalsize is the
          total available data space on the disks, minus some sectors reserved
          for internal management and rounded to the next lower multiple of
          the stripe size.

          If -f is not specified, a yes/no confirmation will be required from
          the standard input.

          See the section on formatting for more information.

     -c [-m] [-f] [path]
          Check a RAID for down disks or inconsistent configuration.  If no
          path is specified, all RAIDs connected to the system will be
          checked.

          Conditions checked for include:  a down disk, a disk that was
          replaced while the RAID was inactive, and a controller that was
          replaced while the RAID was inactive.

          If the -m option is given, attempt to correct simple configuration
          errors.  If a disk was replaced while the RAID was inactive, that
          disk will be marked as down.  If a controller was replaced while the
          RAID was inactive, the controller will be reprogrammed to match the
          disks in the RAID.

          If the -m option is specified, changes are required, and the -f
          option is not given, a yes/no confirmation will be required from the
          standard input.

          See the sections on formatting, disk replacement, and preventive
          maintenance for more information.

     -d driveno [-f] path
          Force the disk in slot driveno on the indicated RAID to be marked as
          down.  This can be used on a disk that has begun to fail, but has
          not yet been marked as down, so that the disk can easily be
          identified and removed with a reduced risk of system interruption.

          If -f is not specified, a yes/no confirmation will be required from
          the standard input.

          CAUTION: Two or more disks can be marked as down with this option.
          If a second disk is inadvertently marked as down, the entire
          contents of the RAID will be lost.  The RAID must be reformatted and
          the contents restored from a back-up tape.

          See the section on disk replacement for more information.

     -i [path]
          Check the integrity of the parity information for a RAID.  This
          involves reading all of the data on the RAID and comparing it with
          the stored parity information for that data.  Any inconsistencies
          will be corrected by generating new parity from the existing data
          sectors and overwriting the old or mis-matched parity information.

          If a path is specified, only that RAID will be checked, otherwise
          all RAIDs connected to the system will be checked in sequence.

          See the section on preventive maintenance for more information.

     -l firmware [-f] path
          Download a new release of firmware for the RAID controller into the
          flash memory on the RAID controller.  firmware is the pathname of
          the file containing the firmware image.

          If -f is not specified, a yes/no confirmation will be required from
          the standard input.

     -p [-v] [path]
          Print configuration and status information for a RAID.  The pathname
          of the RAID, the configured RAID level, the stripe size, the total
          size of the RAID, and some status information will be printed to
          stderr.  The status information includes indications of when a disk
          is down, and when the RAID needs initializing or formatting.

          The -v option will additionally print the SCSI inquiry string
          information (vendor name, product name, and firmware revision level)
          for the firmware in the controller and for each of the disks in the
          RAID.

          If a path is specified, only that RAID will be shown, otherwise all
          RAIDs connected to the system will be shown.

     -r path
          Rebuild a failed disk drive in a RAID.  The RAID must have a slot
          that has been marked as down (failed) and that disk must have been
          replaced by a known good drive.  The contents of the existing data
          and parity sectors are used to reconstruct the data that was stored
          in the corresponding sectors on the failed drive.  The reconstructed
          data is then written to the replacement disk.  When that is
          complete, the new drive is marked as being up and the RAID is fully
          functional again.

          See the section on disk replacement for more information.

     -L   Send any error messages produced by raid to /var/adm/SYSLOG as well
          as stderr.  This option is useful when invoking raid from within a
          shell script.  Some error messages from raid, and all messages from
          IRIX , are unconditionally copied to /var/adm/SYSLOG.

FORMATTING
     RAID level 3 is optimized for support of a few very large serial data
     requests.  It is best suited for support of raw disk access with large
     transfers or for filesystems containing a smaller number of very large
     files that will be accessed via very large transfers.

     The process of formatting a RAID involves programming the RAID
     controller, zeroing out the contents of the drives so that we know that
     parity and data match, and writing configuration information to reserved
     areas on each disk.  A signature is generated that is unique to both this
     RAID and this format operation.

     The controller has a Non-Volatile RAM (NVRAM) that is used to store
     configuration information and some information required for error
     recovery.  The format signature is also stored in the NVRAM.
     There are a small number of sectors on each of the disks in a RAID that
     are reserved for administrative use, and are not user accessible.  At
     format time, the configuration parameters and the format signature are
     written into those reserved sectors.

     Since the signature is stored in the controller and on each disk, it
     allows us to examine a RAID and determine if any of the hardware has been
     replaced.

     Since fx doesn't know about the reserved sectors or programming the RAID
     controller, the raid program must be used to format a RAID device.  The
     fx command to format a device will only zero the data and parity disks in
     the RAID, it cannot be used to change the configuration parameters of the
     RAID.

DISK REPLACEMENT
     A feature of the disks in a RAID is that they will try to provide an
     early warning that they are about to fail.  If a disk fails outright, or
     is predicting that it will fail shortly, it should be replaced as soon as
     possible.  A message to that effect will be sent to the console and to
     /var/adm/SYSLOG in either case.  During the time that a disk is marked
     down, the RAID cannot protect the system from data loss due to hard, or
     even transient, disk errors.

     When a disk fails outright, it will be marked as down and the
     corresponding failed LED will be lit on the controller and in the front
     bezel of the disk.

     When a disk is predicting its own failure, it will not yet be marked as
     down.  It may fail shortly, but it has not yet failed.  As a result, the
     failed LED has not been lit and the task of removing the correct disk can
     be error prone.  In this case, it may be advantageous to use the -d
     option to force the failing disk down.  That will light the failed LED,
     and make finding the correct disk easier.  It will also eliminate the
     system delays associated with the RAID controller determining that a disk
     has been pulled out of the RAID.

     CAUTION: Note that removing a disk other than the failed disk will
     completely invalidate all data in the RAID.  The RAID will need to be
     reformatted and the contents reloaded from backup tape.

     Once the correct disk has been identified, simply pull it out of the RAID
     enclosure, whether the RAID is active or not.  Plug in the replacement
     disk in the same manner.

     After the new disk has spun up, which may take up to 1 minute, execute
     the raid -r path command to rebuild the data on the new disk.  This will
     take some time, and will be impacted by heavy data requests to the RAID.
     Once the rebuild completes, the disk will be marked as up and the RAID
     will again be ready to ensure data availability.
PREVENTIVE MAINTENANCE
     Preventive maintenance should be performed on all RAIDs at regular
     intervals in order to ensure data integrity in the event of a disk
     failure.  In this case, preventive maintenance means running a check that
     the parity information in the RAID matches the data in the RAID.  That
     check can be invoked via executing the raid -i command.

     If any discrepancies are found between the data and the corresponding
     parity information, the parity is regenerated from the existing data and
     a message is sent to the console and to /var/adm/SYSLOG.

     The -i option will check each RAID in sequence.  By invoking several
     copies of raid, each with the pathname of a RAID, several RAIDs can be
     checked in parallel.  Note that checking the integrity of a RAID will
     impact the available bandwidth from the RAID, and heavy I/O requests to
     the RAID will impact the time to perform the integrity check.

     The system performs a raid -cmf command at every boot and system
     shutdown.

ERROR HANDLING
     Occasionally, a SCSI bus may be reset in order to clear an error on one
     of the devices attached to that bus.  A bus reset effectively resets all
     of the devices on the bus, including the RAID controller.  If the RAID
     controller was in the middle of writing to the RAID, it is possible for
     data and the corresponding parity information to become inconsistent.
     The usraid driver checks for this after all bus resets and will
     regenerate the parity information for all stripes left inconsistent for
     that reason.  A message will be sent to the console and to
     /var/adm/SYSLOG detailing which stripes had to be fixed.  The system will
     then retry any operations that were in progress at the time of the reset.

     If the controller ever reports an internal hardware error, the system
     will invoke the controller's internal diagnostics in an attempt to either
     verify the error or show that it was transient.  All hardware errors are
     detailed in messages sent to the console and to /var/adm/SYSLOG.

DEVICE NODES
     The MAKEDEV script will create device nodes for all possible disk drives
     on all configured SCSI busses.  After a RAID device has been configured
     and formatted, it responds to read and write commands just like a normal
     SCSI disk.  In fact, the dksc (standard SCSI disk) driver could be used
     to access the RAID.  Unfortunately, the maintenance and error recovery
     operations detailed above will not be performed by the dksc driver, while
     they would be performed by the usraid driver.  Without those maintenance
     and error recovery operations, there is a likelihood of data loss.  As a
     result, the MAKEDEV script will create device nodes for all currently
     connected RAID devices, but not all possible RAID devices, and will
     remove the dksc nodes for the SCSI bus and target ID corresponding to
     those RAIDs.
FILES
     /dev/dsk/rad*,
     /dev/rdsk/rad*

SEE ALSO
     usraid(7M), fx(1M), hinv(1M), mknod(1M), mount(1M), dvhtool(1M),
     MAKEDEV(1M), Add_disk(1), and vh(7M).