intro_blas(3F)

INTRO_BLAS - Introduction to Basic Linear Algebra Subprograms

As shipped in IRIX 6.5.7. Added in IRIX 6.5.5.

NAME
     INTRO_BLAS - Introduction to Basic Linear Algebra Subprograms

IMPLEMENTATION
     See the individual man pages for implementation details

DESCRIPTION
     BLAS is a library of routines that perform basic operations involving
     matrices and vectors. They were designed as a way of achieving
     efficiency in the solution of linear algebra problems. The BLAS, as
     they are now commonly called, have been very successful and have been
     used in a wide range of software, including LINPACK, LAPACK and many
     of the algorithms published by the ACM Transactions on Mathematical
     Software. They are an aid to clarity, portability, modularity and
     maintenance of software, and have become the de facto standard for
     elementary vector and matrix operations.

     The BLAS promote modularity by identifying frequently occurring
     operations of linear algebra and by specifying a standard interface to
     these operations.  Efficiency is achieved through optimization within
     the BLAS without altering the higher-level code references them.

     There are three levels of BLAS:

     * Level 1: The original set of BLAS, commonly referred as the Level 1
       BLAS, perform low-level operations such as dot-product and the
       adding of a multiple of one vector to another.

       Typically these operations involve O(n) floating point operations
       and O(n) data items moved (loaded or stored), where n is the length
       of the vectors. The Level 1 BLAS permit efficient implementation on
       scalar machines, but the ratio of floating-point operations to data
       movement is too low to be effective on most vector or parallel
       hardware.

     * Level 2: The Level 2 BLAS perform matrix-vector operations that
       occur frequently in the implementation of many of the most common
       linear algebra algorithms.
                                 2
       These routines involve O(n ) floating point operations. Algorithms
       that use Level 2 BLAS can be very efficient on vector computers, but
       are not well suited to computers with a hierarchy of memory (such as
       cache memory).

     * Level 3:  The Level 3 BLAS are targeted at matrix-matrix operations.
                       3                                                2
       They involve O(n ) floating point operations, but only create O(n )
       data movement. These operations permit efficient reuse of data that
       resides in cache and create what is often called the surface-to-
       volumne effect for the ratio of computations to data movement. In
       addition, matrices can be partitioned into blocks, and operations on
       distinct blocks can be performed in parallel, and within the
       operations on each block, scalar or vector operations may be
       performed in parallel.

     BLAS2 and BLAS3 modules are optimized and parallelized to take
     advantage of Silicon Graphics' RISC parallel architecture. The best
     performances are achieved for BLAS3 routines (for exmaple, DGEM) where
     outer-loop unrolling and blocking techniques were applied to take
     advantage of the memory cache. The performance of BLAS2 routines (for
     example, DGEMV) is sensitive to the size of the problem; for large
     sizes the high rate of cache miss slows down the algorithms.

     LAPACK algorithms use (preferably_ BLAS3 modules and are the most
     efficient.  LINPACK uses only BLAS1 modules and therefore is less
     efficient than LAPACK.

     To link with libblas, f77 to load all the Fortran Libraries required;
     otherwise include -lftn in your link line.  For R8000 and R10000 based
     machines, use the MIPS4 version by using the -mips4 option when
     linking, as in this example:

          f77 -mips4
          -o foobar.out foo.o bar.o
          -lblas

     To use the parallelized version, use the -mips4 option as follows:

          f77 -mips4 -mp
          -o foobar.out foo.o bar.o
          -lblas_mp

   Increment arguments
     A vector's description consists of the name of the array (x or y)
     followed by the storage spacing (increment) in the array of vector
     elements (incx or incy).  The increment can be positive or negative.
     When a vector x consists of n elements, the corresponding actual array
     arguments must be of a length at least 1+(n-1)*|incx| .  For a
     negative increment, the first element of x is assumed to be x(1+(n-1)*
     |incx|) .  The standard specification of _SCAL, _NRM2, _ASUM, and
     I_AMAX does not define their behavior for negative increments, so this
     functionality is an extension to the standard BLAS.

     Setting an increment argument to 0 can cause unpredictable results.

   Multiple routine man pages
     Many of the routines are available in real (single-precision),
     complex, double precision and double complex versions.  Often little
     or no difference exists between these versions, other than the data
     types of some inputs and outputs.  In this case, the routines are
     described on the same man page, and that man page is named after the
     real or complex routine.

     The following data types are used in these routines:

     * REAL: Fortran "real" data type, 32-bit floating point; these routine
       names begin with S.

     * COMPLEX: Fortran "complex" data type, two 32-bit floating point
       reals; these routine names begin with C.

     * DOUBLE PRECISION: Fortran "double precision" data type, 64-bit
       floating point; these routine names begin with D.

     * DOUBLE COMPLEX: Fortran "double complex" data type, two 64-bit
       floating point doubles; these routine names begin with Z.

     The man(1) command can find a man page online by either the real,
     complex, double precision, or double complex name.

     The following table describes the naming conventions for these
     routines:

     -------------------------------------------------------------
                                                       64-bit
                                                       complex
                           64-bit real                 (double
                           (double       32-bit        complex
             32-bit real   precision)    complex       precision)
     -------------------------------------------------------------
     form:   Sname         Dname        Cname         Zname
     example:SAXPY         DAXPY        CAXPY         ZAXPY
     -------------------------------------------------------------

   Fortran type declaration for functions
     Always declare the data type of external functions.  Declaring the
     data type of the complex Level 1 BLAS functions is particularily
     important because, based on the first letter of their names and the
     Fortran data typing rules, the default implied data type would be
     REAL.

Summary of routines
     The following tables list the available BLAS routines.

   BLAS Level 1
 -------------------------------------------------------------------------
 Function                Prefix and suffix (if provided)     Man page name
 -------------------------------------------------------------------------
 dot product              s-   d-   c-u   c-c  z-u   z-c   dot
 y = a*x + y              s-   d-         c-         z-    axpy
 setup Givens rotation    s-   d-                          rotg
 apply Givens rotation    s-   d-         cs-        zd-   rot
 copy x into y            s-   d-         c-         z-    copy
 swap x and y             s-   d-         c-         z-    swap
 Euclidean norm           s-   d-         sc-        dz-   nrm2
 sum of absolute values   s-   d-         sc-        dz-   asum
 x = a*x                  s-   d-   cs-   c-   zd-   z-    scal
 index of max abs value  is-   id-        ic-        iz-   amax
 -------------------------------------------------------------------------

   BLAS Level 2
     In the following tables, these abbreviations are used:

     MV   Matrix vector multiply

     R    Rank one update to a matrix

     R2   Rank two update to a matrix

     SV   Solving certain triangular matrix problems.

  single precision Level 2 BLAS     |     Double precision Level 2 BLAS
  -----------------------------------------------------------------------
          MV   R    R2  SV          |             MV   R    R2  SV
  SGE     x    x                    |     DGE     x    x
  SGB     x                         |     DGB     x
  SSP     x    x    x               |     DSP     x    x    x
  SSY     x    x    x               |     DSY     x    x    x
  SSB     x                         |     DSB     x
  STR     x              x          |     DTR     x              x
  STB     x              x          |     DTB     x              x
  STP     x              x          |     DTP     x              x

  complex  Level 2 BLAS             | Double precision complex Level 2 BLAS
  -----------------------------------------------------------------------
          MV   R     RC   RU  R2  SV|          MV   R     RC   RU  R2  SV
  CGE     x          x    x         |  ZGE     x          x    x
  CGB     x                         |  ZGB     x
  CHE     x    x              x     |  ZHE     x    x              x
  CHP     x    x              x     |  ZHP     x    x              x
  CHB     x                         |  ZHB     x
  CTR     x                       x |  ZTR     x                       x
  CTB     x                       x |  ZTB     x                       x
  CTP     x                       x |  ZTP     x                       x

   BLAS Level 3
     In the following tables, these abbreviations are used:

     MM   Matrix matrix multiply

     RK   Rank-k update to a matrix

     R2K  Rank-2k update to a matrix

     SM   Solving triangular matrix with many right-hand-sides.

  single precision Level 3 BLAS     |     Double precision Level 3 BLAS
  -----------------------------------------------------------------------
          MM   RK   R2K SM          |             MM   RK   R2K SM
  SGE     x                         |     DGE     x
  SSY     x    x    x               |     DSY     x    x    x
  STR     x              x          |     DTR     x              x

  complex  Level 3 BLAS             | Double precision complex Level 3 BLAS
  -----------------------------------------------------------------------
          MM   RK   R2K SM          |             MM   RK   R2K SM
  CGE     x                         |     ZGE     x
  CSY     x    x    x               |     ZSY     x    x    x
  CHE     x    x    x               |     ZHE     x    x    x
  CTR     x              x          |     ZTR     x              x

FILES
     /usr/lib/libblas.a
     /usr/lib/libblas_mp.a
     /usr/include/cblas.h

NOTES
     libblas does not currently support reshaped arrays.

SEE ALSO
     S.P. Datardina, J.J. Du Croz, S.J. Hammarling and M.W. Pont, "A
     Proposed Specification of BLAS Routines in C", NAG Technical Report
     TR6/90.

     Lawson, C., Hanson, R., Kincaid, D., and Krogh, F., "Basic Linear
     Algebra Subprograms for Fortran Usage," ACM Transactions on
     Mathematical Software,
      5 (1979),
     pp. 308 - 325.

     J.Dongarra, J.DuCroz, S.Hammarling, and R.Hanson, "An extended set of
     Fortran Basic Linear Algebra Subprograms", ACM Trans. on Math. Soft.
     14, 1(1988) 1-32

     J.Dongarra, J.DuCroz, I.Duff,and S.Hammarling, "An set of level 3
     Basic Algebra Subprograms", ACM Trans on Math Soft( Dec 1989)

     This man page is available only online.