intro_shmem(3)

intro_shmem - Introduction to logically shared memory access routines

As shipped in IRIX 6.5.19. Added in IRIX 6.5.19.

NAME
     intro_shmem - Introduction to logically shared memory access routines

DESCRIPTION
     The logically shared, distributed memory access (SHMEM) routines
     provide low-latency, high-bandwidth communication for use in highly
     parallelized scalable programs.

     The SHMEM routines are data passing library routines similar to
     message passing library routines.  They can be used as an alternative
     to message passing routines such as Message Passing Interface (MPI) or
     Parallel Virtual Machine (PVM).  Like the message passing routines,
     the SHMEM routines pass data between cooperating parallel processes.

     SHMEM routines can be used in programs that perform computations in
     separate address spaces and that explicitly pass data to and from
     different processes in the program.  These processes are also called
     processing elements (PEs).

     The SHMEM routines minimize the overhead associated with data passing
     requests, maximize bandwidth, and minimize data latency.  Data latency
     is the period of time that starts when a PE initiates a transfer of
     data and ends when a PE can use the data.

     SHMEM routines support remote data transfer through put operations,
     which transfer data to a different PE, and get operations, which
     transfer data from a different PE.  Other operations supported are
     work-shared broadcast and reduction, barrier synchronization, and
     atomic memory operations.  An atomic memory operation is an atomic
     read-and-update operation, such as a fetch-and-increment, on a remote
     or local data object.  The value read is guaranteed to be the value of
     the data object just prior to the update.

   SHMEM Routines
     The following SHMEM-related routines enhance the portabiliy of SHMEM
     programs across platforms.

     * PE queries:

          C/C++ only:         _num_pes(3I), _my_pe(3I)

          Fortran only:       NUM_PES(3I), MY_PE(3I)

     * Block data put routines:

          C/C++ and Fortran:  shmem_put32, shmem_put64, shmem_put128

          C/C++ only:         shmem_double_put, shmem_float_put,
                              shmem_int_put, shmem_long_put,
                              shmem_short_put

          Fortran only:       shmem_complex_put, shmem_integer_put,
                              shmem_logical_put, shmem_real_put

     * Block data get routines:

          C/C++ and Fortran:  shmem_get32, shmem_get64, shmem_get128

          C/C++ only:         shmem_double_get, shmem_float_get,
                              shmem_int_get, shmem_long_get,
                              shmem_short_get

          Fortran only:       shmem_complex_get, shmem_integer_get,
                              shmem_logical_get, shmem_real_get

     * Strided put routines:

          C/C++ and Fortran:  shmem_iput32, shmem_iput64, shmem_iput128

          C/C++ only:         shmem_double_iput, shmem_float_iput,
                              shmem_int_iput, shmem_long_iput,
                              shmem_short_iput

          Fortran only:       shmem_complex_iput, shmem_integer_iput,
                              shmem_logical_iput, shmem_real_iput

     * Strided get routines:

          C/C++ and Fortran:  shmem_iget32, shmem_iget64, shmem_iget128

          C/C++ only:         shmem_double_iget, shmem_float_iget,
                              shmem_int_iget, shmem_long_iget,
                              shmem_short_iget

          Fortran only:       shmem_complex_iget, shmem_integer_iget,
                              shmem_logical_iget, shmem_real_iget

     * Point-to-point synchronization routines:

          C/C++ only:         shmem_int_wait, shmem_int_wait_until,
                              shmem_long_wait, shmem_long_wait_until,
                              shmem_longlong_wait,
                              shmem_longlong_wait_until, shmem_short_wait,
                              shmem_short_wait_until

          Fortran:            shmem_int4_wait, shmem_int4_wait_until,
                              shmem_int8_wait, shmem_int8_wait_until

     * Barrier synchronization routines:

          C/C++ and Fortran:  shmem_barrier_all, shmem_barrier

     * Atomic memory fetch-and-operate (fetch-op) routines:

          C/C++ and Fortran:  shmem_swap

     * Reduction routines:

          C/C++ only:         shmem_int_and_to_all, shmem_long_and_to_all,
                              shmem_longlong_and_to_all,
                              shmem_short_and_to_all,
                              shmem_double_max_to_all,
                              shmem_float_max_to_all, shmem_int_max_to_all,
                              shmem_long_max_to_all,
                              shmem_longlong_max_to_all,
                              shmem_short_max_to_all,
                              shmem_double_min_to_all,
                              *shmem_float_min_to_all,
                              shmem_int_min_to_all, shmem_long_min_to_all,
                              shmem_longlong_min_to_all,
                              shmem_short_min_to_all,
                              shmem_double_sum_to_all,
                              shmem_float_sum_to_all, shmem_int_sum_to_all,
                              shmem_long_sum_to_all,
                              shmem_longlong_sum_to_all,
                              shmem_short_sum_to_all,
                              shmem_double_prod_to_all,
                              shmem_float_prod_to_all,
                              shmem_int_prod_to_all,
                              shmem_long_prod_to_all,
                              shmem_longlong_prod_to_all,
                              shmem_short_prod_to_all, shmem_int_or_to_all,
                              shmem_long_or_to_all,
                              shmem_longlong_or_to_all,
                              shmem_short_or_to_all, shmem_int_xor_to_all
                              shmem_long_xor_to_all
                              shmem_longlong_xor_to_all
                              shmem_short_xor_to_all,

          Fortran only:       shmem_int4_and_to_all, shmem_int8_and_to_all,
                              shmem_real4_max_to_all,
                              shmem_real8_max_to_all,
                              shmem_int4_max_to_all, shmem_int8_max_to_all,
                              shmem_real4_min_to_all,
                              shmem_real8_min_to_all,
                              shmem_int4_min_to_all, shmem_int8_min_to_all,
                              shmem_real4_sum_to_all,
                              shmem_real8_sum_to_all,
                              shmem_int4_sum_to_all, shmem_int8_sum_to_all,
                              shmem_real4_prod_to_all,
                              shmem_real8_prod_to_all,
                              shmem_int4_prod_to_all,
                              shmem_int8_prod_to_all, shmem_int4_or_to_all,
                              shmem_int8_or_to_all, shmem_int4_xor_to_all,
                              shmem_int8_xor_to_all

     * Broadcast routines:

          C/C++ and Fortran:  shmem_broadcast32, shmem_broadcast64

     * Generalized barrier synchronization routine:

          C/C++ and Fortran:  shmem_barrier

     * Cache management routines:

          C/C++ and Fortran:  shmem_udcflush, shmem_udcflush_line

     * Byte-granularity block put routines:

          C/C++ and Fortran:  shmem_putmem and shmem_getmem

          Fortran only:       shmem_character_put and shmem_character_get

     * Collect routines:

          C/C++ and Fortran:  shmem_collect32, shmem_collect64,
                              shmem_fcollect32, shmem_fcollect64

     * Atomic memory fetch-and-operate (fetch-op) routines:

          C/C++ only:         shmem_double_swap, shmem_float_swap,
                              shmem_int_cswap, shmem_int_fadd,
                              shmem_int_finc, shmem_int_swap,
                              shmem_long_cswap, shmem_long_fadd,
                              shmem_long_finc, shmem_long_swap,
                              shmem_longlong_cswap, shmem_longlong_fadd,
                              shmem_longlong_finc, shmem_longlong_swap

          Fortran only:       shmem_int4_cswap, shmem_int4_fadd,
                              shmem_int4_finc, shmem_int4_swap,
                              shmem_int8_swap, shmem_real4_swap,
                              shmem_real8_swap, shmem_int8_cswap

     * Atomic memory operation routines:

          Fortran only:       shmem_int4_add, shmem_int4_inc

     * Remote memory pointer function:

          C/C++ and Fortran:  shmem_ptr

     * Reduction routines:

          C/C++ only:         shmem_longdouble_max_to_all,
                              shmem_longdouble_min_to_all,
                              shmem_longdouble_prod_to_all,
                              shmem_longdouble_sum_to_all

          Fortran only:       shmem_real16_max_to_all,
                              shmem_real16_min_to_all,
                              shmem_real16_prod_to_all,
                              shmem_real16_sum_to_all

   Remotely Accessible Data Objects
     Typically, target or source arrays that reside on remote processing
     elements (PEs) are identified by passing the address of the
     corresponding data object on the local PE.  The local existence of a
     corresponding data object implies that a data object is symmetric as
     described on this man page.

     Symmetric data objects passed to SHMEM routines can be arrays or
     scalars.  A symmetric data object is one for which the local and
     remote addresses have a known relationship.  You can use SHMEM
     routines to access remote symmetric data objects by using the address
     of the corresponding data object on the local PE.

     The following data objects are symmetric:

     * Fortran data objects in common blocks or with the SAVE attribute.
       These data objects must not be defined in a dynamic shared object
       (DSO).

     * Non-stack C and C++ variables.  These data objects must not be
       defined in a DSO.

     * Fortran arrays allocated with shpalloc(3F)

     * C and C++ data allocated by shmalloc(3C)

     SHMEM collective routines that operate on the same data object on
     multiple PEs require that symmetric data objects be passed.  This
     restriction is for algorithm simplicity and efficiency.  These
     routines define the set of target PEs by the following triplet of
     arguments:  PE_start, logPE_stride, and PE_size.

   Collective Routines
     Some SHMEM routines, for example, shmem_broadcast(3) and
     shmem_float_sum_to_all(3), are classified as collective routines
     because they distribute work across a set of PEs.  They must be called
     concurrently by all PEs in the active set defined by the PE_start,
     logPE_stride, PE_size argument triplet.  The following man pages
     describe the SHMEM collective routines:

     * shmem_and(3)

     * shmem_barrier(3)

     * shmem_broadcast(3)

     * shmem_collect(3)

     * shmem_max(3)

     * shmem_min(3)

     * shmem_or(3)

     * shmem_prod(3)

     * shmem_sum(3)

     * shmem_xor(3)

   Using the Symmetric Work Array, pSync
     Multiple pSync arrays are often needed if a particular PE calls a
     SHMEM collective routine twice without intervening barrier
     synchronization.  Problems would occur if some PEs in the active set
     for call 2 arrive at call 2 before processing of call 1 is complete by
     all PEs in the call 1 active set.  You can use shmem_barrier() or
     shmem_barrier_all(3) to perform a barrier synchronization between
     consecutive calls to SHMEM collective routines.

     There are two special cases:

     * The shmem_barrier(3) routine allows the same pSync array to be used
       on consecutive calls as long as the active PE set does not change.

     * If the same collective routine is called multiple times with the
       same active set, the calls may alternate between two pSync arrays.
       The SHMEM routines guarantee that a first call is completely
       finished by all PEs by the time processing of a third call begins on
       any PE.

     Because the SHMEM routines restore pSync to its original contents,
     multiple calls that use the same pSync array do not require that pSync
     be reinitialized after the first call.

   SHMEM Function Inlining
     Some SHMEM functions that can be called from C/C++ are defined in the
     form of macros in the mpp/shmem.h header file.  These functions are
     inlined by default on some platforms.  To deactivate the automatic
     inlining of SHMEM functions from C/C++, add the following option to
     your C/C++ command line:  -D_SHMEM_MACRO_OPT=0.

   SHMEM Application Placement on NUMA Systems
     On non-uniform memory access (NUMA) systems, such as Origin series
     systems, SHMEM start-up processing ensures that the process associated
     with a SHMEM PE executes on a processor near the memory associated
     with a SHMEM PE.

     The following environment variables allow you to control the placement
     of the SHMEM application on the system:

     Variable            Description

     PAGESIZE_DATA       Specifies the desired page size in kilobytes for
                         program data areas.  Specify an integer value.  On
                         Origin series systems, supported values include
                         16, 64, 256, 1024, and 4096.

     SMA_BAR_COUNTER     Specifies the use of a simple counter barrier
                         algorithm.  By default, this variable is not
                         enabled for jobs with PE counts of 64 or more.

     SMA_BAR_DISSEM      Specifies the use of the alternate barrier
                         algorithm, the dissemination/butterfly, within the
                         shmem_barrier_all(3) function. This alternate
                         algorithm provides better performance on jobs with
                         larger PE counts.  The SMA_BAR_DISSEM option is
                         enabled for jobs with PE counts of 64 or higher.
                         By default, this variable is not enabled for jobs
                         with PE counts below 64.

     SMA_DBX             Specifies the PE number to be debugged.  If you
                         set SMA_DBX to n, PE n prints a message during
                         program startup, describing how to attach to it
                         with the DBX debugger.  PE n sleeps for seven
                         seconds.  If you set SMA_DBX to n,s, PE n will
                         sleep for s seconds.

     SMA_DPLACE_INTEROP_OFF
                         Disables a SHMEM/dplace interoperability feature
                         available beginning with IRIX 6.5.13.  By setting
                         this variable, you can obtain the behavior of
                         SHMEM with dplace on older releases of IRIX.  By
                         default, this variable is not enabled.

     SMA_DSM_CPULIST     Specifies a list of CPUs on which to run a SHMEM
                         application. To ensure that processes are linked
                         to CPUs, this variable should be used in
                         conjunction with SMA_DSM_MUSTRUN.

                         For an explanation of the syntax for this
                         environment variable, see the section entitled
                         "Using a CPU List."

     SMA_DSM_MUSTRUN     Enforces memory locality for SHMEM processes.  Use
                         of this feature ensures that each SHMEM process
                         will get a CPU and physical memory on the node to
                         which it was originally assigned.  This variable
                         has been observed to improve program performance
                         on IRIX systems running release 6.5.7 and earlier,
                         when running a program on a quiet system.  With
                         later IRIX releases, under certain circumstances,
                         setting this variable is not necessary.
                         Internally, this feature directs the library to
                         use the process_cpulink(3) function instead of
                         process_mldlink(3) to control memory placement.

                         SMA_DSM_MUSTRUN should not be used when the job is
                         submitted to miser (see miser_submit(1)) because
                         program hangs may result. By default, this
                         variable is not enabled.

                         The process_cpulink(3) function is inherited
                         across process fork(2) or sproc(2). For this
                         reason, when using mixed SHMEM/OpenMP
                         applications, it is recommended either that this
                         variable not be set, or that _DSM_MUSTRUN also be
                         set (see p_environ(5)).

     SMA_DSM_OFF         When set to any value, deactivates
                         processor-memory affinity control.  When set,
                         SHMEM processes run on any available processor,
                         whether or not it is near the memory associated
                         with that process.

     SMA_DSM_PPM         When set to an integer value, specifies the number
                         of processors to be mapped to every memory.  The
                         default is 2 on Origin 2000 systems. The default
                         is 4 on Origin 3000 systems.

     SMA_DSM_TOPOLOGY    Specifies the shape of the set of hardware nodes
                         on which the PE memories are allocated.  Set this
                         variable to one of the following values:

                         Value          Action

                         cube           A group of memory nodes that form a
                                        perfect hypercube. NPES/SMA_DSM_PPM
                                        must be a power of 2.  If a perfect
                                        hypercube is unavailable, a less
                                        restrictive placement will be used.

                         cube_fixed     A group of memory nodes that form a
                                        perfect hypercube.
                                        NPES/SMA_DSM_PPM must be a power of
                                        2.  If a perfect hypercube is
                                        unavailable, the placement will
                                        fail, disabling NUMA placement.

                         cpucluster     Any group of memory nodes.  The
                                        operating system attempts to place
                                        the group numbers close to one
                                        another, taking into account nodes
                                        with disabled processors.  (Default
                                        for IRIX 6.5.11 and higher).

                         free           Any group of memory nodes.  The
                                        operating system attempts to place
                                        the group numbers close to one
                                        another.  (Default for IRIX 6.5.10
                                        and earlier releases).

     SMA_DSM_VERBOSE     When set to any value, writes information about
                         process and memory placement to stderr.

     SMA_INFO            Prints information about environment variables
                         that can control libsma execution.

     SMA_SYMMETRIC_SIZE  Specifies the size, in bytes, of symmetric memory.
                         This is the size of static space plus per-PE
                         symmetric heap size.

     SMA_VERSION         Prints the libsma library release version.

   Using a CPU List
     You can manually select CPUs to use for a SHMEM application by setting
     the SMA_DSM_CPULIST shell variable.  This is treated as a comma and/or
     hyphen delineated ordered list, specifying a mapping of SHMEM
     processes to CPUs.  The shepherd process is not included in this list.

     Examples:

               Value          CPU Assignment

               8,16,32        Place three SHMEM processes on CPUs 8, 16,
                              and 32.

               32,16,8        Place the SHMEM process rank zero on CPU 32,
                              one on 16, and two on CPU 8.

               8-15,32-39     Place the SHMEM processes 0 through 7 on CPUs
                              8 to 15.  Place the SHMEM processes 8 through
                              15 on CPUs 32 to 39.

               39-32,8-15     Place the SHMEM processes 0 through 7 on CPUs
                              39 to 32.  Place the SHMEM processes 8
                              through 15 on CPUs 8 to 15.

     Note that the process rank is the value returned by _my_pe(3I).  CPUs
     are associated with the cpunum values given in the hardware
     graph(hwgraph(4)).

     The number of processors specified must equal the number of SHMEM
     processes (excluding the shepherd process) that will be used.  If an
     error occurs in processing the CPU list, the default placement policy
     is used.

   Using dplace(1)
     The environment variables described previously allow you to map SHMEM
     processes and memories with hardware processors and nodes.  The
     dplace(1) command, which is available on Origin series systems, can
     give you additional control over application placement.

     Perform the following steps to use the dplace(1) command with SHMEM
     programs:

     * Create file placefile with these contents:

          threads $NPES + 1
          memories ($NPES +1)/2 in topology cube
          distribute threads 1:$NPES across memories

     * Execute your program with NPES set to the number of PEs.  For
       example, to run with 4 PEs, invoke your program this way:

          env NPES=4 dplace -place placefile a.out

   Interoperability
     SHMEM routines can be used in conjunction with MPI message passing
     routines in the same application.  Programs that use both MPI and
     SHMEM should call MPI_Init and MPI_Finalize but omit the call to the
     start_pes routine.  SHMEM PE numbers are equal to the MPI rank within
     the MPI_COMM_WORLD environment variable.

     On IRIX clustered systems, you can use SHMEM to comunicate only with
     processes running on the same host.  Use the shmem_pe_accessible
     function to determine if a remote PE is accessible via SHMEM
     communication from the local PE.

   Compiling SHMEM Programs
     The SHMEM routines reside in libsma.so.

     The following sample command lines compile programs that include SHMEM
     routines:

     * IRIX systems:
          cc -64 c_program.c -lsma
          CC -64 cplusplus_program.c -lsma
          f90 -64 -LANG:recursive=on fortran_program.f -lsma
          f77 -64 -LANG:recursive=on fortran_program.f -lsma

     * IRIX systems with Fortran 90 version 7.2.1 available:
          f90 -64 -LANG:recursive=on -auto_use shmem_interface
          fortran_program.f -lsma

     The shmem_interface module is intended for use only with the -auto_use
     option.  This module provides compile-time checking of interfaces.
     The keyword=arg actual argument format is not supported for SHMEM
     subroutines defined in the shmem_interface procedure interface module.

     The IRIX N32 ABI, selected by the -n32 compiler option, is also
     supported by SHMEM, but is recommended only for small process counts
     and program memory sizes, due to the limitation in the size of virtual
     addresses imposed by the N32 ABI.  The use of the N64 ABI, selected by
     the -64 compiler option, is recommended for most SHMEM programs.

   Program Start-up
     The SHMEM implementation uses mapped files to render static memory
     remotely accessible on IRIX systems.  The result is that enough file
     space must be available in /var/tmp to accommodate a file of size
     npes * staticsz, where npes is the number of PEs and staticsz is the
     size of the program's static data area.  Static data includes Fortran
     common blocks and C/C++ static data.

     If a SHMEM program's memory requirements exceed available file space
     in /var/tmp, a SHMEM run-time error message is generated.  You can use
     the TMPDIR environment variable to select a directory in a file system
     with sufficient file space.

     To minimize SHMEM program start-up time, use symmetric memory
     allocated by the SHPALLOC(3F) or shmalloc(3C) routines instead of
     static memory.  Memory allocated by these routines does not require a
     corresponding file space allocation in /var/tmp.  This avoids problems
     when file space is low and executes more quickly when start-up
     processing needs to handle large static memory areas.

NOTES
     The SHMEM software is packaged with the Message Passing Toolkit.

ENVIRONMENT VARIABLES
     For information on environment variables that affect SHMEM routines,
     see the "SHMEM Application Placement on NUMA Systems" section of this
     man page.

EXAMPLES
     Example 1.  The following Fortran SHMEM program runs on IRIX systems:

          PROGRAM REDUCTION
          REAL VALUES, SUM
          COMMON /C/ VALUES
          REAL WORK
          CALL START_PES(0)
          VALUES = MY_PE()
          CALL SHMEM_BARRIER_ALL       ! Synchronize all PEs
          SUM = 0.0
          DO I = 0,NUM_PES()-1
             CALL SHMEM_REAL_GET(WORK, VALUES, 1, I)   ! Get next value
             SUM = SUM + WORK                          ! Sum it
          ENDDO
          PRINT*,'PE ',MY_PE(),' COMPUTED       SUM=',SUM
          CALL SHMEM_BARRIER_ALL
          END

     Since start_pes(3) is called with a value of 0, the number of PEs used
     to run the program is specified by the NPES environment variable.

     This Fortran program directs all PEs to sum simultaneously the numbers
     in the VALUES variable across all PEs.  By executing the program using
     the following command line, you can you can run the program with 4
     PEs:

          env NPES=4 a.out

     Example 2.  The following C SHMEM program runs on IRIX systems:

          #include <mpp/shmem.h>
          main()
          {
             long source[10] = { 1, 2, 3, 4, 5,
                                 6, 7, 8, 9, 10 };
          static long target[10];
          start_pes(0);
          if (_my_pe() == 0) {
             /* put 10 words into target on PE 1 */
             shmem_long_put(target, source, 10, 1);
          }
          shmem_barrier_all();  /* sync sender and receiver */
          if (_my_pe() == 1)
             shmem_udcflush();         /* needed on T90  */
             printf("target[0] on PE %d is %d\n", _my_pe(), target[0]);
          }

     In this C program, PE 0 sends 10 integers to the target array on PE 1.
     By executing the program using the following command line, you can you
     can run the program with 2 PEs:

          env NPES=2 a.out

SEE ALSO
     dplace(1)

     The following man pages also contain information on SHMEM routines.
     See the specific man pages for implementation information.

     CC(1), cld(1), f90(1), f90(1M), mpprun(1)

     shmem_add(3), shmem_and(3), shmem_barrier(3), shmem_barrier_all(3),
     shmem_broadcast(3), shmem_cache(3), shmem_collect(3), shmem_cswap(3),
     shmem_event(3), shmem_fadd(3), shmem_fence(3), shmem_finc(3),
     shmem_get(3), shmem_iget(3), shmem_inc(3), shmem_iput(3),
     shmem_ixput(3), shmem_lock(3), shmem_max(3), shmem_min(3),
     shmem_mswap(3), shmem_my_pe(3), shmem_or(3), shmem_prod(3),
     shmem_put(3), shmem_quiet(3), shmem_short_g(3) shmem_short_p(3),
     shmem_stack(3), shmem_sum(3), shmem_swap(3), shmem_wait(3),
     shmem_xor(3), start_pes(3)

     shmalloc(3C)

     shpalloc(3F)

     MY_PE(3I), NUM_PES(3I)

     For information on using SHMEM routines with message passing routines,
     see the Message Passing Toolkit: PVM Programmer's Manual, or the
     Message Passing Toolkit: MPI Programmer's Manual.