وو

وحید آنلاین . آرشیو وبلاگ وحیدمی دات آی آر . شرکت بیان. vahidmy.blog.ir

وو

وحید آنلاین . آرشیو وبلاگ وحیدمی دات آی آر . شرکت بیان. vahidmy.blog.ir

PMADDWD

PMADDWD 

Usage:  PMADDWD    dest,src                        Modifies flags: None

Multiplies the individual signed words of the 'dest' by the corresponding signed words of the 'src', producing temporary signed, doubleword results. The adjacent doubleword results are then summed and stored in the 'dest' operand.

MMX Packed Multiply and Add


PMADDWD mm1,mm2/m64           ; 0F F5 /r             [PENT,MMX]

PMADDWD xmm1,xmm2/m128        ; 66 0F F5 /r     [WILLAMETTE,SSE2]


PMADDWD treats its two inputs as vectors of signed words. It multiplies corresponding elements of the two operands, giving doubleword results. These are then added together in pairs and stored in the destination operand.


The operation of this instruction is:

   dst[0-31]   := (dst[0-15] * src[0-15])

                               + (dst[16-31] * src[16-31]);

   dst[32-63]  := (dst[32-47] * src[32-47])

                               + (dst[48-63] * src[48-63]);

The following applies to the SSE version of the instruction:

   dst[64-95]  := (dst[64-79] * src[64-79])

                               + (dst[80-95] * src[80-95]);

   dst[96-127] := (dst[96-111] * src[96-111])

                               + (dst[112-127] * src[112-127]).

The 'src' operand can be an MMX technology register or a 64- bit memory location, or it can be an XMM register or a 128-bit memory location. The 'dest' operand can be an MMX technology register or an XMM register.


EXAMPLE:

pmaddwd Label xmm1 


PINSRW

PINSRW 

Usage:  PINSRW    dest,src,count                Modifies flags: None

Copies a word from the 'src' (second operand) and inserts it in the 'dest' (first operand) at the location specified with the 'count' operand (third operand). (The other words in the 'dest'  register are left untouched.)

Insert Word


PINSRW mm,r16/r32/m16,imm8    ;0F C4 /r ib      [KATMAI,MMX]

PINSRW xmm,r16/r32/m16,imm8   ;66 0F C4 /r ib   [WILLAMETTE,SSE2]


PINSRW loads a word from a 16-bit register (or the low half of a 32-bit register), or from memory, and loads it to the word position in the destination register, pointed at by the count operand (third operand). If the destination is an MMX register, the low two bits of the count byte are used, if it is an XMM register the low 3 bits are used. The insertion is done in such a way that the other words from the destination register are left untouched.


The 'src' operand can be a general-purpose register or a 16-bit memory location. (When the 'src' operand is a general-purpose register, the low word of the register is copied.) The 'dest'  operand can be an MMX technology register or an XMM register. 

The 'count' operand is an 8-bit immediate. When specifying a word location in an MMX technology register, the 2 least-significant bits of the count operand specify the location; for an XMM register, the 3 least-significant bits specify the location.


EXAMPLE:

pinsrw xmm1 ecx 4

 

PF2IW

PF2IW 

Usage:  PF2IW  dest,src                                         Modifies flags: None

Converts a register containing single-precision floating-point operands to 16-bit signed integers using truncation. Arguments outside the range representable by signed 16-bit integers are saturated to the largest and smallest 16-bit integer, depending on their sign. All results are sign-extended to 32-bits. 

Packed Word Integer to Single-Precision FP Convert


PI2FW mm1,mm2/m64             ; 0F 0F /r 0C          [PENT,3DNOW]


PF2IW converts two signed 16-bit integers in the source operand to single-precision FP values, and stores them in the destination operand. The input values are in the low word of each doubleword.


See the PF2ID, PI2FW, and PI2FD instructions.


Example:

pf21w mm1 Label


PI2FW

PI2FW

Usage: PI2FW  dest,src                                   Modifies flags: None

Conversion of signed integers to single precision floating point.

Packed 16 Bit Integer to FP Conversion


PI2FW  mm1, mm2/mem64            ; 0F 0F /0Ch         [3DNow]


PI2FW converts a register containing signed, 16-bit integers to single-precision,

floating-point operands.


PI2FW  mm1, mm2  performs the following operations:

mm1 [31:0 ]     =float (mm2 [15:0 ])

mm1 [63:32 ]   =float (mm2 [47:32 ])


PI2FW  mm1, mem64  performs the following operations:

mmreg [31:0 ]  =float (mem64 [15:0 ])

mmreg [63:32 ]=float (mem64 [47:32 ]) 


 See also the PI2FD, PF2IW, and PF2ID instructions.


EXAMPLE:

pi2fw mm1 mm2


PI2FD

PI2FD 

Usage: PI2FD  dest,src                               Modifies flags: None

PI2FD is a vector instruction that converts a vector register containing signed, 32-bit integers to single-precision, floating-point operands. When PI2FD converts an input operand with more significant digits than are available in the output, the output is truncated.

Packed Doubleword Integer to Single-Precision FP Convert


PI2FD mm1,mm2/m64             ; 0F 0F /r 0D          [PENT,3DNOW]


PF2ID converts two signed 32-bit integers in the source operand to single-precision FP values, using truncation of significant digits, and stores them in the destination operand.


Example:

pi2fd mm1 mm2   

PFSUBR

PFSUBR 

Usage:  PFSUBR    dest,src                         Modifies flags: None

Performs subtraction of the destination operand from the source operand. Both operands are single-precision, floating-point operands with 24-bit significands. 

Packed Single-Precision FP Reverse Subtract



PFSUBR mm1,mm2/m64            ; 0F 0F /r AA          [PENT,3DNOW]


PFSUBR subtracts the single-precision FP values in the destination from those in the source, and stores the result in the destination operand.


   dst[0-31]  := src[0-31]  - dst[0-31],

   dst[32-63] := src[32-63] - dst[32-63].


EXAMPLE:

pfsubr mm1 mm2


PFSUB

PFSUB 

Usage:  PFSUB    dest,src                           Modifies flags: None

Performs subtraction of the source operand from the destination operand. Both operands are single-precision, floating-point operands with 24-bit significands.

Packed Single-Precision FP Subtract


PFSUB mm1,mm2/m64             ; 0F 0F /r 9A          [PENT,3DNOW]


PFSUB subtracts the single-precision FP values in the source from those in the destination, and stores the result in the destination operand.


   dst[0-31]  := dst[0-31]  - src[0-31],

   dst[32-63] := dst[32-63] - src[32-63].


EXAMPLE:

pfsub mm1 mm2

PFRSQRT

PFRSQRT 

Usage:  PFRSQRT    dest,src                       Modifies flags: None

Returns a low-precision estimate of the reciprocal square root of the source operand. The single result value is duplicated in both high and low halves of this instruction's 64-bit result.

Packed Single-Precision FP Reciprocal Square Root Approximation


PFRSQRT mm1,mm2/m64           ; 0F 0F /r 97          [PENT,3DNOW]


PFRSQRT performs a low precision estimate of the reciprocal square root of the low-order single-precision FP value in the source operand, storing the result in both halves of the destination register. The result is accurate to 15 bits. Negative operands are treated as positive operands for purposes of reciprocal square root computation, with the sign of the result the same as the sign of the source operand. 


For higher precision reciprocals, this instruction should be followed by two more instructions: PFRSQIT1 and PFRCPIT2. This will result in a 24-bit accuracy. For more details, see the AMD 3DNow! technology manual.


EXAMPLE:

pfrsqrt mm1 mm2 

PFRSQIT1

PFRSQIT1

Usage:  PFRSQRT1    dest,src                      Modifies flags: None

Performs the first intermediate step in the Newton-Raphson iteration to refine the reciprocal square root approximation produced by the PFSQRT instruction (the second and final step completes the iteration and is accurate to 24 bits).

Packed Single-Precision FP Reciprocal Square Root,

First Iteration Step


PFRSQIT1 mm1,mm2/m64          ; 0F 0F /r A7          [PENT,3DNOW]


PFRSQIT1 performs the first intermediate step in the calculation of the reciprocal square root of a single-precision FP value. The first source value mm1 is the square of the result of a PFRSQRT  instruction, and the second source value mm2/m64 is the original value.


For the final step in a calculation, returning the full 24-bit accuracy of a single-precision FP value, see PFRCPIT2. For more details, see the AMD 3DNow! technology manual.


EXAMPLE:

pfrsqrt1 mm1 mm2 

 

PFRCPIT2

PFRCPIT2 

Usage:  PFRCPIT2  dest,src                         Modifies flags: None

Performs the second and final intermediate step in the Newton-Raphson iteration to refine the reciprocal or reciprocal square root approximation produced by the PFRCP and PFSQRT instructions, respectively.

Packed Single-Precision FP Reciprocal/ Reciprocal Square Root, Second Iteration Step


PFRCPIT2 mm1,mm2/m64          ; 0F 0F /r B6          [PENT,3DNOW]


PFRCPIT2 performs the second and final intermediate step in the calculation of a reciprocal or reciprocal square root, refining the values returned by the PFRCP and PFRSQRT instructions, respectively.


The first source value (mm1) is the output of either a PFRCPIT1  or a PFRSQIT1 instruction, and the second source is the output of either the PFRCP or the PFRSQRT instruction. For more details, see the AMD 3DNow! technology manual.


Example:

pfrcpit2 mm1 mm2