XADD
Usage: XADD dest,src Modifies flags: AF CF OF PF SF ZF
Exchange and Add
XADD r/m8,reg8 ; 0F C0 /r [486]
XADD r/m16,reg16 ; o16 0F C1 /r [486]
XADD r/m32,reg32 ; o32 0F C1 /r [486]
XADD exchanges the values in its two operands, and then adds them together and writes the result into the destination (first) operand. This instruction can be used with a LOCK prefix for multi-processor synchronisation purposes.
EXAMPLE:
mov eax 00_0000_1111
mov ecx 10
xadd eax ecx ; eax = 25d > 19h
WRMSR
Usage: WRMSR Modifies flags: None
Writes the contents of EDX:EAX into the 64-bit model specific register (MSR) specified in the ECX register.
Write Model-Specific Registers
WRMSR ; 0F 30 [PENT]
WRMSR writes the values in EDX:EAX to the processor Model-Specific Register (MSR) whose index is stored in ECX. See also RDMSR.
The input value loaded into the ECX register is the address of the MSR to be written to. The contents of the EDX register are copied to high-order 32 bits of the selected MSR and the contents of the EAX register are copied to low-order 32 bits of the MSR. Undefined or reserved bits in an MSR should be set to the values previously read.
The MSRs and the ability to read them with the WRMSR instruction were introduced into the IA-32 architecture with the Pentium processor. Execution of this instruction by an IA-32 processor earlier than the Pentium processor results in an invalid opcode exception #UD.
Example:
mov ecx MsrAddr
mov eax ValueLo
mov edx ValueHi
wrmsr
WBINVD
Usage: WBINVD Modifies flags: None
Flushes internal cache, then signals the external cache to write back current data followed by a signal to flush the external cache.
Write Back and Invalidate Cache
WBINVD ; 0F 09 [486]
WBINVD invalidates and empties the processor's internal caches, and causes the processor to instruct external caches to do the same. It writes the contents of the caches back to memory first, so no data is lost.
To flush the caches quickly without bothering to write the data back first, use INVD.
Example:
wbinvd
WAIT
Usage: WAIT Modifies flags: None
FWAIT
CPU enters wait state until the coprocessor signals it has finished its operation. This instruction is used to prevent the CPU from accessing memory that may be temporarily in use by the coprocessor. WAIT and FWAIT are identical.
Wait for Floating-Point Processor
WAIT ; 9B [8086]
FWAIT ; 9B [8086]
WAIT, on 8086 systems with a separate 8087 FPU, waits for the FPU to have finished any operation it is engaged in before continuing main processor operations, so that (for example) an FPU store to main memory can be guaranteed to have completed before the CPU tries to read the result back out.
On higher processors, WAIT is unnecessary for this purpose, and it has the alternative purpose of ensuring that any pending unmasked FPU exceptions have happened before execution continues.
All the FPU Instructions that would actually require a WAIT are directely encoded with a 09B Prefix. These Instructions are FSAVE, FINIT, FCLEX, FSTCW, FSTSW, FSTENV.
EXAMPLE:
wait
fwait
VERW
Usage: VERW src Modifies flags: ZF
Verifies the specified segment selector is valid and is ratable at the current privilege level. If the segment is writable, the Zero Flag is set, otherwise it is cleared.
Verify Segment Readability/Writability
VERR r/m16 ; 0F 00 /4 [286,PRIV]
VERW r/m16 ; 0F 00 /5 [286,PRIV]
VERR sets the zero flag if the segment specified by the selector in its operand can be read from at the current privilege level. Otherwise it is cleared.
VERW sets the zero flag if the segment can be written.
Example:
verw Label
VERR
Usage: VERR src Modifies flags: ZF
Verifies the specified segment selector is valid and is readable at the current privilege level. If the segment is readable, the Zero Flag is set, otherwise it is cleared.
Verify Segment Readability/Writability
VERR r/m16 ; 0F 00 /4 [286,PRIV]
VERW r/m16 ; 0F 00 /5 [286,PRIV]
VERR sets the zero flag if the segment specified by the selector in its operand can be read from at the current privilege level. Otherwise it is cleared.
VERW sets the zero flag if the segment can be written.
Example:
verr Label
LTJ UTJ
Likely Taken Jump / Unlikely Taken Jump
LTJ Jcc ; 02E [Pentium4 SSE2]
UTJ Jcc ; 03E [Pentium4 SSE2]
These instructions have no Mnemonic defined by Intel. Their Encodages are nothing but the older 'CS:' and 'DS:' Segment overrides.
These 'Branch Hint Prefixes' give information to the Processor about the more likely Code path that will be taken at a given Branching (Branch Predictions override).
These Prefixes can only be used with the Conditional Branch Instructions ( Jcc ).
See Intel Documentation for more info about the Branch Predictions mechanism.
UNPCKLPS
Usage: UNPCKLPS dest,src Modifies flags: None
Interleaved unpacking of the low-order single-precision floating-point values from the 'src' and the 'dest'.
Unpack and Interleave Low Packed Single-Precision FP Data
UNPCKLPS xmm1,xmm2/m128 ; 0F 14 /r [KATMAI,SSE]
UNPCKLPS performs an interleaved unpack of the low-order data elements of the source and destination operands, saving the result in xmm1. It ignores the lower half of the sources.
The operation of this instruction is:
dst[31-0] := dst[31-0];
dst[63-32] := src[31-0];
dst[95-64] := dst[63-32];
dst[127-96] := src[63-32].
The source operand can be an XMM register or a 128-bit memory location; the destination operand is an XMM register.
EXAMPLE:
unpcklps xmm1 Label
UNPCKLPD
Usage: UNPCKLPD dest,src Modifies flags: None
Interleaved unpacking of the low double-precision floating-point values from the 'src' operand and the 'dest' .
Unpack and Interleave Low Packed Double-Precision FP Data
UNPCKLPD xmm1,xmm2/m128 ; 66 0F 14 /r [WILLAMETTE,SSE2]
UNPCKLPD performs an interleaved unpack of the low-order data elements of the source and destination operands, saving the result in xmm1. It ignores the lower half of the sources.
The operation of this instruction is:
dst[63-0] := dst[63-0];
dst[127-64] := src[63-0].
The 'src' operand can be an XMM register or a 128-bit memory location; the 'dest' operand is an XMM register.
EXAMPLE:
unpcklpd xmm1 Label
UNPCKHPS
Usage: UNPCKHPS dest,src Modifies flags: None
Interleaved unpacking of the high-order single-precision floating-point values from the 'src' operand and the 'dest'.
Unpack and Interleave High Packed Single-Precision FP Values
UNPCKHPS xmm1,xmm2/m128 ; 0F 15 /r [KATMAI,SSE]
UNPCKHPS performs an interleaved unpack of the high-order data elements of the source and destination operands, saving the result in xmm1. It ignores the lower half of the sources.
The operation of this instruction is:
dst[31-0] := dst[95-64];
dst[63-32] := src[95-64];
dst[95-64] := dst[127-96];
dst[127-96] := src[127-96].
The 'src' operand can be an XMM register or a 128-bit memory location; the 'dest' operand is an XMM register.
EXAMPLE:
unpckhps xmm1 Label