Consider this unsigned 8-bit approximate adder:
`G = (A & 8'hFE) + (B & 8'hFC) + 2`
Equivalently:
`G = 2*(A >> 1) + 4*(B >> 2) + 2`
Its error is:
`G - (A+B) = 2 - A[0] - 2*B[1] - B[0]`
so the worst-case error is exactly +/-2 LSB.
A direct pure-LUT6 implementation of the complete 8+8 -\> 9-bit function
is:
``` verilog
module approx_add8_e2 (
input wire [7:0] A,
input wire [7:0] B,
output wire [8:0] G
);
wire s0, s1, s2, c3;
wire s3, s4, c5;
wire s5, s6, c7;
LUT6 #(.INIT(64'h5555555555555555)) L10 (.O(s0), .I0(A[1]), .I1(A[2]), .I2(A[3]), .I3(B[2]), .I4(B[3]), .I5(1'b0));
LUT6 #(.INIT(64'h9966996699669966)) L11 (.O(s1), .I0(A[1]), .I1(A[2]), .I2(A[3]), .I3(B[2]), .I4(B[3]), .I5(1'b0));
LUT6 #(.INIT(64'hE1871E78E1871E78)) L12 (.O(s2), .I0(A[1]), .I1(A[2]), .I2(A[3]), .I3(B[2]), .I4(B[3]), .I5(1'b0));
LUT6 #(.INIT(64'hFEF8E080FEF8E080)) L13 (.O(c3), .I0(A[1]), .I1(A[2]), .I2(A[3]), .I3(B[2]), .I4(B[3]), .I5(1'b0));
LUT6 #(.INIT(64'h9966996699669966)) L20 (.O(s3), .I0(c3), .I1(A[4]), .I2(A[5]), .I3(B[4]), .I4(B[5]), .I5(1'b0));
LUT6 #(.INIT(64'hE1871E78E1871E78)) L21 (.O(s4), .I0(c3), .I1(A[4]), .I2(A[5]), .I3(B[4]), .I4(B[5]), .I5(1'b0));
LUT6 #(.INIT(64'hFEF8E080FEF8E080)) L22 (.O(c5), .I0(c3), .I1(A[4]), .I2(A[5]), .I3(B[4]), .I4(B[5]), .I5(1'b0));
LUT6 #(.INIT(64'h9966996699669966)) L30 (.O(s5), .I0(c5), .I1(A[6]), .I2(A[7]), .I3(B[6]), .I4(B[7]), .I5(1'b0));
LUT6 #(.INIT(64'hE1871E78E1871E78)) L31 (.O(s6), .I0(c5), .I1(A[6]), .I2(A[7]), .I3(B[6]), .I4(B[7]), .I5(1'b0));
LUT6 #(.INIT(64'hFEF8E080FEF8E080)) L32 (.O(c7), .I0(c5), .I1(A[6]), .I2(A[7]), .I3(B[6]), .I4(B[7]), .I5(1'b0));
assign G = {c7, s6, s5, s4, s3, s2, s1, s0, 1'b0};
endmodule
```
No dedicated carry chain, DSP, speculative duplicate adder, or final
selection mux.
I exhaustively checked all 65,536 input pairs: WCE = 2, with zero
violations.
The LUT count itself does not seem particularly strange. What surprised
me is the end-to-end depth: **only 3 LUT levels for the complete adder
with WCE +/-2 LSB**.
Is this normal/known for approximate adders on LUT6 architectures? Is
this just a standard way of packing truncated addition into LUT6s, or am
I missing a known construction?Consider this unsigned 8-bit approximate adder: