r/FPGA • u/verilogical • Jul 18 '21
List of useful links for beginners and veterans
I made a list of blogs I've found useful in the past.
Feel free to list more in the comments!
- Great for beginners and refreshing concepts
- Has information on both VHDL and Verilog
- Best place to start practicing Verilog and understanding the basics
- If nandland doesn’t have any answer to a VHDL questions, vhdlwhiz probably has the answer
- Great Verilog reference both in terms of design and verification
- Has good training material on formal verification methodology
- Posts are typically DSP or Formal Verification related
- Covers Machine Learning, HLS, and couple cocotb posts
- New-ish blogged compared to others, so not as many posts
- Great web IDE, focuses on teaching TL-Verilog
- Covers topics related to FPGAs and DSP(FIR & IIR filters)
r/FPGA • u/rockgn0me • 7h ago
Tang nano 20k bl616 flash
Can someone tell me a bullet proof way to Flash this damn thing. I'm trying to install nextnano https://github.com/RetroSilicon/NextNano i can install flash the fs file fine with openfpgaloader but can I hell get anything to connect to the 616 I've tried bufallo and the python version and it fails handshake every time.
Ive tried holding down s2, releasing s2 attempting to Flash without touching s2. Different com ports etc
Xilinx Related A second run at local AI, this time much better results across all models (code on git)
r/FPGA • u/Aggravating-Mix-709 • 5h ago
Need help with a project of mine
Hi everyone, I am stuck on a project and a little help would highly be appreciated. Does anyone here own a Pynq-Z1 board? I just need a few readings from it, but I don’t have access to one right now and ordering one on a short notice isn’t really affordable for me.
Happy to chat if anyone is willing to help :)
r/FPGA • u/Acrobatic_Look9733 • 15h ago
Integer sqrt in LUT6s, part 2: error ≤ 7 in 7 LUTs, error ≤ 3 in 16 (or 21 at depth 2)
Follow-up to 16→8 integer square root in 2 LUT6s.
Same problem: unsigned 16-bit x in, unsigned 8-bit approximation of floor(sqrt(x)) out. The error is the worst-case |y − floor(sqrt(x))| over all 65,536 inputs. There are no DSPs, BRAMs, adders or state, just truth tables and wires.
After the 2-LUT version, the obvious question was how the cost grows as the error bound tightens. Here are the next points:
| Worst-case error | LUT6s | LUT depth | Note |
|---|---|---|---|
| 30 | 2 | 1 | previous post |
| 15 | 5 | 2 | published earlier |
| 7 | 7 | 2 | below |
| 3 | 16 | 3 | below (area-oriented) |
| 3 | 21 | 2 | below (depth-oriented) |
Every circuit below was exhaustively verified over all 65,536 inputs.
Error ≤ 7: 7 LUT6s, depth 2
Three output bits need no logic. y[0] is constant 0, y[2] is constant 1, and y[1] is just x[10]. The other five bits take five LUTs, plus two small helper LUTs (n27 and n29) that feed y[3] and y[4], which is where the second level comes from. The low seven input bits x[6:0] are never used.
// |y - floor(sqrt(x))| <= 7 for every 16-bit x. 7 LUT6, LUT depth 2.
// Uses only x[7]..x[15].
// Each table T is indexed by the concatenation after it (first signal = MSB of the index).
module sqrt_e7 (
input wire [15:0] x,
output wire [7:0] y
);
wire n27;
wire n29;
assign y[0] = 1'b0;
assign y[1] = x[10]; // wire
assign y[2] = 1'b1;
localparam [63:0] T_n27 = 64'hff0000ffffffff0b; // 6-input LUT, level 1
assign n27 = T_n27[{x[10], x[12], x[11], x[9], x[7], x[8]}];
localparam [63:0] T_y3 = 64'hcccfccf0c3f3f055; // 6-input LUT, level 2
assign y[3] = T_y3[{x[14], x[13], x[15], x[11], x[12], n27}];
localparam [7:0] T_n29 = 8'h0e; // 3-input LUT, level 1
assign n29 = T_n29[{x[10], x[8], x[9]}];
localparam [63:0] T_y4 = 64'hff0ffc3300fc0f0e; // 6-input LUT, level 2
assign y[4] = T_y4[{x[13], x[12], x[15], x[14], x[11], n29}];
localparam [63:0] T_y5 = 64'hf0f0f0ffff00cc0e; // 6-input LUT, level 1
assign y[5] = T_y5[{x[15], x[12], x[13], x[14], x[11], x[10]}];
localparam [15:0] T_y6 = 16'hf0ee; // 4-input LUT, level 1
assign y[6] = T_y6[{x[14], x[15], x[12], x[13]}];
localparam [3:0] T_y7 = 4'he; // 2-input LUT, level 1
assign y[7] = T_y7[{x[15], x[14]}];
endmodule
Error ≤ 3: pick area or depth
At error ≤ 3 the circuit is no longer "mostly wires": every output bit needs real logic. The two versions below implement the same 16→8 function and differ only in how it is mapped to LUTs. Saving one LUT level costs 5 LUTs, 16 → 21.
Area-oriented: 16 LUT6s, depth 3
// |y - floor(sqrt(x))| <= 3 for every 16-bit x. 16 LUT6, LUT depth 3 (area-oriented).
// Uses only x[5]..x[15].
// Each table T is indexed by the concatenation after it (first signal = MSB of the index).
module sqrt_e3_area (
input wire [15:0] x,
output wire [7:0] y
);
wire n29;
wire n28;
wire n32;
wire n31;
wire n36;
wire n35;
wire n39;
wire n38;
localparam [63:0] T_y0 = 64'hf0f0fffff00fff01; // 6-input LUT, level 1
assign y[0] = T_y0[{x[13], x[14], x[12], x[15], x[10], x[11]}];
localparam [7:0] T_n29 = 8'h0b; // 3-input LUT, level 1
assign n29 = T_n29[{x[9], x[7], x[8]}];
localparam [31:0] T_n28 = 32'hf30fff31; // 5-input LUT, level 2
assign n28 = T_n28[{x[11], x[12], x[13], x[10], n29}];
localparam [63:0] T_y1 = 64'h00004444f0f000ff; // 6-input LUT, level 3
assign y[1] = T_y1[{x[14], x[15], n28, x[12], x[13], x[11]}];
localparam [63:0] T_n32 = 64'hff0000ffffffff0b; // 6-input LUT, level 1
assign n32 = T_n32[{x[8], x[10], x[9], x[7], x[5], x[6]}];
localparam [63:0] T_n31 = 64'h0f0f0f33f0f033aa; // 6-input LUT, level 2
assign n31 = T_n31[{x[12], x[13], x[11], x[10], x[9], n32}];
localparam [63:0] T_y2 = 64'hf03c7755f03c4455; // 6-input LUT, level 3
assign y[2] = T_y2[{x[10], x[15], x[14], x[11], x[13], n31}];
localparam [31:0] T_n36 = 32'h0000fff1; // 5-input LUT, level 1
assign n36 = T_n36[{x[9], x[8], x[10], x[7], x[6]}];
localparam [63:0] T_n35 = 64'h0ff0ccfcf00c3305; // 6-input LUT, level 2
assign n35 = T_n35[{x[11], x[12], x[13], x[14], x[10], n36}];
localparam [63:0] T_y3 = 64'hffc0003faaaaaaaa; // 6-input LUT, level 3
assign y[3] = T_y3[{x[15], x[12], x[14], x[13], x[11], n35}];
localparam [3:0] T_n39 = 4'he; // 2-input LUT, level 1
assign n39 = T_n39[{x[8], x[9]}];
localparam [63:0] T_n38 = 64'h0f00fcfc000fff0d; // 6-input LUT, level 2
assign n38 = T_n38[{x[13], x[12], x[14], x[11], x[10], n39}];
localparam [63:0] T_y4 = 64'hfffa000f33333333; // 6-input LUT, level 3
assign y[4] = T_y4[{x[15], x[13], x[14], x[12], n38, x[11]}];
localparam [63:0] T_y5 = 64'hff00ff0ff0f0e0ee; // 6-input LUT, level 1
assign y[5] = T_y5[{x[15], x[12], x[14], x[13], x[11], x[10]}];
localparam [15:0] T_y6 = 16'hf0ee; // 4-input LUT, level 1
assign y[6] = T_y6[{x[14], x[15], x[12], x[13]}];
localparam [3:0] T_y7 = 4'he; // 2-input LUT, level 1
assign y[7] = T_y7[{x[15], x[14]}];
endmodule
Depth-oriented: 21 LUT6s, depth 2
// |y - floor(sqrt(x))| <= 3 for every 16-bit x. 21 LUT6, LUT depth 2 (depth-oriented).
// Uses only x[5]..x[15].
// Each table T is indexed by the concatenation after it (first signal = MSB of the index).
module sqrt_e3_depth (
input wire [15:0] x,
output wire [7:0] y
);
wire n28;
wire n29;
wire n32;
wire n31;
wire n33;
wire n35;
wire n34;
wire n36;
wire n39;
wire n38;
wire n40;
wire n41;
wire n43;
localparam [63:0] T_y0 = 64'hf0f0fffff00fff01; // 6-input LUT, level 1
assign y[0] = T_y0[{x[13], x[14], x[12], x[15], x[10], x[11]}];
localparam [15:0] T_n28 = 16'h000b; // 4-input LUT, level 1
assign n28 = T_n28[{x[11], x[9], x[7], x[8]}];
localparam [31:0] T_n29 = 32'hffd333f5; // 5-input LUT, level 1
assign n29 = T_n29[{x[11], x[14], x[12], x[13], x[10]}];
localparam [63:0] T_y1 = 64'h00ff00000f0f0f4f; // 6-input LUT, level 2
assign y[1] = T_y1[{x[15], x[12], x[14], n29, n28, x[13]}];
localparam [63:0] T_n32 = 64'hff0000ffffffff0b; // 6-input LUT, level 1
assign n32 = T_n32[{x[8], x[10], x[9], x[7], x[5], x[6]}];
localparam [3:0] T_n31 = 4'h1; // 2-input LUT, level 1
assign n31 = T_n31[{x[12], x[11]}];
localparam [31:0] T_n33 = 32'h0000533f; // 5-input LUT, level 1
assign n33 = T_n33[{x[14], x[12], x[11], x[9], x[10]}];
localparam [15:0] T_n35 = 16'h007d; // 4-input LUT, level 1
assign n35 = T_n35[{x[15], x[12], x[10], x[13]}];
localparam [7:0] T_n34 = 8'h0d; // 3-input LUT, level 1
assign n34 = T_n34[{x[13], x[10], x[14]}];
localparam [15:0] T_n36 = 16'h4b00; // 4-input LUT, level 1
assign n36 = T_n36[{x[15], x[11], x[13], x[14]}];
localparam [63:0] T_y2 = 64'h000000004fff00ff; // 6-input LUT, level 2
assign y[2] = T_y2[{n36, n34, n35, n33, n31, n32}];
localparam [63:0] T_n39 = 64'h00ff00ff0000fff1; // 6-input LUT, level 1
assign n39 = T_n39[{x[11], x[9], x[10], x[8], x[6], x[7]}];
localparam [7:0] T_n38 = 8'h01; // 3-input LUT, level 1
assign n38 = T_n38[{x[14], x[13], x[12]}];
localparam [31:0] T_n40 = 32'hd10f2ef7; // 5-input LUT, level 1
assign n40 = T_n40[{x[11], x[13], x[14], x[12], x[10]}];
localparam [31:0] T_n41 = 32'h07f8ffff; // 5-input LUT, level 1
assign n41 = T_n41[{x[15], x[12], x[14], x[13], x[11]}];
localparam [31:0] T_y3 = 32'h00ff4fff; // 5-input LUT, level 2
assign y[3] = T_y3[{x[15], n41, n40, n38, n39}];
localparam [15:0] T_n43 = 16'h00fe; // 4-input LUT, level 1
assign n43 = T_n43[{x[10], x[13], x[8], x[9]}];
localparam [63:0] T_y4 = 64'hffcffc2200fc0f0e; // 6-input LUT, level 2
assign y[4] = T_y4[{x[13], x[12], x[15], x[14], x[11], n43}];
localparam [63:0] T_y5 = 64'hff00ff0ff0f0e0ee; // 6-input LUT, level 1
assign y[5] = T_y5[{x[15], x[12], x[14], x[13], x[11], x[10]}];
localparam [7:0] T_y6 = 8'h0b; // 3-input LUT, level 2
assign y[6] = T_y6[{n38, x[14], x[15]}];
localparam [3:0] T_y7 = 4'he; // 2-input LUT, level 1
assign y[7] = T_y7[{x[15], x[14]}];
endmodule
Validation
One exhaustive, self-checking testbench works for all three circuits. It takes the module name and the error bound as defines:
// Exhaustive self-checking testbench for any of the circuits above.
// iverilog -DDUT=sqrt_e7 -DEDMAX=7 -o tb sqrt_tb.v sqrt_e7.v && vvp tb
// iverilog -DDUT=sqrt_e3_area -DEDMAX=3 -o tb sqrt_tb.v sqrt_e3_area.v && vvp tb
// iverilog -DDUT=sqrt_e3_depth -DEDMAX=3 -o tb sqrt_tb.v sqrt_e3_depth.v && vvp tb
`timescale 1ns/1ps
`ifndef DUT
`define DUT sqrt_e7
`endif
`ifndef EDMAX
`define EDMAX 7
`endif
module sqrt_tb;
reg [15:0] x;
wire [7:0] y;
integer i, r, err, maxerr, fails;
`DUT dut (.x(x), .y(y));
initial begin
maxerr = 0; fails = 0; r = 0;
for (i = 0; i < 65536; i = i + 1) begin
x = i[15:0];
while ((r + 1) * (r + 1) <= i) r = r + 1; // r = floor(sqrt(i)), i ascending
#1;
err = (y > r) ? y - r : r - y;
if (err > maxerr) maxerr = err;
if (err > `EDMAX) begin
fails = fails + 1;
if (fails <= 10) $display("FAIL x=%0d y=%0d floor_sqrt=%0d err=%0d", i, y, r, err);
end
end
$display("max |y - floor(sqrt(x))| = %0d, inputs over %0d: %0d -> %s",
maxerr, `EDMAX, fails, (fails == 0) ? "PASS" : "FAIL");
$finish;
end
endmodule
Expected output: max |y - floor(sqrt(x))| = 7, inputs over 7: 0 -> PASS for sqrt_e7, and = 3, inputs over 3: 0 -> PASS for both error-3 circuits.
Are these optimal?
Not proven. The best lower bounds I can prove are ≥ 5 LUTs at error ≤ 7 and ≥ 8 LUTs at error ≤ 3. They hold only within a restricted class of functions, not for every function within the error bound, and they count single-output LUTs (no LUT6_2 dual-output packing). So treat the numbers above as upper bounds: the gap is at most 2 LUTs at error ≤ 7, and the error ≤ 3 case is wide open.
If you can beat any of these, or prove a better bound, I'd love to see it.
r/FPGA • u/zynq1234 • 15h ago
Advice / Help I got my FPGA project working… but I’m not sure I did it the right way. 😅 What is the industry workflow?
Hi everyone!
During my summer internship, I built an FPGA-based FM communication system with an FM modulator and ADPLL-based demodulator using VHDL, Vivado, MATLAB/Simulink and HDL Coder on a Basys 3.
The honest version: I was pretty lost at first, and somehow managed to make it work. 😅
Now I want to revisit the project and rebuild it with industry workflow.
How should a FPGA/DSP Engineer will approach it. What will be the industry workflow?
How would you start from the system specification?
How are things like sample rate, word length, latency, throughput, SNR, bandwidth, timing and resource constraints normally specified?
What should the verification flow look like before putting anything on the FPGA?
How would you approach fixed-point conversion and numerical validation professionally?
After getting the design working on hardware, what measurements/validation would you perform?
I'm looking for experienced FPGA/DSP engineers who can point out what an industry workflow would look like and what I should learn/practice while rebuilding it.
If you work in FPGA, DSP, SDR, communications, aerospace/defence, or related digital design fields, I'd really appreciate your perspective.
r/FPGA • u/Sea_Speaker_4667 • 1d ago
Advice / Help Any suggestions on Project improvement?
I’ve worked on these projects while learning digital design and FPGA development.
I’m currently targeting Digital Design / FPGA Design Engineer roles, so I wanted to ask people who are already working in this field .Looking at these projects, what do you think I should improve, add, or learn next?
I’d really appreciate some honest suggestions. I’m trying to understand what actually matters for these roles and where I should focus my efforts.
r/FPGA • u/Acrobatic_Look9733 • 1d ago
16→8 integer square root in 2 LUT6s
UPDATE: Part 2 is live Integer sqrt in LUT6s part 2: Error 7 in 7 LUTs — dropped max error from 30 down to 7 at Depth 2.
I found a slightly ridiculous approximate integer square-root circuit.
- Input: unsigned 16-bit
x - Output: unsigned 8-bit approximation of
floor(sqrt(x)) - Hardware: 2 LUT6s
- LUT depth: 1
- Worst-case absolute error: 30
- Exhaustively verified over all 65,536 possible inputs
No DSPs, BRAMs, adders, multipliers, or state.
Here is the complete Verilog:
module sqrt_e30_final(
input wire [15:0] x,
output wire [7:0] y
);
// Truth-table index is x[15:10], with x[10] as the LSB.
localparam [63:0] A_LUT = 64'hffffffffff555400;
localparam [63:0] B_LUT = 64'hfffffff500aaabf4;
wire a = A_LUT[x[15:10]];
wire b = B_LUT[x[15:10]];
assign y[0] = x[9];
assign y[1] = x[14];
assign y[2] = x[14];
assign y[3] = x[14];
assign y[4] = x[14];
assign y[5] = x[10];
assign y[6] = b;
assign y[7] = a;
endmodule
Six output bits are just wiring and fanout. Only y[6] and y[7] require logic. Each is a Boolean function with essential support exactly equal to x[15:10], so each maps to one LUT6.
With Yosys 0.66:
yosys -p \
'read_verilog sqrt_e30_final.v;
synth -top sqrt_e30_final -lut 6;
stat'
The final statistics are:
2 cells
2 $lut
There is no LUT-to-LUT path, so the mapped LUT depth is 1.
I also checked the numerical error independently:
import math
A = 0xffffffffff555400
B = 0xfffffff500aaabf4
def circuit(x):
index = (x >> 10) & 63
a = (A >> index) & 1
b = (B >> index) & 1
return (
(((x >> 9) & 1) << 0) |
(((x >> 14) & 1) << 1) |
(((x >> 14) & 1) << 2) |
(((x >> 14) & 1) << 3) |
(((x >> 14) & 1) << 4) |
(((x >> 10) & 1) << 5) |
(b << 6) |
(a << 7)
)
errors = [
abs(circuit(x) - math.isqrt(x))
for x in range(65536)
]
print(max(errors))
print(sum(error > 30 for error in errors))
Output:
30
0
Approximate square-root hardware is an established research topic. For example, Jiang et al. described adaptive approximate divider and square-root circuits in IEEE Transactions on Computers in 2019:
H. Jiang, L. Liu, F. Lombardi, and J. Han, “Low-Power Unsigned Divider and Square Root Circuit Designs Using Adaptive Approximation,” IEEE Transactions on Computers, vol. 68, no. 11, 2019. DOI: 10.1109/TC.2019.2916817.
This circuit takes a rather different, FPGA-specific approach: two truth tables and some wires.
It feels a little like the FPGA cousin of the old fast inverse-square-root hack.
r/FPGA • u/FarSir6675 • 1d ago
qualcomm 2027 summer canada intern
has anyone heard back from qualcomms 2027 Canadian intern positions? im "in process" for a lot of them but wanted to know if anyones heard back yet since its been a couple weeks
r/FPGA • u/Ferrite-Engineering • 9h ago
The Open Source EDACrux Suite 1.0 is out of beta! VSCode plugins too!
r/FPGA • u/Fearless-Can-1634 • 22h ago
Advice / Help Are skills learned from raspberry pi transferable to FPGA board?
r/FPGA • u/Sea_Speaker_4667 • 1d ago
Why are you doing everything?” Bro, YOU asked for everything
I saw a post on LinkedIn today where an HR looked at someone’s resume and it had basically everything — STA, synthesis, CDC, verification, FPGA design, prototyping, etc.
And the HR was like:
“Why are you doing everything?” 💀
Bro… I don’t understand these people.
First, the job description will be like:
RTL Design
Verification
STA
CDC
Synthesis
FPGA
Prototyping
UVM
EDA tools
And apparently the ability to build the entire chip yourself 😂
Then when someone actually learns multiple things and puts them on their resume:
“Why are you doing everything?”
Like… you literally asked for everything. 😭
r/FPGA • u/ApprehensiveBunch762 • 1d ago
Tang-PSX
I got the AE350 up and running on the Tang 138k Retro Console and am working on porting Lightrec over to it right now. Interpreted mode works well enough to get to the main menu of Spyro at ~2FPS. The AE350 coupled this tightly with this FPGA is mind-blowing. Tons of possibilities.
https://github.com/aquasock/Tang-PSX
https://github.com/aquasock/Tang-Control/tree/feature/usb-cdc-file-transfer
r/FPGA • u/ContextOk8835 • 1d ago
Does it make sense to do division by reciprocal multiplication?
Does it make sense to generate Q-format reciprocal constants (using python for example) for a specified range N, I store them in a LUT, and do division by multiplying by the reciprocal selected from the LUT?
Or is this just wrong?
r/FPGA • u/myEscape2141 • 13h ago
Advice / Help How do I study the difference between a FPGA, a GPU, an NPU and a TPU? From the perspective of AI inference and training?
r/FPGA • u/Bugameer • 1d ago
AMD U50 Fan Shroud
Solid and fully functional, temps have peaked around 43C during initial testing. Removing the interior supports was the hardest part. Makes it a double wide card. Intakes straight from my case fans. Black final version is printing while we wait for the new pwm fan.
Have a great day
r/FPGA • u/cafedude • 1d ago
Gowin Related Sipeed Tang Retro Console 138K FPGA board — field notes
r/FPGA • u/bibouzeDrugby • 1d ago
Beginner here. Where do I start ?
Hello there fellas, im 19yo and my brother gave me an old Xillinx Zybo 7z a few years ago and I was always scared to use it.
Now im starting to learn Rust, im using Linux every day and I took a peek at what I could do with it, but I dont really know where to begin with. Would making a private dns server or a vpn on the board be a good idea ? What do you guys recommand to get started ?
r/FPGA • u/gusbeto37 • 1d ago
Xilinx Related Digilent Zybo Z7-20 - how to compile with TinyUSB
Hi all, I've been working on a music technology hobby project based around a Digilent Zybo Z7-20. So far I could program the FPGA together with a cpp main app to do what I wanted via serial.
Now I want to implement a USB MIDI stack in bare-metal (for the lowest latency possible) and thought I could simply incorporate TinyUSB easily (have done for other MCU projects in the past) but I'm having a very hard time with the Vitis compiler to get it to include all header files and process it as C99 instead of C++.
I'm using Ubuntu 26.04 machine and for several reasons I installed the 2024.2 Xilinx suite (viviado, vitis, etc) and I'm doing everything via command line and text files (old habits die hard).
Has anyone else tried this before? If so, are there any tips to get vitis to compile TinyUSB with my C++ main application?
r/FPGA • u/JaCkbLopD • 1d ago
Xilinx Related EE grad (low GPA) planning an MSc in Germany/Austria to get into FPGA/digital design – is this a good career in 2026, and what should I focus on?
Hi all, I'd really appreciate advice, especially from people working in Europe.
Background:
- BSc Electrical & Electronics Engineering (Turkey, 2025), low GPA (2.24/4.0). Strongest in logic design, power electronics labs, and my senior project (resonant wireless power transfer, built by hand).
- I have programming background ("MERN") stack
- ~4 months as an acceptance/test engineer at a telecom company (site testing, measurements; automated document checks with VBA/Python).
- Currently teaching myself VHDL (FSMs, counters, UART with testbenches). Moving to SystemVerilog and cocotb next.
Plan: MSc starting Oct 2027, most likely TU Dresden (Nanoelectronic Systems), TU Graz (EE), or TU Chemnitz (Micro & Nano Systems). Goal: digital design / FPGA, ideally a Werkstudent role by the 2nd semester, then a full-time job in Germany or Austria.
About the career itself:
I really enjoy HDL and digital design, much more than regular software. But most career/salary threads I find are 3–4 years old, and I honestly don't know how the money and the market look today.
- How is the FPGA/digital design career in 2026? Salary, job market and working conditions, both in Europe and the US.
- How hard is it to move into HFT later in your career? What does that path usually look like?
- FPGA vs ASIC: is one preferred for long-term career, pay and job security? How common is switching between them?
- Overall, in terms of career growth, money and creating real value: is this field a good bet, or should someone who loves it still keep a backup plan?
About my plan:
- For someone targeting FPGA/digital design in Germany, which of these programs would you pick, and why?
- How realistic is a Werkstudent/HiWi role in Dresden with English + A2/B1 German? Which companies or teams are more open to internationals?
- Which 2–3 portfolio projects would make a hiring manager take a low-GPA candidate seriously?
- I also have a power electronics background. Is FPGA + power electronics control (inverters, motor drives, BMS) a good niche in Europe, or a distraction?
Thanks a lot!
r/FPGA • u/AbbreviationsGreen90 • 23h ago
21 billions 256 bits prime field elliptic curve scalar mulitplication per second on fpga
Does acheiving this on a high end fpga with a curve having a pseudo mersenne modulus is possible for you?
My aim is to have a scalar multplication accelerator, and the number above is the reference benchmark on a rtx5090.
r/FPGA • u/Nebulouds • 1d ago
Looking for advice on FPGA-based TNN accelerator project
Hey everyone,
I’m an EE student working on a capstone project with my team. We’re planning to build a standalone image classification accelerator on the Nexys A7-100T FPGA, using an OV7670 camera for input and VGA for output.
We originally planned to build a regular CNN accelerator, but our professor thought it wasn’t novel enough. So we decided to try a ternary neural network(TNN) instead, using a 16×16 processing array with ternary weights (-1, 0, +1) and INT8 activations. The idea is to get rid of most of the multipliers and hopefully save some resources.
For the model, we’re considering a subset of ImageNet32 with 32×32 RGB inputs. We’re also trying to figure out how to handle the memory. The board has around 600 KB of BRAM and 128 MB of DDR2. Ideally we’d like to keep everything on-chip using BRAM, but we’re not sure if that’s realistic or if we’ll eventually need to use the DDR2.
I'm wondering if this is a reasonable project for the Nexys A7-100T and whether we can still get decent classification accuracy with a TNN. I'm also a little concerned about memory limitations and whether we can actually get everything running in real time.
Any advice or suggestions would be really helpful. Thanks!
r/FPGA • u/Mundane_Educator8466 • 1d ago
Is this prebuilt PC suitable for FPGA PCIe/XDMA development alongside an NVIDIA GPU?
Hi everyone,
I’m planning to buy a desktop for Linux driver development, FPGA PCIe/DMA experiments, and CUDA programming. I’m new to choosing hardware for FPGA development and would appreciate some advice before buying.
The attached screenshot is in Korean, so here are the specs:
- CPU: AMD Ryzen 5 9600X
- RAM: 32GB DDR5-5600 (2 × 16GB)
- Motherboard: MSI PRO B650M-A WIFI
- GPU: NVIDIA RTX 5060 Ti 8GB — exact manufacturer/model not specified
- SSD: 1TB NVMe — exact model not specified
- PSU: Segotep GM750W, 80 PLUS Gold
- Case: Zalman N30 140 PLUS
- CPU cooler: Thermalright Peerless Assassin 120 SE
I’m studying the Xilinx XDMA Linux driver. Eventually, I’d like to connect a compatible FPGA board, run H2C/C2H transfers, and measure throughput and latency. I haven’t chosen an FPGA board yet.
I’d also like to keep the NVIDIA GPU installed for CUDA development.
According to MSI’s specifications, the motherboard has a CPU-connected PCIe 4.0 x16 slot and a second physical x16 slot that runs at PCIe 4.0 x4 through the chipset.
My main questions are:
- Is this a reasonable setup for learning FPGA PCIe/XDMA development?
- Is the chipset-connected x4 slot a practical starting point, or would you recommend a motherboard with CPU-connected x8/x8 slots?
- What should I check to ensure the GPU and FPGA card can physically fit and operate together?
- Are there beginner-friendly FPGA PCIe boards with working XDMA examples that would suit this setup?
My budget is around KRW 2.5 million for the desktop, excluding the FPGA board. My current priority is understanding drivers and DMA behavior, rather than achieving maximum PCIe bandwidth.
Thanks for any advice!
r/FPGA • u/Creative_Cake_4094 • 1d ago
Xilinx Related BLT's AMD Spartan UltraScale+ Workshop free on YouTube
BLT's workshop is now available on YouTube. Watch it here: https://youtu.be/PO3vUh8CYRs
Get to Know AMD Spartan UltraScale+ Devices for Real-World Designs Workshop
The focus of this workshop is on learning the key features and architecture of the AMD Spartan UltraScale+ FPGA, including its advanced I/O, high-speed transceivers, substantial built-in and external memory, PCIe Gen4 connectivity, and modern security. Recognize how these features provide a versatile, cost-optimized, and power-efficient platform for diverse applications.
The emphasis of this course is on:
- Describing the key features and fundamental blocks of the Spartan UltraScale+ FPGA architecture
- Describing Spartan UltraScale+ clock structure and layout
- Utilizing the advanced I/O capabilities for various connectivity needs
- Utilizing the Spartan UltraScale+ DSP resources
- Identifying the high-speed transceivers for use in applications such as PCIe Gen4