Is your MATLAB code too slow for the job?
Does any of this sound familiar?
- A single run takes minutes or hours, so continuous measurement or large batch analysis is out of reach.
- You ran the Profiler and found the slow spots, but that did not lead to a fix.
- You tried parfor or gpuArray, and the gain was far smaller than expected.
- You want the code to run faster, but the numerical results must not change.
- The person who wrote the code has left, and nobody is comfortable touching it.
If so, our MATLAB Code Speed-Up Service may help. We start by reading your code and telling you where the bottleneck is and how much speed-up is realistic. If your budget is limited, we can restrict the work to the parts that consume the most time.

Why the usual fixes do not make it faster
When parallelization or a GPU does not help, the cause is usually not a lack of skill. In most cases, the remedy does not match the cause.
Slow MATLAB code falls into three broad categories: it is limited by the amount of computation, by memory bandwidth, or by the way it is implemented. Each calls for a different remedy. For example, if the program is limited by moving data through memory, porting it to a GPU as it is will not make it faster.
That is why we start with a diagnosis, before choosing a technique.
Our standard process
We follow the same four steps in every speed-up project.
- 1. Profiling the existing code: We locate the bottleneck by measurement, not by guesswork. We determine whether the code is limited by computation, by memory bandwidth, or by implementation, and choose the remedy that fits.
- 2. Implementation: We apply the standard techniques step by step: vectorization, parallelization (parfor and others), GPU computing (gpuArray), MEX functions and custom kernels in C/CUDA, and memory and I/O optimization. We implement several candidates, compare them by measurement, and report the rejected ones as well, with the reasons.
- 3. Numerical-equivalence check: Before we change anything, we save the output of your original code as reference data. We agree with you, in numbers, on the tolerance (bit-identical, double-precision relative error, or single-precision equivalent). We then compare the output before and after, and show in numbers that the results have not changed.
- 4. Timing: We measure the run time before and after, and report it together with the conditions, the data and the environment.
Case study: signal processing code for a measurement system
This is how the process above worked in a real project.
Client: The R&D division of a major company
Request: In a measurement technique under development, one MATLAB computation took about 70 seconds per measurement. The client wanted it faster. Changing the algorithm or losing accuracy was not acceptable.
Challenges:
- Continuous measurement requires a processing time of a few seconds or less.
- The algorithm and its output must stay the same.
What we did:
- Phase 1: We implemented and measured 13 candidates. Approximation, vectorization and a straightforward GPU port gave little gain (all reported with reasons). We adopted fine-grained parallelization with parfor (68.5 s to 32 s).
- Phase 2: We found that the bottleneck was not the computation itself but memory transfer of arrays on the order of 10 million points. We reduced the transfers by merging the FFT steps, keeping the data on the GPU, and writing a custom CUDA kernel that fuses several operations into one.
- The kernel is delivered as a PTX file, so no CUDA compiler is needed on the client's machines.
Results:

- 68.5 s to 1.46 s per measurement (about 47 times faster), on the same laptop PC
- No additional error relative to the reference data (single precision)
- We also provided a formula to estimate the processing time on other GPU models, as a guide for hardware selection
After delivery (in-house verification): We kept working on the same code internally with newer AI coding agents. The run time went from 1.46 s to 0.915 s, then to 0.74 s with identical results (about 93 times faster than the original), and to 0.44 s when small rounding differences are allowed (about 156 times faster). All figures were measured on the same laptop PC and checked against the same reference data.
What the AI did, and what it did not: AI coding agents rewrite the code, run MATLAB, measure the time and try again, so we can test far more candidates than before. They still need supervision. At one point the agent reported that the code slowed down to 2.5 to 4.6 s in consecutive runs. When we ran it five times ourselves, it took 1.44 to 1.50 s. The agent had measured right after starting MATLAB, and with a GPU monitoring tool running at the same time. We decide what counts as correct, and how it is measured.
Track record
- The R&D division of a major company: speed-up of signal processing code for a measurement system (the case above)
- The research division of a major electronic components manufacturer: speed-up and GPU porting of radar 3D imaging (3D point cloud and compressed sensing), provided as technical consulting (case page, in Japanese)
- A measurement instrument manufacturer: speed-up of an existing reconstruction process without changing its results, from about 26 minutes to about 3 minutes, as part of developing feature extraction software for 3D volume data (case page, in Japanese)
To discuss your code, please use the inquiry form at the bottom of this page.
Who we are
Systems Laboratories Corporation (Algorithm Development Center) is based in Yokohama, Japan. Since the center opened in 2013, we have developed algorithms and MATLAB software for more than 70 organizations: corporate R&D groups, national research institutes and universities. Our work covers signal processing, image processing, machine learning and data analysis. More than 40 case studies are published on this site (in Japanese).
Our MATLAB development service is listed in the MathWorks Connections Program as a consulting service (listing, in Japanese). MathWorks is the developer of MATLAB and Simulink.
Our lead engineer has been developing in MATLAB since 2006, and worked in the United States for eight years as a manager at a U.S. company. We work in English by email and video call. Our business hours are in Japan Standard Time (UTC+9), and we reply within one business day.
How it works
- 1. Inquiry: Tell us about your code through the form below. If you know them, include the current run time, your environment and your target.
- 2. Online meeting: We talk by Zoom or Teams. If you need a non-disclosure agreement, we sign it before you send anything.
- 3. Free diagnosis: You send the code and a sample dataset. We profile it and report where the bottleneck is and how much speed-up is realistic, typically within one week.
- 4. Fixed-price quote: If the diagnosis makes sense to you, we send a fixed-price quote. If not, you can stop there at no cost.
- 5. Speed-up work: We compare candidate techniques by measurement and report what we adopt and what we reject, with reasons.
- 6. Verification and delivery: We confirm in numbers that the results have not changed, and deliver the source code, a technical note and the numerical-equivalence report.
What you can expect from us
- We show in numbers that the results have not changed: We agree on the tolerance with you in advance, and show the agreement between the outputs before and after.
- You order only after you are convinced: Before any order, we analyze the code and present the likely bottleneck and the realistic speed-up. If you are not convinced, you can decline at that point.
- We tell you when it is not worth it: If the analysis shows that a speed-up is difficult, or does not justify the cost, we say so and explain why.
- The work can be split to fit your budget: You do not have to speed up everything at once. We can start with the most time-consuming part and proceed in phases.
- You can verify the delivery afterwards: We deliver the comparison of the outputs before and after, together with the timing conditions and results. On request, we also write a report with the before/after evaluation and a breakdown of each measure.
- We deliver in the form your environment needs: If MATLAB is not available where the code runs, we can deliver a standalone executable. We also offer technical consulting on an hourly basis.
- You are free to publish: You can present the results at conferences and in papers. You do not need to credit us. If a joint research arrangement suits you better, we are open to that.
Frequently asked questions
How much faster will my code run?
It depends on the code, the data and the environment, so we cannot promise a figure in advance. That is why we analyze the code before any order and present the likely bottleneck and a realistic estimate. When the diagnosis is right, as in the case above, the improvement can reach several tens of times.
Will the speed-up change my results?
This is the point we care about most. Before we start, we save the output of the original code as reference data, and agree with you on the tolerance in numbers. We compare the output after each change and show in numbers that the results have not changed.
How much does it cost?
The diagnosis is free. After the diagnosis, we send a fixed-price quote. Typical projects range from USD 5,000 to USD 20,000, depending on the size of the code and the target.
How do I pay?
By international bank transfer (SWIFT) against our invoice, in USD or JPY.
Can I ask for only part of the code?
Yes. We identify the most time-consuming parts by profiling, and can limit the work to those parts. We have carried out projects in two phases in this way.
I am not comfortable showing our code to an outside company.
We sign a non-disclosure agreement on request, in English. Algorithms and code are often sensitive, and it is common for us to sign an agreement before we receive any code.
Which MATLAB versions and toolboxes do you support?
We work with the release and the toolboxes you use. Tell us your MATLAB release, your toolboxes, and whether a GPU is available. If the code must run without Parallel Computing Toolbox, or without MATLAB, we plan the work for that from the start.
Do you also work on code that is not MATLAB?
Yes. We also speed up C/C++ and Python code. Combining MATLAB with C/CUDA (MEX functions and CUDA kernels) is one of our strengths.
We are in a different time zone. How does communication work?
We work by email and video call in English. We reply within one business day, Japan Standard Time. Meetings are usually held in the morning, Japan time, which is the afternoon or evening of the previous day in the United States.
Contact us
If your MATLAB code is too slow, tell us about it. Please fill in the form below. We reply within one business day (Japan Standard Time).
Systems Laboratories Corporation (Algorithm Development Center), Yokohama, Japan