Skip to content

Repository files navigation

rwkv-mobile

An inference runtime with multiple backends supported.

Goal:

  • Easy integration on different platforms using flutter or native cpp, including mobile devices.
  • Support inference using different hardware like Qualcomm Hexagon NPU, or general CPU/GPU.
  • Provide easy-to-use C apis
  • Provide an api server compatible with AI00_server(openai api)

Supported or planned backends:

  • WebRWKV (WebGPU): Compatible with most PC graphics cards, as well as macOS Metal. Doesn't work on Qualcomm's proprietary Adreno GPU driver though.
  • llama.cpp: Run on Android devices with CPU inference.
  • ncnn: Initial support for rwkv v6/v7 unquantized models (suitable for running tiny models everywhere).
  • Qualcomm Hexagon NPU: Based on Qualcomm's QNN SDK 2.42.0.
  • MLX: Running RWKV on Apple Silicon devices using Apple's MLX framework.
  • MediaTek Neuropilot7: Running RWKV on Dimensity 9300 NPU.
  • MediaTek Neuropilot9: Running RWKV on Dimensity 9500 NPU through Neuron Adapter/uSDK.
  • CoreML: Running RWKV with Apple Neural Engine. Based on Apple's CoreML framework.
  • To be continued...

How to build:

Build for Android:

git clone https://github.com/MollySophia/rwkv-mobile.git
cd rwkv-mobile
mkdir build && cd build
cmake .. -DENABLE_NCNN_BACKEND=ON -DENABLE_WEBRWKV_BACKEND=ON -DENABLE_QNN_BACKEND=ON -DENABLE_MTK_NP7_BACKEND=ON \
    -DENABLE_TTS=ON -DENABLE_VISION=ON -DENABLE_WHISPER=ON -DBUILD_EXAMPLES=ON -DENABLE_SERVER=ON \
    -DANDROID_ABI=arm64-v8a -DANDROID_PLATFORM=android-28 -DANDROID_NDK=$HOME/android-ndk-r25c \
    -DCMAKE_TOOLCHAIN_FILE=$HOME/android-ndk-r25c/build/cmake/android.toolchain.cmake -DCMAKE_BUILD_TYPE=Release -DCMAKE_POLICY_VERSION_MINIMUM=3.5 \
    -G Ninja
ninja

For Dimensity 9500/NP9 builds, use -DENABLE_MTK_NP9_BACKEND=ON. NP7 and NP9 can be enabled together; the MediaTek runtime archives are packaged as separate librwkv_mtk_np7.so and librwkv_mtk_np9.so libraries and loaded with dlopen. See MTK_NP9_SUPPORT.md for the current NP9 integration notes.

TODO:

  • Better tensor abstraction for different backends
  • Batch inference for all backends

About

Inference RWKV with multiple supported backends.

Resources

Stars

95 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages